all_lessons/ World Models/ 18 · what "trained" meanslesson 18 / 31

What "trained" means for a simulator

A video model and a world model can have identical architectures, identical losses, and identical validation curves — and one of them is useless. The difference is not quality. It is that a simulator has to answer questions, and questions have a different acceptance test than pictures.

Where we are
Lesson 17 located the object: the forward conditional p(o′|o,a), trained from tuples whose action column costs four orders of magnitude more than its observation column. Before we spend that money we need to know what we are buying. "Trained" cannot mean "validation loss stopped falling," because the loss is measured on a task nobody will ever ask the model to perform.
Forced by 17The object is the forward conditional, and its training data is expensive. This stepDefine acceptance: controllability, consistency, calibration — three tests no pixel metric implies. Forces 19Acceptance is stated per-quantity, so we must choose which quantity to predict.

1 · The only use that matters is being asked a question

Here is the thing a planner does with your model, in full. It holds a history. It invents an action sequence it has never taken. It asks: if I did this, what would happen? It compares the answer against other invented sequences. It commits to the first action of the winner and throws the rest away.

Read that again and notice what is absent. Nobody ever asks the model to reproduce a recorded clip. Every single query is counterfactual — an action sequence that is not in the data — and every answer is consumed as a comparison, not as an image. So the training objective (reconstruct what actually happened next) and the deployment query (rank things that never happened) are different tasks that merely share a parameter vector.

Write the deployment contract down, because it is the specification the rest of this part builds against:

The simulator contract
Given a history o≤t and a proposed action sequence at:t+H, return a distribution over futures p(ot+1:t+H | o≤t, at:t+H) such that (i) changing the actions changes the answer in the right direction, (ii) the answer stays a coherent single world for H steps, (iii) the spread it reports matches the spread that actually occurs — and all of it inside the caller's latency budget.

Those three clauses are not stylistic. They are three separable failure modes, each with its own cause and its own fix, and they map one-to-one onto the three acceptance tests.

2 · Test one: controllability

Does the action column do anything? The honest measurement is a paired intervention: hold the history fixed, feed two different action sequences, and measure how much the futures differ.

C = Eo≤t[ d( f(o≤t, a), f(o≤t, a′) ) ] / E[ d( f(o≤t, a), ot+1:t+Htrue ) ]

Read the ratio: the numerator is how much your action mattered, the denominator is how much error you have anyway. If C ≈ 0 the model is a very expensive video player — it produces plausible continuations that ignore the controller entirely. This failure is common, it is invisible to every fidelity metric, and lesson 22 shows exactly why gradient descent prefers it.

Why the loss actively rewards ignoring you
Suppose 80% of the next frame is predictable from history alone (it is — backgrounds persist, motion is smooth). Then a model that ignores the action forfeits at most 20% of the achievable likelihood, and it gains a much easier optimisation problem. Under a pixel-weighted loss, "ignore the handle" is not a bug the model stumbles into. It is a good trade, and the model takes it.

3 · Test two: consistency

Does it stay one world? Two distinct things break here, and conflating them wastes months.

drift · compounding
Errors feed forward
Each step's small mistake becomes the next step's input. Governed by ek+1≈J·ek+δ — derived in lesson 6. Fixed by the schedule (lesson 21).
forgetting · permanence
Things stop existing
Look away, look back, the chair is a different chair. Nothing to do with drift; it is a memory-window problem. Fixed by context and retrieval (lesson 25).

Measure them separately or you will "fix" the wrong one. Drift shows up as a monotone error-vs-horizon curve. Forgetting shows up as a revisit test: leave a region for k steps, return, and score whether the scene matches what you left. A model can have beautiful drift curves and total amnesia, and vice versa.

4 · Test three: calibration

The model reports a spread. Is the spread true? If the model says "the mug tips over 10% of the time" and it tips 60% of the time, a planner that trusts it will pick the plan that quietly relies on the 90%. Calibration is what makes the reported uncertainty load-bearing.

The measurement is the standard one, applied to rollouts rather than labels: bin predicted probabilities, compare to observed frequencies, sum the gaps.

ECE = Σb (nb/n) · | p̂b − freqb |

And note which architectural choice this test is sensitive to. A model with a deterministic head has no spread to be right about — its ECE is undefined or trivially terrible. Calibration therefore forces a distributional head, which is lesson 24, which costs sampling steps, which is lesson 29. The tests chain into the syllabus.

Fidelity rises smoothly; acceptance does not
Sweep training compute and change what the objective actually included. Pixel fidelity improves with compute under every setting. The other three tests only move if the corresponding mechanism was trained — and consistency additionally decays with the horizon you demand.
Interactive comparison of four acceptance scores.
pixel fidelity
—
controllability
—
consistency@H
—
calibration
—
tests cleared
—

The widget's point is narrow and important: with the objective set to pixel loss only, you can push compute to the maximum and still clear exactly one of four tests. Fidelity is not a proxy for acceptance. It is one of four coordinates, and it is the one that matters least.

5 · Why FVD and friends cannot stand in

Distributional video metrics score whether your generated clips look like the clip distribution. That is a genuine measurement, and it is the wrong one here for three structural reasons.

They are unconditional or weakly conditional — an action-ignoring model scores wellmisses test 1
They are computed on short clips — drift and permanence live past the clipmisses test 2
They score marginal realism, not whether reported probabilities match frequenciesmisses test 3
What they do measure, correctlytest 0: does it look right

Keep the fidelity metric. It catches real regressions cheaply and it is the only one you can compute every hundred steps. Just never let it be the release gate. The gate is the three tests, and lesson 30 builds the harness that runs them without lying to you.

6 · The acceptance test is a budget, not a wish

One more clause hides in the contract: inside the caller's latency budget. This is not a footnote; it changes the training target. A model that answers in 3 seconds per step is a research artifact. A model that answers in 20 ms is an environment you can act inside.

Work the arithmetic backwards from the loop. A 20 Hz controller has 50 ms per tick. If your generative head needs N sampling steps at ℓ milliseconds each, and the planner wants K candidate rollouts of H frames:

latency = N · ℓ · H · K   ≤   50 ms

With ℓ = 9 ms, that leaves N·H·K ≤ 5. Five. That number is why lesson 29 exists, and why "just add more sampling steps" is not available as an answer to lesson 24's problem. Every acceptance test is scored under this constraint or it is not scored at all.

Takeaway
"Trained" for a simulator means it clears three tests that fidelity does not imply: controllability (paired intervention — does changing the action change the future, measured against your baseline error), consistency (drift and permanence, measured separately, over the horizon you actually demand), and calibration (reported spread matches observed frequency), all inside the caller's latency budget. Pixel loss reliably improves with compute while all three stay flat, because ignoring the handle is a good trade under a pixel-weighted objective. Keep fidelity as a cheap regression alarm; never as the gate.

Where this points next

Each test is stated about a quantity — the thing whose error we measure, whose spread we calibrate, whose consistency we check. We have been vague about what that quantity is: raw pixels? discrete tokens? an abstract latent? The choice is not cosmetic. Under a fixed bit budget it decides which parts of the world get represented at all, and lesson 19 shows — with a rate-distortion calculation you can run yourself — that a pixel-weighted loss spends almost nothing on the millimetres that decide the task.

Interview prompts