What "trained" means for a simulator
A video model and a world model can have identical architectures, identical losses, and identical validation curves — and one of them is useless. The difference is not quality. It is that a simulator has to answer questions, and questions have a different acceptance test than pictures.
1 · The only use that matters is being asked a question
Here is the thing a planner does with your model, in full. It holds a history. It invents an action sequence it has never taken. It asks: if I did this, what would happen? It compares the answer against other invented sequences. It commits to the first action of the winner and throws the rest away.
Read that again and notice what is absent. Nobody ever asks the model to reproduce a recorded clip. Every single query is counterfactual — an action sequence that is not in the data — and every answer is consumed as a comparison, not as an image. So the training objective (reconstruct what actually happened next) and the deployment query (rank things that never happened) are different tasks that merely share a parameter vector.
Write the deployment contract down, because it is the specification the rest of this part builds against:
Those three clauses are not stylistic. They are three separable failure modes, each with its own cause and its own fix, and they map one-to-one onto the three acceptance tests.
2 · Test one: controllability
Does the action column do anything? The honest measurement is a paired intervention: hold the history fixed, feed two different action sequences, and measure how much the futures differ.
C = Eo≤t[ d( f(o≤t, a), f(o≤t, a′) ) ] / E[ d( f(o≤t, a), ot+1:t+Htrue ) ]Read the ratio: the numerator is how much your action mattered, the denominator is how much error you have anyway. If C ≈ 0 the model is a very expensive video player — it produces plausible continuations that ignore the controller entirely. This failure is common, it is invisible to every fidelity metric, and lesson 22 shows exactly why gradient descent prefers it.
3 · Test two: consistency
Does it stay one world? Two distinct things break here, and conflating them wastes months.
Measure them separately or you will "fix" the wrong one. Drift shows up as a monotone error-vs-horizon curve. Forgetting shows up as a revisit test: leave a region for k steps, return, and score whether the scene matches what you left. A model can have beautiful drift curves and total amnesia, and vice versa.
4 · Test three: calibration
The model reports a spread. Is the spread true? If the model says "the mug tips over 10% of the time" and it tips 60% of the time, a planner that trusts it will pick the plan that quietly relies on the 90%. Calibration is what makes the reported uncertainty load-bearing.
The measurement is the standard one, applied to rollouts rather than labels: bin predicted probabilities, compare to observed frequencies, sum the gaps.
ECE = Σb (nb/n) · | p̂b − freqb |And note which architectural choice this test is sensitive to. A model with a deterministic head has no spread to be right about — its ECE is undefined or trivially terrible. Calibration therefore forces a distributional head, which is lesson 24, which costs sampling steps, which is lesson 29. The tests chain into the syllabus.
The widget's point is narrow and important: with the objective set to pixel loss only, you can push compute to the maximum and still clear exactly one of four tests. Fidelity is not a proxy for acceptance. It is one of four coordinates, and it is the one that matters least.
5 · Why FVD and friends cannot stand in
Distributional video metrics score whether your generated clips look like the clip distribution. That is a genuine measurement, and it is the wrong one here for three structural reasons.
Keep the fidelity metric. It catches real regressions cheaply and it is the only one you can compute every hundred steps. Just never let it be the release gate. The gate is the three tests, and lesson 30 builds the harness that runs them without lying to you.
6 · The acceptance test is a budget, not a wish
One more clause hides in the contract: inside the caller's latency budget. This is not a footnote; it changes the training target. A model that answers in 3 seconds per step is a research artifact. A model that answers in 20 ms is an environment you can act inside.
Work the arithmetic backwards from the loop. A 20 Hz controller has 50 ms per tick. If your generative head needs N sampling steps at ℓ milliseconds each, and the planner wants K candidate rollouts of H frames:
latency = N · ℓ · H · K ≤ 50 msWith ℓ = 9 ms, that leaves N·H·K ≤ 5. Five. That number is why lesson 29 exists, and why "just add more sampling steps" is not available as an answer to lesson 24's problem. Every acceptance test is scored under this constraint or it is not scored at all.
Where this points next
Each test is stated about a quantity — the thing whose error we measure, whose spread we calibrate, whose consistency we check. We have been vague about what that quantity is: raw pixels? discrete tokens? an abstract latent? The choice is not cosmetic. Under a fixed bit budget it decides which parts of the world get represented at all, and lesson 19 shows — with a rate-distortion calculation you can run yourself — that a pixel-weighted loss spends almost nothing on the millimetres that decide the task.
Interview prompts
- Why is next-frame validation loss a poor definition of "trained" for a world model? Every deployment query is a counterfactual action sequence consumed as a comparison, not a reconstruction of a recorded clip, so the training and deployment tasks merely share parameters. (§1)
- How do you measure controllability without a downstream policy? A paired intervention: fix the history, feed two action sequences, and divide the resulting divergence by the model's own baseline prediction error. (§2)
- Give the argument that a pixel-weighted loss rewards ignoring actions. If most of the next frame is predictable from history, the likelihood forfeited by ignoring the action is small while the optimisation gain is large, so it is a favourable trade. (§2)
- Distinguish drift from forgetting, and give the test for each. Drift is error feeding forward, governed by e(k+1)≈J·e(k)+δ, seen in an error-versus-horizon curve; forgetting is a memory-window failure, seen in a leave-and-revisit test. (§3)
- Which acceptance test forces a distributional head, and what does that cost? Calibration — a deterministic head has no spread to be correct about — and it costs sampling steps, which the latency budget then constrains. (§4, §6)
- A model's FVD improves 30% and its planner win-rate is unchanged. Explain. FVD scores marginal short-clip realism, so it is blind to controllability, to long-horizon drift and permanence, and to calibration — the three things the planner consumes. (§5)
- Derive how many sampling steps a 20 Hz control loop can afford at 9 ms per step. 50 ms per tick divided by 9 ms gives N·H·K ≤ 5 across steps, horizon, and candidate count combined. (§6)