Why predict? The world-model contract
A reconstruction explains a scene that was recorded. It has no actions and no future, so it cannot say what would happen if the agent did something. This lesson asks what an agent gains from a model that can, and puts a price on it. In a small courtyard a ball must be tapped once into a goal. Trying nudges for real takes dozens of trials; trying them in one's head takes none, provided a function says what a nudge does. We build that planner with a perfect model, measure the exchange rate between imagined and real trials, name the four things any model of consequences needs, and then hand the planner what a real agent is given: a noisy reading, and nothing at all while the ball is behind a curtain. The model is exact. The input is not.
New idea: a world model turns computation into trials: it answers "if I do this, what happens next?" so that an agent can try many actions in its head and one in the world. The answer needs four ingredients, a state, an action, a transition and a score, and the rest of the series is organised by how each of them fails.
Forces next: A model that answers "if I push this way, where does the ball end up?" makes thinking cheaper than trying: one push in the real world instead of dozens, provided the model is handed the true state of the ball. An agent is never handed the state. It sees a noisy position, and nothing at all while the ball is behind the curtain. What should the agent carry in its head instead of the state it cannot see?
1 · A courtyard, and the price of trying
Every lesson in this series happens in one small world, the Courtyard: a floor 8 m wide and 5 m deep (x to the right, y up) with a ball of radius 0.1 m on it. Time moves in steps of Δt = 0.1 s. Friction removes 3.4 % of the ball's velocity each step (a rate γ = 0.35 s−1), so on open floor a ball pushed at speed v rolls to a stop a distance v/γ from where it was pushed, 94 % of the way there after 8 s. The walls return 90 % of the speed on impact. A curtain between x = 2.6 m and 3.4 m hides the ball from the sensor, and a goal disc of radius 0.5 m sits at (6.6, 2.5). The ball's state is four numbers, s = (x, y, vx, vy). An action is an impulse a = (ax, ay) added to the velocity at the start of a step.
The task, called the nudge, is this. A launcher fires the ball from the left wall at 2–3 m/s in a direction that varies from launch to launch. The agent may add one impulse of size at most 3 m/s at t = 1.2 s (step 12); after that everything coasts for 80 more steps, which is 8 s. It succeeds if the ball ends inside the goal. At the moment of the nudge the ball is rolling at about 1.6 m/s, and in 80 % of launches it is behind the curtain.
Suppose the agent knows no physics and learns by trial and error. A trial is a real event: launch, nudge, wait 8 s, look. Choose the nudge at random and the chance that a trial succeeds is some p. We can measure it by trying 512 random nudges on each of 60 launches (all computed, which is the point of the next section) and counting the successes: p = 2.6 %. The number of trials until the first success is then geometric with mean 1/p ≈ 39 and a long tail: the chance of no success in 100 blind trials is (1 − p)100 = 7.4 %. In the widget below the 60 launches average 35.8 trials; the gap to 39 is sampling noise.
A smarter learner would use what each trial shows. It keeps the best nudge so far, tries perturbations of it, widens the step when a perturbation wins and narrows it when not (a (1+1) evolution strategy, restarted after 25 stalled trials). It needs a median of 16.5 real trials per launch and a mean of 21. Learning from outcomes roughly halves the bill. It does not change its order of magnitude, which is dozens.
Nothing carries over, either. The next launch has a different state, so the nudge that worked is the wrong one, and the learner begins again. Over 100 launches blind trial and error costs about 3894 real trials.
2 · Replace trials with thought
Suppose we had a function M that takes the state at the nudge and a nudge and returns where the ball ends up: if I do this, what happens? Then a nudge need not be tried to learn what it does; it can be computed. Draw K candidate nudges, compute the ending of each, keep the one that ends closest to the goal, and do only that one for real. For the Courtyard we can give the agent the best M there is, the simulator itself. Later lessons take that gift away; here it isolates what a model is for.
How many imagined nudges? Each candidate succeeds independently with probability p, so at least one of K does with probability
P(K) = 1 − (1 − p)K, and P(K) ≥ 0.99 once K ≥ ln 0.01 / ln(1 − p) ≈ 4.6 / p
For p = 2.6 % that is 177 candidates. Imagination buys certainty: the 39 real trials of blind search are replaced by one real trial and about 177 imagined ones.
When does that pay? Let cr be the cost of a real trial and ci the cost of an imagined one. Trial and error costs cr/p in expectation; planning costs cr + K·ci. Planning is cheaper when K·ci < cr(1/p − 1), that is
cr / ci > K·p / (1 − p) = 4.7 for K = 177
A model pays whenever imagining a nudge is more than about five times cheaper than trying it. An imagined nudge here is 81 simulator steps, microseconds on a laptop; a real one is 8 s of world time and a reset. The ratio is in the millions. Real systems show the same exchange on harder tasks: PlaNet (Hafner et al., 2019) reports beating the model-free methods A3C, and in some cases D4PG, with on average 200 times less environment interaction, and PETS (Chua et al., 2018) matches the asymptotic performance of Soft Actor-Critic and PPO on the half-cheetah task with 8 and 125 times fewer samples.
One more property is easy to miss. M takes the state as an argument, so the same M serves the next launch and the one after. Trial and error learns the answer to one situation. A model learns the question.
3 · The oracle planner
The widget runs exactly this. For each of 60 fixed launches it rolls the ball to the nudge time, draws 512 candidate nudges, and computes the ending of every candidate with the simulator. A planner with budget K looks at the first K candidates, takes the one with the smallest imagined miss, and the nudge is then executed in the real world, which here means the same simulator started from the true state. The success curve on the right is that experiment for every K from 1 to 512.
What to try. Start with the planner handed the true state and K = 128. It succeeds on 98.3 % of the 60 launches, where 1 − (1 − p)128 predicts 96.4 %. Slide K down: K = 16 gives 35 % (formula 34 %) and K = 1, a blind guess, 3.3 % (2.6 %). Each green path in the fan is an imagined nudge that would have won, the amber one is the planner's pick, and the heavy black path lands where the amber one did, because the model is exact. Step through other launches with the launch slider: the same planner, a different state, one real push each time. Now switch to what the sensor gives. The dashed ring is the state the planner now believes. On launch 28 the ball is behind the curtain, so the last two readings are older than the nudge; the believed position is 0.31 m off and the believed velocity 1.45 m/s off. The planner imagines landing 0.11 m from the goal's centre, and the real push lands 3.27 m away. Raise K to 512: the imagined misses keep shrinking (averaged over the 60 launches the planner's imagined miss is 0.12 m) and the real miss does not (2.44 m on average); not one of the 60 launches succeeds, 0 %. Finally move the sensor noise down. At 0 the planner is as good as it was with the true state (98.3 %), although on 80 % of the launches the ring then sits behind the ball (§5 says why that costs nothing); at 5 mm it is 91.7 %, at 1 cm 50 %, at 2 cm 23 %, at 5 cm 6.7 %, and at the Courtyard's real 100 mm 0 % (all at K = 128). For 90 % success the sensor would have to be accurate to about 5.5 mm, which is 18 times better than the one the agent has.
4 · The contract
The planner above was handed four things, and it is worth listing them, because a "world model" in the narrow sense is only one of them.
| Ingredient | What we handed the planner | What a real agent has to supply | Where the series goes |
|---|---|---|---|
| a state s | the true four numbers | something built from noisy, partial readings, or from pictures | lessons 2–4 |
| an action a | any of K freely chosen nudges | actions it has seen, or may only be able to take where it has tried them | lesson 7 |
| a transition M(s, a) | the exact simulator | a learned function: wrong in places, one future or many, chained again and again | lessons 4–6 |
| a score | distance of the ending from the goal | a value for imagined futures, and a way to know the model is good enough for the decision | lessons 8–9, 16 |
A transition is meaningless on its own: it needs a state to start from, an action to vary, and a score to compare the endings by. Each of the next fifteen lessons takes one of these gifts away, or the assumption behind it, and measures what breaks.
The exam also changes. The 3D track graded a model of a scene by predicting a photograph from a camera nobody used, and that camera was an action which left the world alone. Here the action acts. The exam is predict what the world does when I do this, and in the end are the decisions it supports better ones.
5 · Hand the planner what an agent gets
Everything in the oracle planner is exact except one thing, the input. A real agent is not handed the ball's state. Its sensor reports the position with noise, σO = 0.1 m on each axis, and reports nothing while the ball is behind the curtain. The velocity, on which the future depends even more than on the position, is not measured at all. The simplest thing to do is to infer it from the last reading z1 and the one before it, z2, taken Δt apart:
v̂ = (z1 − z2) / Δt, so Var(v̂) = 2σO² / Δt² and error(v̂) = √2·σO / Δt = 1.41 m/s
That error is nearly as large as the ball's speed (1.6 m/s). A velocity error δv costs far more than a position error, because the ball keeps going: it moves the stopping point by δv/γ = 2.86 m for every m/s, where a position error of one metre moves it by one metre. An error of 1.41 m/s on each axis therefore moves the ending by 4.0 m on each axis, against a goal 1 m across. On the widget's 60 launches (K = 128), a planner handed the true state with only the position wrong by the sensor's 0.1 m still succeeds on 96.7 %; handed the true state with only the velocity wrong by what two readings give, it succeeds on 8.3 %. The planner imagines worlds that are metres away from the real one and, being a good optimiser, finds a nudge that wins in the wrong world. Adding candidates makes it better at winning there (imagined miss 0.12 m at K = 512) and no better at winning here (real miss 2.44 m).
Behind the curtain the two readings are the last ones taken before the ball vanished, so the planner starts its imagination from where the ball was, not where it is. By itself that costs nothing: a coasting ball's stopping point, x + v/γ, does not change while nothing touches it, so an old state points at nearly the same ending as the current one. With noise-free readings the planner scores 98.3 %, as with the true state. The damage is the noise.
This is not a failure of the planner, nor of the model, which is exact. It is the first failure of the contract: the transition is a function of a state, and the agent has none.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
A model that answers "if I push this way, where does the ball end up?" turns computation into trials: one push in the world where blind search needs about 39 and a good learner about 21. It does so for any launch, because the state is its input. But the exact model failed the moment its input was what a sensor gives. With the real sensor the planner succeeded on 0 % of launches: a velocity read from two readings is wrong by 1.41 m/s on each axis. Averaging more readings does not repair that by itself (§5), and in 80 % of launches the ball is behind the curtain at the moment of the nudge, so whatever replaces the state has to be built from what the agent saw a few steps earlier. What should the agent carry in its head instead of the state it cannot see?
Interview prompts
- Why is a model worth building at all, and when is it not? (§2 — it turns computation into trials; planning beats trial and error when an imagined trial is more than K·p/(1 − p) times cheaper than a real one.)
- If a random action succeeds with probability p, how many imagined actions do you need to find a success with probability 1 − ε? (§2 — K ≥ ln ε / ln(1 − p) ≈ (1/p) ln(1/ε).)
- Why does trial and error start again for every new situation while a model does not? (§2 — the model takes the state as an argument, so it answers the question rather than one instance of it.)
- Name the ingredients of a model of consequences and say what can go wrong with each. (§4 — state: hidden or noisy; action: confounded or uncovered; transition: wrong or compounding; score: unknown or exploited.)
- Why does a perfect transition function still fail for a real agent? (§5 — it needs the state, and two noisy readings give a velocity wrong by √2·σ/Δt, which moves the ending by metres.)
- With a wrong state, why does a bigger planning budget not help? (§5 — it finds a nudge that wins in the imagined world; the imagined miss shrinks while the real one does not.)
- How does the exam of this track differ from the exam of the 3D track? (§4 — there the camera was an action that left the world alone; here the action acts, and the exam is the consequence of the action and ultimately the decision it supports.)
Companion reads: 3D Vision, from first principles (the track that ends where this one begins: a scene that has no actions and no future), Reinforcement Learning · 07 Planning (model-based planning in the RL setting), Reinforcement Learning · 03 Observability (why the observation is not the state) and Lesson 17 · Two seats, one tuple (where this series ends: how such a model is trained).