all_lessons/World Models/01 · Why predict?lesson 1 / 31

Why predict? The world-model contract

A reconstruction explains a scene that was recorded. It has no actions and no future, so it cannot say what would happen if the agent did something. This lesson asks what an agent gains from a model that can, and puts a price on it. In a small courtyard a ball must be tapped once into a goal. Trying nudges for real takes dozens of trials; trying them in one's head takes none, provided a function says what a nudge does. We build that planner with a perfect model, measure the exchange rate between imagined and real trials, name the four things any model of consequences needs, and then hand the planner what a real agent is given: a noisy reading, and nothing at all while the ball is behind a curtain. The model is exact. The input is not.

The thesis, here
A world model is a function from what the agent knows and what it does to what happens next. Its value is economic: it turns computation into trials. With an exact model and the true state, one real push does the work of about 39 blind ones, and it does so again for every new launch. But the function is only half of a contract. It has to be given a state, and a real agent is never given one.
Linear position
Forced by: A scene that changes can be reconstructed from what has already happened. An agent that has to act needs the other direction: what will happen next, and what would happen if it did something else. Nothing in a reconstruction answers that, because it has no actions and no future. What must a model carry in its head to answer "if I do this, what happens next?", and why would an agent want one at all?
New idea: a world model turns computation into trials: it answers "if I do this, what happens next?" so that an agent can try many actions in its head and one in the world. The answer needs four ingredients, a state, an action, a transition and a score, and the rest of the series is organised by how each of them fails.
Forces next: A model that answers "if I push this way, where does the ball end up?" makes thinking cheaper than trying: one push in the real world instead of dozens, provided the model is handed the true state of the ball. An agent is never handed the state. It sees a noisy position, and nothing at all while the ball is behind the curtain. What should the agent carry in its head instead of the state it cannot see?
The plan
Five moves. (1) Pose the task in the Courtyard and put a price on trying. (2) Replace real trials by imagined ones, derive how many are needed and when that pays. (3) Run the planner with a perfect model. (4) Name the four ingredients of the contract. (5) Give the planner what a real agent gets, and watch the perfect model fail.

1 · A courtyard, and the price of trying

Every lesson in this series happens in one small world, the Courtyard: a floor 8 m wide and 5 m deep (x to the right, y up) with a ball of radius 0.1 m on it. Time moves in steps of Δt = 0.1 s. Friction removes 3.4 % of the ball's velocity each step (a rate γ = 0.35 s−1), so on open floor a ball pushed at speed v rolls to a stop a distance v/γ from where it was pushed, 94 % of the way there after 8 s. The walls return 90 % of the speed on impact. A curtain between x = 2.6 m and 3.4 m hides the ball from the sensor, and a goal disc of radius 0.5 m sits at (6.6, 2.5). The ball's state is four numbers, s = (x, y, vx, vy). An action is an impulse a = (ax, ay) added to the velocity at the start of a step.

The task, called the nudge, is this. A launcher fires the ball from the left wall at 2–3 m/s in a direction that varies from launch to launch. The agent may add one impulse of size at most 3 m/s at t = 1.2 s (step 12); after that everything coasts for 80 more steps, which is 8 s. It succeeds if the ball ends inside the goal. At the moment of the nudge the ball is rolling at about 1.6 m/s, and in 80 % of launches it is behind the curtain.

Suppose the agent knows no physics and learns by trial and error. A trial is a real event: launch, nudge, wait 8 s, look. Choose the nudge at random and the chance that a trial succeeds is some p. We can measure it by trying 512 random nudges on each of 60 launches (all computed, which is the point of the next section) and counting the successes: p = 2.6 %. The number of trials until the first success is then geometric with mean 1/p ≈ 39 and a long tail: the chance of no success in 100 blind trials is (1 − p)100 = 7.4 %. In the widget below the 60 launches average 35.8 trials; the gap to 39 is sampling noise.

A smarter learner would use what each trial shows. It keeps the best nudge so far, tries perturbations of it, widens the step when a perturbation wins and narrows it when not (a (1+1) evolution strategy, restarted after 25 stalled trials). It needs a median of 16.5 real trials per launch and a mean of 21. Learning from outcomes roughly halves the bill. It does not change its order of magnitude, which is dozens.

Nothing carries over, either. The next launch has a different state, so the nudge that worked is the wrong one, and the learner begins again. Over 100 launches blind trial and error costs about 3894 real trials.

2 · Replace trials with thought

Suppose we had a function M that takes the state at the nudge and a nudge and returns where the ball ends up: if I do this, what happens? Then a nudge need not be tried to learn what it does; it can be computed. Draw K candidate nudges, compute the ending of each, keep the one that ends closest to the goal, and do only that one for real. For the Courtyard we can give the agent the best M there is, the simulator itself. Later lessons take that gift away; here it isolates what a model is for.

How many imagined nudges? Each candidate succeeds independently with probability p, so at least one of K does with probability

P(K) = 1 − (1 − p)K, and P(K) ≥ 0.99 once K ≥ ln 0.01 / ln(1 − p) ≈ 4.6 / p

For p = 2.6 % that is 177 candidates. Imagination buys certainty: the 39 real trials of blind search are replaced by one real trial and about 177 imagined ones.

When does that pay? Let cr be the cost of a real trial and ci the cost of an imagined one. Trial and error costs cr/p in expectation; planning costs cr + K·ci. Planning is cheaper when K·ci < cr(1/p − 1), that is

cr / ci > K·p / (1 − p) = 4.7 for K = 177

A model pays whenever imagining a nudge is more than about five times cheaper than trying it. An imagined nudge here is 81 simulator steps, microseconds on a laptop; a real one is 8 s of world time and a reset. The ratio is in the millions. Real systems show the same exchange on harder tasks: PlaNet (Hafner et al., 2019) reports beating the model-free methods A3C, and in some cases D4PG, with on average 200 times less environment interaction, and PETS (Chua et al., 2018) matches the asymptotic performance of Soft Actor-Critic and PPO on the half-cheetah task with 8 and 125 times fewer samples.

One more property is easy to miss. M takes the state as an argument, so the same M serves the next launch and the one after. Trial and error learns the answer to one situation. A model learns the question.

3 · The oracle planner

The widget runs exactly this. For each of 60 fixed launches it rolls the ball to the nudge time, draws 512 candidate nudges, and computes the ending of every candidate with the simulator. A planner with budget K looks at the first K candidates, takes the one with the smallest imagined miss, and the nudge is then executed in the real world, which here means the same simulator started from the true state. The success curve on the right is that experiment for every K from 1 to 512.

Think first, push once
Left: the Courtyard at the moment of the nudge. The black line is the ball so far; the purple fan is the planner's imagined endings for the first K candidate nudges (green: would succeed); amber is the one it picks; the heavy black line is what really happens. Right: the share of the 60 launches the one real push puts in the goal, against K, with the formula 1 − (1 − p)K dashed. Below: how many real trials each approach needs per launch. Switch what the planner is handed and move the sensor noise to see the dashed ring, the state it believes.
believed position error
—
believed velocity error
—
best imagined miss
—
real miss
—
this launch
—
success, all 60 launches
—
blind trials per launch
—
adaptive search, median
—
break-even cost ratio
—
Show the core JS
function outcome(s, a) {
  var q = CY.step(W0, s, a);                                  // the nudge
  for (var u = 0; u < TF; u++) q = CY.step(W0, q, null);      // then everything coasts
  return q;
}
function missOf(f) { return Math.hypot(f[0] - W0.goal.x, f[1] - W0.goal.y); }
function bestOf(row, K) {                                     // the planner: smallest imagined miss among the first K candidates
  var best = Infinity, bi = 0;
  for (var k = 0; k < K; k++) if (row[k] < best) { best = row[k]; bi = k; }
  return bi;
}
function curveOf(miss) {                                      // success of the ONE real push for every budget K = 1 .. KMAX
  var succ = new Float64Array(KMAX);
  for (var n = 0; n < N; n++) {
    var best = Infinity, bi = 0;
    for (var k = 0; k < KMAX; k++) {
      if (miss[n][k] < best) { best = miss[n][k]; bi = k; }   // best imagined candidate among the first k + 1
      if (TRUE_MISS[n][bi] < GR) succ[k] += 1 / N;            // does it really win?
    }
  }
  return succ;
}

What to try. Start with the planner handed the true state and K = 128. It succeeds on 98.3 % of the 60 launches, where 1 − (1 − p)128 predicts 96.4 %. Slide K down: K = 16 gives 35 % (formula 34 %) and K = 1, a blind guess, 3.3 % (2.6 %). Each green path in the fan is an imagined nudge that would have won, the amber one is the planner's pick, and the heavy black path lands where the amber one did, because the model is exact. Step through other launches with the launch slider: the same planner, a different state, one real push each time. Now switch to what the sensor gives. The dashed ring is the state the planner now believes. On launch 28 the ball is behind the curtain, so the last two readings are older than the nudge; the believed position is 0.31 m off and the believed velocity 1.45 m/s off. The planner imagines landing 0.11 m from the goal's centre, and the real push lands 3.27 m away. Raise K to 512: the imagined misses keep shrinking (averaged over the 60 launches the planner's imagined miss is 0.12 m) and the real miss does not (2.44 m on average); not one of the 60 launches succeeds, 0 %. Finally move the sensor noise down. At 0 the planner is as good as it was with the true state (98.3 %), although on 80 % of the launches the ring then sits behind the ball (§5 says why that costs nothing); at 5 mm it is 91.7 %, at 1 cm 50 %, at 2 cm 23 %, at 5 cm 6.7 %, and at the Courtyard's real 100 mm 0 % (all at K = 128). For 90 % success the sensor would have to be accurate to about 5.5 mm, which is 18 times better than the one the agent has.

4 · The contract

The planner above was handed four things, and it is worth listing them, because a "world model" in the narrow sense is only one of them.

IngredientWhat we handed the plannerWhat a real agent has to supplyWhere the series goes
a state sthe true four numberssomething built from noisy, partial readings, or from pictureslessons 2–4
an action aany of K freely chosen nudgesactions it has seen, or may only be able to take where it has tried themlesson 7
a transition M(s, a)the exact simulatora learned function: wrong in places, one future or many, chained again and againlessons 4–6
a scoredistance of the ending from the goala value for imagined futures, and a way to know the model is good enough for the decisionlessons 8–9, 16

A transition is meaningless on its own: it needs a state to start from, an action to vary, and a score to compare the endings by. Each of the next fifteen lessons takes one of these gifts away, or the assumption behind it, and measures what breaks.

The exam also changes. The 3D track graded a model of a scene by predicting a photograph from a camera nobody used, and that camera was an action which left the world alone. Here the action acts. The exam is predict what the world does when I do this, and in the end are the decisions it supports better ones.

Road not taken · skip the model and learn a policy by trial and error
A model-free learner needs no model of consequences, and when trials are nearly free it is the right choice: if a real trial costs about as much as an imagined one (cr/ci below the break-even above), the model's overhead is not repaid. Where trials are slow, dangerous or scarce it pays the bill of §1 until what it has learned carries over from one situation to the next. The series comes back to this trade in lesson 8, where a policy learns from imagined trials, and in lesson 16, which asks how to know when the imagination can be trusted.

5 · Hand the planner what an agent gets

Everything in the oracle planner is exact except one thing, the input. A real agent is not handed the ball's state. Its sensor reports the position with noise, σO = 0.1 m on each axis, and reports nothing while the ball is behind the curtain. The velocity, on which the future depends even more than on the position, is not measured at all. The simplest thing to do is to infer it from the last reading z1 and the one before it, z2, taken Δt apart:

v̂ = (z1 − z2) / Δt, so Var(v̂) = 2σO² / Δt² and error(v̂) = √2·σO / Δt = 1.41 m/s

That error is nearly as large as the ball's speed (1.6 m/s). A velocity error δv costs far more than a position error, because the ball keeps going: it moves the stopping point by δv/γ = 2.86 m for every m/s, where a position error of one metre moves it by one metre. An error of 1.41 m/s on each axis therefore moves the ending by 4.0 m on each axis, against a goal 1 m across. On the widget's 60 launches (K = 128), a planner handed the true state with only the position wrong by the sensor's 0.1 m still succeeds on 96.7 %; handed the true state with only the velocity wrong by what two readings give, it succeeds on 8.3 %. The planner imagines worlds that are metres away from the real one and, being a good optimiser, finds a nudge that wins in the wrong world. Adding candidates makes it better at winning there (imagined miss 0.12 m at K = 512) and no better at winning here (real miss 2.44 m).

Behind the curtain the two readings are the last ones taken before the ball vanished, so the planner starts its imagination from where the ball was, not where it is. By itself that costs nothing: a coasting ball's stopping point, x + v/γ, does not change while nothing touches it, so an old state points at nearly the same ending as the current one. With noise-free readings the planner scores 98.3 %, as with the true state. The damage is the noise.

This is not a failure of the planner, nor of the model, which is exact. It is the first failure of the contract: the transition is a function of a state, and the agent has none.

Road not taken · average more readings
The instinct is right, and the lesson is in why it falls short. A least-squares line through the last 10 readings removes almost all the noise: its slope error is √12·σO/(Δt·√(m(m² − 1))) = 0.11 m/s per axis, which would move the stopping point by only 0.31 m. But a straight line has no friction in it. Its slope is the velocity in the middle of the window, 0.45 s before the last reading, when the ball was 17 % faster than it was at the last reading, and averaging cannot remove a bias. A planner handed that line (the last ten readings, or all it has) succeeds on 11.7 % of the 60 launches at K = 128, and on only 6.7 % when the readings are noise-free: what is left is the model, not the noise. The summary has to contain the ball's dynamics; a summary of the past that does, and says how sure it is, is what the next lesson derives.
What this lesson did not do
It did not say what to do when the state is hidden (lesson 2) or arrives as pictures (lesson 3), how a transition is learned from data rather than given (lesson 4), what to do when the future is not one point (lesson 5) or when errors compound (lesson 6), whether a model fitted to logs of other people's actions tells us what our action does (lesson 7), or how to use a model for more than choosing the best of K (lessons 8–9). It also used the simulator as the model, so it said nothing about how wrong a model may be.

Common mistakes / failure modes

"a world model must predict pixels"
The contract is about consequences: a state, an action, a transition and a score. Pixels are one possible observation, and a costly one (§4; lessons 3 and 11).
"a model is good if its one-step predictions are accurate"
Here the transition is exact and the system still succeeds on 0 % of launches with the real sensor, because the model needs a state the agent does not have (§5).
"more imagination always helps"
With a wrong state, a larger K finds a better imagined ending and no better real one: at K = 512 the success rate is 0 % (§3 widget, §5).
"an imagined success is a success"
At K = 512 the planner imagines landing 0.12 m from the goal's centre and really lands 2.44 m away. Lesson 8 names this effect (§5).
"model-based always beats model-free"
It pays only when imagining a nudge is more than 4.7 times cheaper than trying one, and only if the model and state are good enough (§2).
"the state is what the sensor reports"
The state includes the velocity, which no single reading holds, and the sensor reports nothing at all behind the curtain (§5; lesson 2).

Checkpoint exercise

Try it
A different task has a blind success probability of p = 2 % per trial. (a) How many real trials does blind search need on average? (b) How many imagined candidates make you 95 % sure that at least one succeeds? (c) By what factor must an imagined trial be cheaper than a real one for planning to beat blind search? Answer: (a) 1/p = 50 trials. (b) K ≥ ln 0.05 / ln 0.98 = 149. (c) K·p/(1 − p) = 3.04: imagining must be about three times cheaper than trying. That is below the Courtyard's 4.7 because 95 % certainty is asked for here and 99 % there: for small p, K·p is close to ln(1/0.05) = 3.0 whatever p is.

Where this points next

A model that answers "if I push this way, where does the ball end up?" turns computation into trials: one push in the world where blind search needs about 39 and a good learner about 21. It does so for any launch, because the state is its input. But the exact model failed the moment its input was what a sensor gives. With the real sensor the planner succeeded on 0 % of launches: a velocity read from two readings is wrong by 1.41 m/s on each axis. Averaging more readings does not repair that by itself (§5), and in 80 % of launches the ball is behind the curtain at the moment of the nudge, so whatever replaces the state has to be built from what the agent saw a few steps earlier. What should the agent carry in its head instead of the state it cannot see?

Takeaway
A world model is a function that answers "if I do this, what happens next?" Its value is economic: it turns computation into trials, so an agent can try K actions in its head and one in the world. In the Courtyard a blind push succeeds 2.6 % of the time, so blind trial and error needs about 39 real trials per launch and a good adaptive learner about 21; an exact model needs 177 imagined candidates and one real push, and it pays whenever imagining is more than 4.7 times cheaper than trying. The contract has four ingredients: a state, an action, a transition and a score. The oracle planner was handed the first and the third perfectly; given what a sensor gives, its success fell from 98.3 % to 0 %. The rest of the series removes one gift at a time and measures what breaks.

Interview prompts

Companion reads: 3D Vision, from first principles (the track that ends where this one begins: a scene that has no actions and no future), Reinforcement Learning · 07 Planning (model-based planning in the RL setting), Reinforcement Learning · 03 Observability (why the observation is not the state) and Lesson 17 · Two seats, one tuple (where this series ends: how such a model is trained).