all_lessons/World Models/06 · Compounding errorslesson 6 / 31

Errors compound: rollouts and horizons

Lesson 5 ended with a model that outputs a distribution and a way to use it: draw a sample, treat it as the next state, draw again. That feeds a model its own output, though it was only ever trained on real states. Here a small network whose one-step error is 0.30 mm is 1.55 cm off after ten self-fed steps and 6.56 cm off after eighty. The error obeys a one-line recursion that adds every one-step error, each carried forward by the model's own Jacobians; its growth belongs to the world, bounded in free flight and geometric among round posts. The lesson defines the horizon over which a rollout can be trusted, measures three training fixes (none helps here), and shows what restarts it: a new observation. It cannot say what the model predicts for an action nobody took.

The thesis, here
A model trained on real inputs and run on its own outputs has an error that obeys ek+1 ≈ Jk ek + δk: every one-step error δ is carried to the end by the model's Jacobians J, and teacher forcing trains δ but never asks about J. Whether the sum stays bounded, grows linearly or grows geometrically is decided by the world along the path. So a rollout has a trustworthy horizon, the first step at which its median error crosses a tolerance; a better one-step model lengthens it slowly; only a fresh observation restarts it.
Linear position
Forced by: A model can now output a distribution and draw samples: several futures, each plausible. Using it means feeding each prediction back in as the next input, again and again, though the model was only ever trained on real inputs. A small error becomes the input of the next step. How fast do errors grow when a model predicts from its own predictions, and how far ahead can it be trusted?
New idea: the error of a rollout is the sum of the one-step errors, each carried forward by the model's own Jacobians, so the trustworthy horizon is where that sum crosses a tolerance, and only a new observation restarts it. The growth belongs to the world, so the horizon can be estimated from one-step data.
Forces next: A rollout has a trustworthy horizon that we can now estimate, and re-observing resets it. But every model so far was fitted to logs that someone generated by acting. If whoever acted reacted to something the model cannot see, such as a gust of wind, the logs confound what an action did with why it was taken. Can a model fitted to logs tell what an action causes?
The plan
Seven moves. (1) Feed a trained model its own output and measure. (2) Derive the recursion for the error and test it. (3) Find what sets the growth in the Courtyard. (4) Define the trustworthy horizon. (5) Try three standard training fixes, honestly. (6) Re-observe: the move that restarts the clock. (7) Ask the horizon about an action the log never contained.

1 · The model has only ever seen real states

The task is lesson 1's nudge. A launcher fires the ball, an operator nudges it once at step 12, and everything then coasts. The log is 60 launches of 93 steps (5580 transitions) in the plain Courtyard. The operator aims the ball's stopping point at the goal, so each nudge is a function of the state it met, plus a habitual error of 0.25 m/s. The model is a network of two hidden layers of 12 tanh units that maps a state s = [x, y, vx, vy] (m, m/s) and a nudge to the state one step (0.1 s) later, ŝ′ = s + net(s, a). It is trained by teacher forcing: at every step it is handed the real state and asked for the next one, and its loss is the squared one-step error.

On 100 held-out launches it is excellent. The median one-step error is 0.30 mm (the ball's radius is 100 mm); the step that contains the nudge, which the network saw only 60 times, is worse, 3.4 mm. Now use it as lesson 5 ended: start from the real state at the nudge, give the network the nudge, and feed every output back in as the next input (free running; this task has no fork, so a point prediction is enough, and sampling would only add noise to what follows). Measure the distance between the network's position and the true one after k steps, and take the median over the 100 launches. It is 1.55 cm at k = 10, 3.46 cm at 30 and 6.56 cm at 81 (8.1 s): 52, 116 and 219 times the median one-step error. Nothing in the one-step number says so.

The gap has a name, exposure bias (Ranzato et al., 2016): in training the model sees the data's inputs, in use its own outputs. Cures abound: scheduled sampling (Bengio et al., 2015) moves from true to self-generated inputs during training, Professor Forcing (Lamb et al., 2016) matches the network's behaviour in the two modes, and DAgger (Ross, Gordon and Bagnell, 2011) starts from the observation that sequential prediction breaks the assumption that training and test inputs are alike. GameNGen (Valevski et al., 2024) reports that the shift between teacher-forced training and autoregressive sampling leads to error accumulation and fast degradation of quality, and counters it by adding noise to the context frames. §5 tests three such cures. First: how fast should the error grow?

2 · The recursion

Write f for the world, f̂ for the model, sk for the real state after k steps, ŝk for the model's (ŝ0 = s0), and ek = ŝk − sk. Add and subtract the model's own answer at the real state:

ek+1 = f̂(ŝk) − f(sk) = [f̂(ŝk) − f̂(sk)] + [f̂(sk) − f(sk)]

The second bracket, δk, is the model's one-step error at a real state: what teacher forcing trains and a held-out test measures. The first is how the model reacts to being handed a slightly wrong input, and to first order it is Jk ek with Jk = ∂f̂/∂s at sk, a 4 × 4 matrix. So

ek+1 ≈ Jk ek + δk, en ≈ Σi<n (Jn−1 ⋯ Ji+1) δi + (Jn−1 ⋯ J0) e0

Every one-step error is carried to step n by the product of the Jacobians between. Let ρ be the growth per step of that product along the path.

growth ρ of the productan old error isthe sum of n one-step errors of size δ
ρ < 1forgottenat most δ/(1 − ρ): bounded
ρ = 1keptabout nδ: linear
ρ > 1amplifiedabout ρnδ/(ρ − 1): geometric

Asadi, Misra and Littman (2018) prove the same form for the n-step prediction error, Δ Σ K̄i with K̄ a Lipschitz constant: nΔ when K̄ = 1, geometric when K̄ > 1. The analysis behind Model-Based Policy Optimization (MBPO, Janner et al., 2019) says full rollouts let model errors compound, with a bound that grows with the square of the effective horizon 1/(1 − γ).

Test it. Measure δ and J at the real states of the 100 held-out launches (J is the network's exact input gradient, by backpropagation), run the recursion, take the median. It reproduces the free-running error to within 3.0 % at every k, to 0.9996 and 1.0010 of it at 30 and 81, and within 1.9 % at 10, 30 and 81 for all 48 networks of §5. In this world the recursion is not an analogy, it is the error. What it drops is the curvature of f̂ over the distance e, the model's behaviour a few centimetres off the data (the states visited by the network's own 80-step coasting rollouts are on average 3.0 cm from the real ones), and that is invisible here. The recursion is linear in δ, so it can be taken apart: the nudge step's error alone, carried forward, leaves 5.2 cm at step 81 (median over launches), the 80 coasting steps together leave 6.0 cm, and launch by launch the two add as vectors to the whole, whose median is 6.56 cm. The one step that contains an action is the worst single contributor; §7 returns to it.

Teacher forcing minimises δ at real states. Jk appears nowhere in its loss. What decides J?

3 · What decides the growth: the world

The model's J approximates the world's, so ask the world. The Courtyard has three ingredients.

ingredientwhat it does to an errorgrowth
free flight, one axisposition += g v, v ×= d, with d = e−γΔt = 0.9656 and g = (1 − d)/γ = 0.0983 s: J = [[1, g], [0, d]], eigenvalues 1 and da position error is kept; a velocity error is forgotten over 1/(1 − d) = 29 steps and moves the final position by Σ g di Δv = Δv/γ = 2.86 m per m/s. ρ = 1: a start error stays bounded, a steady bias adds up linearly (below)
walla mirror: the position error is reflected, the velocity error is reflected and shrunk by the restitution 0.9|J| ≤ 1: none
round post (elastic)a ball that meets a post at impact parameter b (its distance from the line through the post's centre) leaves at π − 2 arcsin(b/R) from its incoming direction, R = 0.3 m being the two radii added: an error in b becomes an error of direction 2/√(R² − b²) rad per metre7.07 rad/m at b = 0.1 m, 4.05° per cm: 1 cm off at the post is 21 cm off sideways 3 m later. ρ ≫ 1 per collision

The trained network learned the first row: along the real paths its average velocity entry is 0.9653 (true 0.9656) and its position-from-velocity entry 0.0980 (true 0.0983), so its total travel per m/s of velocity is 2.82 m against the world's 2.86, 1.1 % short: a slightly wrong friction. A bias β per step in the velocity, which is roughly what a slightly wrong friction makes while the ball moves, saturates the velocity error at β/(1 − d) and grows the position error by β/γ per step: a straight line, bent because the ball slows. That is the shape of the curve in §1.

Measure the world's growth where nothing is learned: give the model the exact dynamics (δ = 0) and start it ε = 1 cm wrong, in a random direction of the four-dimensional state (m and m/s together), on the 100 launches with no nudge. In the plain Courtyard the median error is 0.90 cm at 10 steps, 1.47 at 30 and 2.00 at 81. In free flight it cannot exceed √(1 + 1/γ²) = 3.03 times ε (a position error plus its velocity part), and walls only mirror an error; the largest ratio in 300 pairs over 81 steps is 2.86. With a lattice of 23 posts (radius 0.2 m, five columns 1 m apart) the same start gives 2.14 cm, 39.5 cm and 1.23 m: geometric growth at a rate λ = 0.091 per step (300 pairs of starts 1 μm apart; plain world 0.0095), an error multiplied by e every 11 steps. The recursion, with finite-difference Jacobians, follows the lattice while the error is tiny and falls under it later, because a collision is not linear: at ε = 1 cm it gives 0.42 of the measured error at step 30.

4 · The trustworthy horizon

Define it so that it can be measured and predicted. Pick a tolerance τ, the error you can live with (5 cm here, half a ball radius). The trustworthy horizon H(τ) is the first step at which the median error over launches exceeds τ: the median says what to expect, not the worst case, and it is what the recursion predicts.

For the network of §1, H(5 cm) = 52 steps (5.2 s), H(2 cm) = 13, and 10 cm is never crossed in 81 steps. The recursion, fed only the δ and J measured at real states, with no rollout, gives 52 for 5 cm: a horizon can be read off one-step data and the model's Jacobian, and a rollout is only a check.

What improves it depends on the regime. Where ρ > 1 the error is ε eλn and H = ln(τ/ε)/λ: a tenfold better start buys ln 10/λ = 25 steps and no more. In the lattice, with a 5 cm tolerance, a start error of 1 cm lasts 14 steps, 1 mm lasts 38, and 0.1 mm is not crossed in 81; in the plain Courtyard 1 cm or less is never crossed. Where ρ = 1 the sum is linear and a better δ helps, but only as far as the systematic part of δ falls (see the road below).

5 · Three fixes, measured honestly

The recursion says what a fix must do: lower δ at the states a rollout visits, or lower the growth of J. Three standard fixes try; all start from the same pre-trained network, get the same 20 further epochs and learning-rate schedule, and differ only in data or loss.

fixwhat it trains on
relabelled visits (DAgger-style)the states the network's own rollouts visit, labelled with the simulator's true next state and added to the log (two rounds)
multi-step lossthe squared error after 1 to 10 self-fed steps, the gradient passed back through the unrolled steps (checked against finite differences); a cousin of PlaNet's latent overshooting (Hafner et al., 2019), which its final agent works without
noisy inputs (GameNGen-style)the logged states plus noise (3 cm, 2 cm/s) as inputs, the clean next state as the target

One trained network proves nothing: another seed of the plain recipe moves the error at 81 steps between 5.9 and 11.7 cm (median 6.6, a factor 2.0 apart). So each fix is compared with the plain recipe on the same seed, over 12 seeds: the median of the paired ratios of the median errors, and the number of seeds on which the fix wins.

fixerror at 81, ratio to plainwins of 12error at 30, ratio to plainwins of 12
relabelled visits1.0161.143
multi-step loss1.0061.112
noisy inputs3.701.90

Against the plain recipe, then: relabelled visits and the multi-step loss leave the 81-step error where it was (median ratios 1.01 and 1.00; ahead on 6 and 6 of the 12 seeds) and make the 30-step error worse (1.14 and 1.11; ahead on 3 and 2); noisy inputs are worse on every seed, 3.7 times at 81 steps. None of the three lengthens the horizon, for this network, this task and these 20 epochs; longer training or other settings were not tried. For relabelling the reason can be checked. It pays when the network is worse at the states its own rollouts reach than at real states, and it is not: the median one-step error at the visited coasting states is 0.29 mm, against 0.30 mm at the real ones (ratio 0.99), which fits §2, where the recursion with real-state δ was already exact. The relabelled rollouts also start after the nudge, so the nudge step, the largest single contributor in §2, is never relabelled. For the multi-step loss there is the null result and no explanation; its gradient is checked against finite differences, so the null result is not a bug in the gradient. Noise does damage because a noisy four-number state is still a state, and a ball there moves differently: trained to map noisy inputs to the clean next state, the network learns to shrink velocities (its velocity entry falls from 0.9653 to 0.9604), so it lets the ball travel 13 % less far. In a picture the trick can work, since pixels are redundant and a network can learn to pull a noisy frame back toward a clean one (GameNGen reports that without noisy context its rollouts degrade fast after 20 to 30 steps). A four-number state has no redundancy to pull toward.

6 · Re-observe: the clock restarts

The sum has a start. Replace the model's state by the real one every m steps (a reading, or a belief: lesson 2's filter doing the job) and the product of Jacobians restarts, so the error depends only on the steps since the last reading. For the network of §1, an exact reading every 10 steps keeps the median error under 1.37 cm, where open loop it reaches 6.56 cm at 81; every 5 steps, 0.72 cm. In the lattice, where the open-loop error from a 1 cm start is 1.23 m at step 81, a reading with a 1 cm error every 10 steps keeps it under 1.53 cm. The cap is the error of an m-step rollout plus the error of the reading: a horizon of m, repeated. MBPO (Janner et al., 2019) does the same with a model, branching short rollouts from real states; its authors found 200-step rollouts worse than short ones. A real reading is noisy (σ = 10 cm in the Courtyard) and sometimes missing (behind the curtain): the reset is then lesson 2's posterior, whose error sets the floor instead of zero.

The widget

Feed a model its own output
Top left: one launch, real path cyan, model amber (centimetres are about a pixel at this scale, so read errors in the other panels). Right: the median error over 100 launches against steps k (log axis), the recursion's prediction (dashed), the tolerance (red) and the horizon where they meet. Bottom left: one dot per launch at its error at step k (log radius; rings at 1 mm, 1 cm, 10 cm, 1 m), the amber ring the median. Amplification is the error at k over the one-step error (learned network) or the start error (exact model).
one-step error (median)
—
error at step k
—
amplification at step k
—
error at 10 steps
—
error at 30 steps
—
error at 81 steps
—
horizon, measured
—
horizon, from the recursion
—
recursion ÷ measured at k
—
Show the core JS
L6.roll = function (step, truth, a1, K, m, kick) {
  var p = [kick ? L6.add(truth[0], kick(0)) : truth[0].slice()], k, s;
  for (k = 1; k <= K; k++) {
    s = step(p[k - 1], k === 1 ? a1 : null);
    if (m > 0 && k % m === 0) s = kick ? L6.add(truth[k], kick(k / m)) : truth[k].slice();
    p.push(s);
  }
  return p;
};
L6.linearise = function (J, dl, K, m, kick) {
  var e = kick ? kick(0).slice() : [0, 0, 0, 0], out = [e], k, c, n, Jk;
  for (k = 0; k < K; k++) {
    Jk = J[k]; n = [dl[k][0], dl[k][1], dl[k][2], dl[k][3]];
    for (c = 0; c < 4; c++) n[c] += Jk[c * 4] * e[0] + Jk[c * 4 + 1] * e[1] + Jk[c * 4 + 2] * e[2] + Jk[c * 4 + 3] * e[3];
    if (m > 0 && (k + 1) % m === 0) n = kick ? kick((k + 1) / m).slice() : [0, 0, 0, 0];
    e = n; out.push(e);
  }
  return out;
};

What to try. Start with the learned network, teacher forcing, k = 30, tolerance 5 cm. The error is 3.46 cm, 116 times the one-step error of 0.30 mm; the dashed recursion lies on the solid curve (ratio 0.9996) and both cross the red line at step 52. Drag k to 81: the dots move out through the rings and the median reaches 6.56 cm. A tolerance of 2 cm gives a horizon of 13; of 10 cm, none. Retrain: noisy inputs give 29 cm at 81, relabelled visits 6.16 and the multi-step loss 7.20, against 6.56 for teacher forcing (one seed, so the last two lie inside the spread that §5 measures over twelve). The exact model on the plain Courtyard with a 1 cm start error: 1.47 cm at 30 and 2.00 cm at 81, the amplification under 3.03. With the 23 posts the twelve amber copies fan out, the error is 39.5 cm at 30 and 1.23 m at 81, the horizon 14; with a 1 mm start 38, with 0.1 mm none. Re-observe every 10 steps: the error is a sawtooth that never passes 1.37 cm (lattice 1.53 cm). Last, set the learned network's nudges to "shifted 1 m/s": the error at 81 becomes 53 cm and the horizon drops from 52 to 11 (§7).

Road not taken · a better model, or the average of several
The obvious cure is a better one-step model. More data shrinks δ: with 15, 60 and 240 launches in the log (five networks each) the median one-step error is 1.00, 0.26 and 0.109 mm, a factor 9.2, but the 5 cm horizon is 20, 50 and 69 steps, a factor 3.45. The sum is driven by the systematic part of δ and by the J that carries it, not by the typical error of one step; and in a chaotic world a factor 9.2 in precision buys ln 9.2/λ = 24 steps. Averaging helps little: the mean of five networks' rollouts is off by 3.30 cm at 30 steps against 3.82 for a member, and by 6.64 against 6.56 cm at 81: the members err the same way and disagree by only 1.89 cm at 30 steps (lesson 5: disagreement is not error). Never rolling out is model-free learning, which gives up lesson 1's economics.

7 · The horizon was measured on the log's own actions

Everything above used nudges like the log's, and those are not arbitrary. The operator aims the stopping point at the goal, so a nudge is nearly a linear function of the state it met: regressing 2000 logged nudges on that state explains 74 % of their variance; the rest is habit. The network has seen 60 nudges, each the right one for its state, and nothing like a nudge the state does not call for.

Ask it about one. Shift every held-out nudge by 0.5 or 1 m/s in a random direction (the last control of the widget) and compare the network's imagined 81-step ending with the simulator's. For the operator's own nudges 1 % of imagined endings are off by more than the goal's radius (50 cm); for a shift of 0.5 m/s 5 %; for 1 m/s 51 % (25 % even among the 51 launches whose true path never touches a wall, so it is not only the walls). The median error at 81 steps is 53 cm against 6.6 cm, 8.1 times larger, and the 5 cm horizon falls from 52 to 31 and 11 steps. The recursion still predicts those horizons to within a step: the mechanism is unchanged, only δ at the nudge step is 4.0 times larger. A held-out test drawn from the log would have passed.

What this lesson did not do
It measured the error of a deterministic model; a sampled rollout (lesson 5) adds the sampling noise of every step to the same recursion, which was not measured. Its re-observation was an exact reading, or one with a stated error; a real reading is noisy and missing behind the curtain, and the reset is then lesson 2's belief, learned in lesson 4. It used a four-number state; for pictures compounding appears as drift (lesson 11), and the fixes behave differently where noisy inputs are not valid inputs. It did not say what to do with a horizon: choosing actions on a model that is wrong in a way an optimiser can exploit is lesson 8, and replanning from fresh observations is lesson 9. And it did not say what a model fitted to a log knows about an action the log never took: lesson 7.

Common mistakes / failure modes

"a small one-step error means a small rollout error"
0.30 mm per step became 6.56 cm at 81 steps, 219 times larger: a rollout's error is a sum carried by Jacobians (§1, §2).
"errors compound exponentially"
Only where ρ > 1. In free flight the error stays under 3.03 times the start error; among posts it grows by e every 11 steps (§3).
"the training loss tells me how far to trust the model"
It measures δ at real states and never J: another seed moves the 81-step error by a factor 2.0, and 16 times the data lowered the one-step error 9.2 times for a 3.45 times longer horizon (§4, §5, road).
"add noise to the inputs: that is what GameNGen does"
It works for pictures, where a noisy frame is not a frame. For a four-number state it made the 81-step error 3.7 times worse on all 12 seeds (§5).
"a model validated on held-out log data is validated"
Only for the actions the log contains: for nudges shifted by 1 m/s, 51 % of imagined endings missed the goal's radius, against 1 % (§7).

Checkpoint exercise

Try it
A model is exact except that its friction is 0.30 s⁻¹ instead of 0.35 s⁻¹. A ball is launched at 2 m/s down a corridor long enough that nothing is hit. (a) By how much do the model's and the real positions differ after one step of 0.1 s? (b) Where does each say the ball stops? (c) How many times larger is the final error than the one-step error? Answer: (a) a step moves the ball g v with g = (1 − e−γΔt)/γ: 0.098515 s for 0.30 and 0.098270 s for 0.35, so the positions differ by 0.000245 s × 2 m/s = 0.49 mm. (b) The ball stops v0/γ from where it was launched: 6.67 m for the model, 5.71 m for the world, 0.95 m apart. (c) 1946 times: the position part of J has eigenvalue 1, so nothing is forgotten, and a velocity error is carried until friction removes it (§3).

Where this points next

A rollout now has a trustworthy horizon that can be estimated from one-step data (52 steps at 5 cm for the trained network, 13 at 2 cm), and re-observing every ten steps holds the error under 1.37 cm. But every model so far was fitted to logs that someone generated by acting, and the horizon was measured on actions like the log's. Shift the nudges by 1 m/s and it falls to 11 steps, 51 % of the imagined endings miss the goal's radius, and a test on held-out log data would not have said so. If whoever acted reacted to something the model cannot see, such as a gust of wind, the logs confound what an action did with why it was taken. Can a model fitted to logs tell what an action causes?

Takeaway
A model trained on real states and run on its own outputs has an error that obeys ek+1 ≈ Jk ek + δk: the sum of its one-step errors, each carried forward by its Jacobians; for the trained network that is the whole story to within 3.0 %. How the sum grows is the world's: bounded in free flight (under 3.03 times a start error), geometric among round posts (25 steps per decade of precision). The trustworthy horizon, the first step at which the median error crosses a tolerance, is 52 steps at 5 cm and can be estimated from δ and J at real states. Relabelling and multi-step training did not move it beyond the seed spread, noisy inputs made it worse, and 16 times the data lengthened it 3.45 times; re-observing every m steps restarts the sum and caps the error at its m-step value. All of this holds for actions like the log's: shift the nudges by 1 m/s and the horizon falls from 52 to 11.

Interview prompts

Companion reads: Lesson 21 · From teacher forcing to rollout (the same fixes as training recipes at scale) and Robot Model Training · 03 Labels on your own states: DAgger and corrections (the policy-side version of exposure bias).