Errors compound: rollouts and horizons
Lesson 5 ended with a model that outputs a distribution and a way to use it: draw a sample, treat it as the next state, draw again. That feeds a model its own output, though it was only ever trained on real states. Here a small network whose one-step error is 0.30 mm is 1.55 cm off after ten self-fed steps and 6.56 cm off after eighty. The error obeys a one-line recursion that adds every one-step error, each carried forward by the model's own Jacobians; its growth belongs to the world, bounded in free flight and geometric among round posts. The lesson defines the horizon over which a rollout can be trusted, measures three training fixes (none helps here), and shows what restarts it: a new observation. It cannot say what the model predicts for an action nobody took.
New idea: the error of a rollout is the sum of the one-step errors, each carried forward by the model's own Jacobians, so the trustworthy horizon is where that sum crosses a tolerance, and only a new observation restarts it. The growth belongs to the world, so the horizon can be estimated from one-step data.
Forces next: A rollout has a trustworthy horizon that we can now estimate, and re-observing resets it. But every model so far was fitted to logs that someone generated by acting. If whoever acted reacted to something the model cannot see, such as a gust of wind, the logs confound what an action did with why it was taken. Can a model fitted to logs tell what an action causes?
1 · The model has only ever seen real states
The task is lesson 1's nudge. A launcher fires the ball, an operator nudges it once at step 12, and everything then coasts. The log is 60 launches of 93 steps (5580 transitions) in the plain Courtyard. The operator aims the ball's stopping point at the goal, so each nudge is a function of the state it met, plus a habitual error of 0.25 m/s. The model is a network of two hidden layers of 12 tanh units that maps a state s = [x, y, vx, vy] (m, m/s) and a nudge to the state one step (0.1 s) later, ŝ′ = s + net(s, a). It is trained by teacher forcing: at every step it is handed the real state and asked for the next one, and its loss is the squared one-step error.
On 100 held-out launches it is excellent. The median one-step error is 0.30 mm (the ball's radius is 100 mm); the step that contains the nudge, which the network saw only 60 times, is worse, 3.4 mm. Now use it as lesson 5 ended: start from the real state at the nudge, give the network the nudge, and feed every output back in as the next input (free running; this task has no fork, so a point prediction is enough, and sampling would only add noise to what follows). Measure the distance between the network's position and the true one after k steps, and take the median over the 100 launches. It is 1.55 cm at k = 10, 3.46 cm at 30 and 6.56 cm at 81 (8.1 s): 52, 116 and 219 times the median one-step error. Nothing in the one-step number says so.
The gap has a name, exposure bias (Ranzato et al., 2016): in training the model sees the data's inputs, in use its own outputs. Cures abound: scheduled sampling (Bengio et al., 2015) moves from true to self-generated inputs during training, Professor Forcing (Lamb et al., 2016) matches the network's behaviour in the two modes, and DAgger (Ross, Gordon and Bagnell, 2011) starts from the observation that sequential prediction breaks the assumption that training and test inputs are alike. GameNGen (Valevski et al., 2024) reports that the shift between teacher-forced training and autoregressive sampling leads to error accumulation and fast degradation of quality, and counters it by adding noise to the context frames. §5 tests three such cures. First: how fast should the error grow?
2 · The recursion
Write f for the world, f̂ for the model, sk for the real state after k steps, ŝk for the model's (ŝ0 = s0), and ek = ŝk − sk. Add and subtract the model's own answer at the real state:
ek+1 = f̂(ŝk) − f(sk) = [f̂(ŝk) − f̂(sk)] + [f̂(sk) − f(sk)]
The second bracket, δk, is the model's one-step error at a real state: what teacher forcing trains and a held-out test measures. The first is how the model reacts to being handed a slightly wrong input, and to first order it is Jk ek with Jk = ∂f̂/∂s at sk, a 4 × 4 matrix. So
ek+1 ≈ Jk ek + δk, en ≈ Σi<n (Jn−1 ⋯ Ji+1) δi + (Jn−1 ⋯ J0) e0
Every one-step error is carried to step n by the product of the Jacobians between. Let ρ be the growth per step of that product along the path.
| growth ρ of the product | an old error is | the sum of n one-step errors of size δ |
|---|---|---|
| ρ < 1 | forgotten | at most δ/(1 − ρ): bounded |
| ρ = 1 | kept | about nδ: linear |
| ρ > 1 | amplified | about ρnδ/(ρ − 1): geometric |
Asadi, Misra and Littman (2018) prove the same form for the n-step prediction error, Δ Σ K̄i with K̄ a Lipschitz constant: nΔ when K̄ = 1, geometric when K̄ > 1. The analysis behind Model-Based Policy Optimization (MBPO, Janner et al., 2019) says full rollouts let model errors compound, with a bound that grows with the square of the effective horizon 1/(1 − γ).
Test it. Measure δ and J at the real states of the 100 held-out launches (J is the network's exact input gradient, by backpropagation), run the recursion, take the median. It reproduces the free-running error to within 3.0 % at every k, to 0.9996 and 1.0010 of it at 30 and 81, and within 1.9 % at 10, 30 and 81 for all 48 networks of §5. In this world the recursion is not an analogy, it is the error. What it drops is the curvature of f̂ over the distance e, the model's behaviour a few centimetres off the data (the states visited by the network's own 80-step coasting rollouts are on average 3.0 cm from the real ones), and that is invisible here. The recursion is linear in δ, so it can be taken apart: the nudge step's error alone, carried forward, leaves 5.2 cm at step 81 (median over launches), the 80 coasting steps together leave 6.0 cm, and launch by launch the two add as vectors to the whole, whose median is 6.56 cm. The one step that contains an action is the worst single contributor; §7 returns to it.
Teacher forcing minimises δ at real states. Jk appears nowhere in its loss. What decides J?
3 · What decides the growth: the world
The model's J approximates the world's, so ask the world. The Courtyard has three ingredients.
| ingredient | what it does to an error | growth |
|---|---|---|
| free flight, one axis | position += g v, v ×= d, with d = e−γΔt = 0.9656 and g = (1 − d)/γ = 0.0983 s: J = [[1, g], [0, d]], eigenvalues 1 and d | a position error is kept; a velocity error is forgotten over 1/(1 − d) = 29 steps and moves the final position by Σ g di Δv = Δv/γ = 2.86 m per m/s. ρ = 1: a start error stays bounded, a steady bias adds up linearly (below) |
| wall | a mirror: the position error is reflected, the velocity error is reflected and shrunk by the restitution 0.9 | |J| ≤ 1: none |
| round post (elastic) | a ball that meets a post at impact parameter b (its distance from the line through the post's centre) leaves at π − 2 arcsin(b/R) from its incoming direction, R = 0.3 m being the two radii added: an error in b becomes an error of direction 2/√(R² − b²) rad per metre | 7.07 rad/m at b = 0.1 m, 4.05° per cm: 1 cm off at the post is 21 cm off sideways 3 m later. ρ ≫ 1 per collision |
The trained network learned the first row: along the real paths its average velocity entry is 0.9653 (true 0.9656) and its position-from-velocity entry 0.0980 (true 0.0983), so its total travel per m/s of velocity is 2.82 m against the world's 2.86, 1.1 % short: a slightly wrong friction. A bias β per step in the velocity, which is roughly what a slightly wrong friction makes while the ball moves, saturates the velocity error at β/(1 − d) and grows the position error by β/γ per step: a straight line, bent because the ball slows. That is the shape of the curve in §1.
Measure the world's growth where nothing is learned: give the model the exact dynamics (δ = 0) and start it ε = 1 cm wrong, in a random direction of the four-dimensional state (m and m/s together), on the 100 launches with no nudge. In the plain Courtyard the median error is 0.90 cm at 10 steps, 1.47 at 30 and 2.00 at 81. In free flight it cannot exceed √(1 + 1/γ²) = 3.03 times ε (a position error plus its velocity part), and walls only mirror an error; the largest ratio in 300 pairs over 81 steps is 2.86. With a lattice of 23 posts (radius 0.2 m, five columns 1 m apart) the same start gives 2.14 cm, 39.5 cm and 1.23 m: geometric growth at a rate λ = 0.091 per step (300 pairs of starts 1 μm apart; plain world 0.0095), an error multiplied by e every 11 steps. The recursion, with finite-difference Jacobians, follows the lattice while the error is tiny and falls under it later, because a collision is not linear: at ε = 1 cm it gives 0.42 of the measured error at step 30.
4 · The trustworthy horizon
Define it so that it can be measured and predicted. Pick a tolerance τ, the error you can live with (5 cm here, half a ball radius). The trustworthy horizon H(τ) is the first step at which the median error over launches exceeds τ: the median says what to expect, not the worst case, and it is what the recursion predicts.
For the network of §1, H(5 cm) = 52 steps (5.2 s), H(2 cm) = 13, and 10 cm is never crossed in 81 steps. The recursion, fed only the δ and J measured at real states, with no rollout, gives 52 for 5 cm: a horizon can be read off one-step data and the model's Jacobian, and a rollout is only a check.
What improves it depends on the regime. Where ρ > 1 the error is ε eλn and H = ln(τ/ε)/λ: a tenfold better start buys ln 10/λ = 25 steps and no more. In the lattice, with a 5 cm tolerance, a start error of 1 cm lasts 14 steps, 1 mm lasts 38, and 0.1 mm is not crossed in 81; in the plain Courtyard 1 cm or less is never crossed. Where ρ = 1 the sum is linear and a better δ helps, but only as far as the systematic part of δ falls (see the road below).
5 · Three fixes, measured honestly
The recursion says what a fix must do: lower δ at the states a rollout visits, or lower the growth of J. Three standard fixes try; all start from the same pre-trained network, get the same 20 further epochs and learning-rate schedule, and differ only in data or loss.
| fix | what it trains on |
|---|---|
| relabelled visits (DAgger-style) | the states the network's own rollouts visit, labelled with the simulator's true next state and added to the log (two rounds) |
| multi-step loss | the squared error after 1 to 10 self-fed steps, the gradient passed back through the unrolled steps (checked against finite differences); a cousin of PlaNet's latent overshooting (Hafner et al., 2019), which its final agent works without |
| noisy inputs (GameNGen-style) | the logged states plus noise (3 cm, 2 cm/s) as inputs, the clean next state as the target |
One trained network proves nothing: another seed of the plain recipe moves the error at 81 steps between 5.9 and 11.7 cm (median 6.6, a factor 2.0 apart). So each fix is compared with the plain recipe on the same seed, over 12 seeds: the median of the paired ratios of the median errors, and the number of seeds on which the fix wins.
| fix | error at 81, ratio to plain | wins of 12 | error at 30, ratio to plain | wins of 12 |
|---|---|---|---|---|
| relabelled visits | 1.01 | 6 | 1.14 | 3 |
| multi-step loss | 1.00 | 6 | 1.11 | 2 |
| noisy inputs | 3.7 | 0 | 1.9 | 0 |
Against the plain recipe, then: relabelled visits and the multi-step loss leave the 81-step error where it was (median ratios 1.01 and 1.00; ahead on 6 and 6 of the 12 seeds) and make the 30-step error worse (1.14 and 1.11; ahead on 3 and 2); noisy inputs are worse on every seed, 3.7 times at 81 steps. None of the three lengthens the horizon, for this network, this task and these 20 epochs; longer training or other settings were not tried. For relabelling the reason can be checked. It pays when the network is worse at the states its own rollouts reach than at real states, and it is not: the median one-step error at the visited coasting states is 0.29 mm, against 0.30 mm at the real ones (ratio 0.99), which fits §2, where the recursion with real-state δ was already exact. The relabelled rollouts also start after the nudge, so the nudge step, the largest single contributor in §2, is never relabelled. For the multi-step loss there is the null result and no explanation; its gradient is checked against finite differences, so the null result is not a bug in the gradient. Noise does damage because a noisy four-number state is still a state, and a ball there moves differently: trained to map noisy inputs to the clean next state, the network learns to shrink velocities (its velocity entry falls from 0.9653 to 0.9604), so it lets the ball travel 13 % less far. In a picture the trick can work, since pixels are redundant and a network can learn to pull a noisy frame back toward a clean one (GameNGen reports that without noisy context its rollouts degrade fast after 20 to 30 steps). A four-number state has no redundancy to pull toward.
6 · Re-observe: the clock restarts
The sum has a start. Replace the model's state by the real one every m steps (a reading, or a belief: lesson 2's filter doing the job) and the product of Jacobians restarts, so the error depends only on the steps since the last reading. For the network of §1, an exact reading every 10 steps keeps the median error under 1.37 cm, where open loop it reaches 6.56 cm at 81; every 5 steps, 0.72 cm. In the lattice, where the open-loop error from a 1 cm start is 1.23 m at step 81, a reading with a 1 cm error every 10 steps keeps it under 1.53 cm. The cap is the error of an m-step rollout plus the error of the reading: a horizon of m, repeated. MBPO (Janner et al., 2019) does the same with a model, branching short rollouts from real states; its authors found 200-step rollouts worse than short ones. A real reading is noisy (σ = 10 cm in the Courtyard) and sometimes missing (behind the curtain): the reset is then lesson 2's posterior, whose error sets the floor instead of zero.
The widget
What to try. Start with the learned network, teacher forcing, k = 30, tolerance 5 cm. The error is 3.46 cm, 116 times the one-step error of 0.30 mm; the dashed recursion lies on the solid curve (ratio 0.9996) and both cross the red line at step 52. Drag k to 81: the dots move out through the rings and the median reaches 6.56 cm. A tolerance of 2 cm gives a horizon of 13; of 10 cm, none. Retrain: noisy inputs give 29 cm at 81, relabelled visits 6.16 and the multi-step loss 7.20, against 6.56 for teacher forcing (one seed, so the last two lie inside the spread that §5 measures over twelve). The exact model on the plain Courtyard with a 1 cm start error: 1.47 cm at 30 and 2.00 cm at 81, the amplification under 3.03. With the 23 posts the twelve amber copies fan out, the error is 39.5 cm at 30 and 1.23 m at 81, the horizon 14; with a 1 mm start 38, with 0.1 mm none. Re-observe every 10 steps: the error is a sawtooth that never passes 1.37 cm (lattice 1.53 cm). Last, set the learned network's nudges to "shifted 1 m/s": the error at 81 becomes 53 cm and the horizon drops from 52 to 11 (§7).
7 · The horizon was measured on the log's own actions
Everything above used nudges like the log's, and those are not arbitrary. The operator aims the stopping point at the goal, so a nudge is nearly a linear function of the state it met: regressing 2000 logged nudges on that state explains 74 % of their variance; the rest is habit. The network has seen 60 nudges, each the right one for its state, and nothing like a nudge the state does not call for.
Ask it about one. Shift every held-out nudge by 0.5 or 1 m/s in a random direction (the last control of the widget) and compare the network's imagined 81-step ending with the simulator's. For the operator's own nudges 1 % of imagined endings are off by more than the goal's radius (50 cm); for a shift of 0.5 m/s 5 %; for 1 m/s 51 % (25 % even among the 51 launches whose true path never touches a wall, so it is not only the walls). The median error at 81 steps is 53 cm against 6.6 cm, 8.1 times larger, and the 5 cm horizon falls from 52 to 31 and 11 steps. The recursion still predicts those horizons to within a step: the mechanism is unchanged, only δ at the nudge step is 4.0 times larger. A held-out test drawn from the log would have passed.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
A rollout now has a trustworthy horizon that can be estimated from one-step data (52 steps at 5 cm for the trained network, 13 at 2 cm), and re-observing every ten steps holds the error under 1.37 cm. But every model so far was fitted to logs that someone generated by acting, and the horizon was measured on actions like the log's. Shift the nudges by 1 m/s and it falls to 11 steps, 51 % of the imagined endings miss the goal's radius, and a test on held-out log data would not have said so. If whoever acted reacted to something the model cannot see, such as a gust of wind, the logs confound what an action did with why it was taken. Can a model fitted to logs tell what an action causes?
Interview prompts
- What is exposure bias, and how does teacher forcing cause it? (§1 — the model is trained on real inputs and used on its own outputs; its loss measures δ at real states and nothing else.)
- Derive the recursion for the error of a rollout and say when it grows linearly and when geometrically. (§2, §3 — add and subtract f̂(sk): ek+1 ≈ Jkek + δk; linear for ρ = 1 (free flight), geometric for ρ > 1 (round posts).)
- Define a trustworthy horizon and say how to estimate it without rolling out. (§4 — the first step at which the median error exceeds a tolerance; run the recursion with δ and J measured at real states.)
- Would DAgger-style relabelling, multi-step training or noise injection help your model, and how would you know? (§5 — compare with the plain recipe on the same seeds; they help only if δ at visited states is worse than at real states, and here it is not.)
- How does re-observation change a rollout's error? (§6 — it restarts the Jacobian product, so the error is capped by its m-step value; MBPO's short branched rollouts are the same lever.)
- A model passes a held-out test drawn from its training log. What has been checked? (§7 — only the actions the log contains; for nudges shifted by 1 m/s the horizon fell from 52 to 11 steps.)
Companion reads: Lesson 21 · From teacher forcing to rollout (the same fixes as training recipes at scale) and Robot Model Training · 03 Labels on your own states: DAgger and corrections (the policy-side version of exposure bias).