Practice in imagination
Lesson 7 ended with a model that says what a nudge causes, as far as the data went. The cheapest use of it is practice: let a policy improve on the model's imagined launches instead of in the world. In the world that costs 46,080 real launches; on the model, the 300 trials that fitted it. But an optimiser seeks whatever the model likes best. After 24 generations the policy's imagined return is 0.88 (92 % hits) and its real return 0.03 (0 % hits), below the 0.18 of doing nothing. This lesson derives why, penalises the model's own disagreement, puts real trials back where the policy goes, measures both repairs, and then changes the goal.
New idea: score a policy on the model by its imagined return minus the model's disagreement, and keep adding real trials from where the policy goes. Wherever the penalty covers the model's error it makes the penalised return a lower bound on the real one, so the optimiser cannot profit from the model's mistakes.
Forces next: Practice produces a policy that is only as good as the model was where it practiced. When the goal changes, or the situation is new, the policy has nothing to say. What if the agent used the model at the moment of decision, searching over what to do next, and what does it do about a model that is trustworthy for only a few steps?
1 · Practice, and the price of a perfect model
The task is lesson 1's nudge. A launcher fires the ball; at step 12 the agent is handed its state s = (x, y, vx, vy) (lessons 2 to 4 showed how it gets one) and chooses one nudge a = (ax, ay) with |a| ≤ 3 m/s; the ball then coasts for 80 more steps (8 s). Its ending e(s, a) is scored by the miss, the distance to the goal's centre (6.6, 2.5): R = exp(−½(miss/0.5)²), which is 1 on the centre, 0.61 at the rim of the 0.5 m goal and 0.14 at 1 m. A hit is a miss under 0.5 m.
A policy is ten numbers θ: a = clip3(A(s − s̄) + b), with A a 2×4 matrix, b two numbers and s̄ = (3.5, 2.5, 1.5, 0) a fixed reference state. The class holds a good policy. Friction (γ = 0.35 s−1) multiplies the velocity by d = 0.9656 per step, so while no wall is touched the ending is the position at the nudge plus c(v + a), with c = (1 − d81)/γ = 2.69 s. Setting that equal to the goal gives a* = (goal − position)/c − v, linear in the state and so in the class: it hits 100 % of 3000 fresh launches, with a median nudge of 0.58 m/s, and 97 % of its nudges are at most 1.2 m/s.
The exam is fixed: the same 60 launches, the policy's mean return and hit rate in the true simulator (the real return) and under a model (the imagined return). Doing nothing has a real return of 0.18 and 7 % hits. To improve θ we use the cross-entropy method, a search that needs no gradient. Draw 96 candidates around a mean μ (start μ = 0, "do nothing", spread 0.5), score each, refit μ and the spread to the best 16, repeat. After g generations it has tried K = 96g candidates, and it reports the best found so far.
Control. Let the model be the simulator itself, so the score is the real return. The optimiser reaches 100 % hits at generation 8: K = 768 candidates on 60 launches each, 46,080 real launches. That is the price of practising in the world, and the reason for a model. Dyna (Sutton, 1991) learns a model from experience and improves the policy on imagined experience as well; Dreamer (Hafner et al., 2020) does it in a learned latent space, with value gradients flowing through imagined trajectories 15 steps long; DreamerV2 (Hafner et al., 2021) trained its policy on 468 billion imagined states, about 10,000 times the 50 million real inputs it saw.
2 · A model fitted to a log, and where the data went
The log is a cautious operator's: 300 real trials, each a launch, a nudge drawn uniformly from the disc |a| ≤ 1.2 m/s (40 % of the actuator's range, 16 % of its area) and the ending. The model is lesson 1's M, learned: five small networks (tanh, two hidden layers of 8) map the state and nudge to the ending, each fitted to its own bootstrap resample of the log (lesson 5). The model's answer is the mean of the five; its doubt u is the largest distance of a member from that mean, in metres, and U is the mean of u over the 60 launches. How wrong is it, by nudge size? 300 fresh launches per column, nudges of that size in a random direction:
| nudge |a| (m/s) | 0–0.6 | 0.6–1.2 | 1.2–1.8 | 1.8–2.4 | 2.4–3 |
|---|---|---|---|---|---|
| in the log? | yes | yes | no | no | no |
| RMS error of the mean (m) | 0.15 | 0.22 | 0.66 | 1.42 | 2.32 |
| disagreement u (m) | 0.19 | 0.26 | 0.48 | 0.79 | 1.05 |
Inside the disc the error is 0.19 m; in the last column it is 16 times the first. The members notice, but not enough: u grows 5.6 times, and in the last column the error is 2.2 times u: the doubt understates the error, as in lesson 7. A policy needs the model only at the nudges it chooses. Nothing so far says it will stay inside the disc.
3 · The optimiser finds the mistakes
Practise against the model, scoring by the imagined return alone (the naive optimiser). At generation 0 the model says 0.17 for a policy that does nothing and the world says 0.18, right where the data went. At generation 24 the model says 0.88 with 92 % hits and the world says 0.03 with 0 % (one log, the one nearest the medians of twelve, which follow below). The real return never rose above where it started, 0.18, and fell to 0.013; all 100 % of the policy's nudges lie outside the log (the smallest is 1.41 m/s, the median 2.74), and the mean disagreement is 1.34 m. Stopping early does not help: the imagined return of the best policy so far cannot fall, it is what is maximised, and the real return peaked at generation 0. The widget opens on the launch where the flattery is largest: a nudge of 2.76 m/s, members that disagree by 1.28 m, a model that puts the ball 0.12 m from the goal, and a world that puts it 2.65 m away, in the corner at (7.82, 4.85).
The cause is selection, the optimiser's curse. Give candidate i a real return Ji and an imagined one Ĵi = Ji + ei, the errors being independent normals of standard deviation δ, and keep the largest Ĵ. (a) If the candidates are equally good, the winner is the one with the largest error and its optimism is δ·E[maxK], with E[maxK] the expected maximum of K standard normals. The textbook bound √(2 ln K) holds for every K but is tight only as K → ∞; at moderate K it overstates: E[maxK] (numerical integration, checked by simulation) is 1.54, 2.51 and 3.24 for K = 10, 100, 1000, against √(2 ln K) = 2.15, 3.03, 3.72 (39 %, 21 %, 15 % too large). (b) If the candidates really differ, with spread σJ, then J and Ĵ are jointly normal, the winner's Ĵ sits √(σJ² + δ²)·E[maxK] above the mean, and given Ĵ the expected real return is the fraction σJ²/(σJ² + δ²) of that. So
optimism = δ²/√(σJ² + δ²) · E[maxK], real gain = σJ²/√(σJ² + δ²) · E[maxK]
Is this what happened? Draw 768 policies blindly from the optimiser's starting distribution and score each by model and world. The errors have a spread of δ = 0.024 and the real returns σJ = 0.052, so 82 % of a blind winner's gain would be real. The blind best-of-K gap (imagined minus real) is 0.033, 0.044, 0.050 at K = 96, 192, 384, between the formula with the real share removed (0.026, 0.028, 0.031) and without (0.061, 0.067, 0.072): the formula brackets the measured gaps without predicting them, and blind selection pays a tax of a few hundredths. The optimiser's own gap at K = 384, generation 4, is 0.66, 13 times that, and 0.85 at generation 24. The law assumes errors drawn independently of the choice. The optimiser draws each generation from where the last one scored well, which selects favourable errors again and again: it does not sample the model's error, it searches for it, and what it finds is an error of metres, worth most of the imagined return and not a few hundredths of it.
Is a narrow log the cause? Not by itself. Repeat the whole procedure on twelve logs (naive optimiser, medians) with the operator's nudges stopping at 0.8, 1.2 and 3.0 m/s. The imagined return is 1.00, 0.88 and 0.99; the real return is 0.98, 0.10 and 0.91. Whether the optimiser's favourite nudges fall where the model's extrapolation happens to be harmless is a property of the model that imagination cannot see. With a log covering the whole actuator none of the twelve fails badly (worst 0.83), and a careful operator cannot supply that coverage. At 1.2, 11 of 12 logs end below 0.4 and none above 0.88.
4 · Be pessimistic where the model is unsure
If the optimiser climbs the model's errors, make them cost. Score a policy by Jλ = Ĵ − λU, the imagined return minus λ times the disagreement (λ in return per metre). MOPO (Yu et al., 2020) penalises each reward by λ times the largest standard deviation any ensemble member predicts; MOReL (Kidambi et al., 2020) sends every state-action pair on which members disagree beyond a threshold to an absorbing halt state with reward −κ. Both bound the real return from below by the penalised one, and the argument is short when stated for whole policies instead of single state-action pairs. Call the penalty admissible if λU(θ) ≥ |Ĵ(θ) − J(θ)| for every policy θ. Then (i) J ≥ Jλ, because Ĵ − J ≤ λU: the penalised return is a lower bound. (ii) The maximiser θ̃ of Jλ has J(θ̃) ≥ Jλ(θ̃) ≥ Jλ(θ) ≥ J(θ) − 2λU(θ) for every θ, using (i), the maximiser, and Ĵ ≥ J − λU: the pessimist is as good as any policy minus twice its own doubt. (A hard version, forbidding nudges outside the log's disc, would forbid part of the solution: the control's best policy has 18 % of its nudges beyond 1.2 m/s, the largest 1.44.)
Is λ = 1 admissible here? On 1225 policies (random ones, and the policies of the naive run) the smallest admissible λ is 0.63; at λ = 0.25 and 0.5, 2.4 % and 1.6 % of policies violate it, at λ = 1 none do. The control's policy has a real return of 1.00 and U = 0.20 m under the model, so (ii) promises 0.60. Practising on J1 reaches 0.98 (imagined 0.99, 100 % hits, U = 0.20 m), and only 18 % of its nudges lie outside the log, against 100 %.
λ is a dial, not a free lunch. Median real return over the twelve logs, with the worst log in brackets: λ = 0.25, 0.98 (0.07); 1, 0.99 (0.98); 4, 0.98 (0.55); 16, 0.63 (0.10) with 56 % hits. Too small a penalty is not admissible and some log still has a hole to find. A large one makes every nudge off the data cost more than it earns and the policy turns timid. At λ = 4 the worst log still ends at 0.55; we did not test why, and a poor local optimum of the rougher penalised surface is only our guess. The same failure is on record. Ha and Schmidhuber (2018) trained a controller inside a learned model of VizDoom Take Cover, scored in steps survived: at sampling temperature τ = 0.1 it scored 2086 in the model and 193 in the real game, at τ = 1.0 1145 and 868; a random policy scores 210. A cold dream lets the controller exploit the model's flaws and a hotter one is harder to exploit; ours is a penalty.
5 · Put the world back in the loop
The second repair is Dyna's own. Every 6 generations the best policy so far runs on 30 fresh real launches, with exploration noise of 0.3 m/s on its nudge; the trials join the log, the five networks are refitted for 40 epochs on new bootstrap resamples, the incumbent is re-scored by the new model, and the optimiser's spread is widened again. Three rounds add 90 trials, 390 in all. The real return after generations 6, 12, 18 and 24 is 0.02, 0.20, 0.43 and 0.90, the imagined 0.01, 0.13, 0.37 and 0.82: each refit turns the optimiser's favourite hole into data. Between rounds the model serves 6 × 96 × 60 = 34,560 imagined launches for 30 real ones, 1,152 to 1; over the run, 354 to 1.
| practice against | real trials | imagined return | real return | worst of 12 logs |
|---|---|---|---|---|
| the simulator (control) | 46,080 launches | — | 1.00 | — |
| the model | 300 | 0.88 | 0.10 | 0.03 |
| the model, penalised (λ = 1) | 300 | 0.99 | 0.99 | 0.98 |
| the model + real trials | 390 | 0.79 | 0.84 | 0.13 |
| both repairs | 390 | 0.97 | 0.97 | 0.95 |
Each repair has a price. The penalty costs no data and wins most of the gap, but it protects only as far as the members' disagreement tells the truth, and it gives up the nudges the log lacks. Real trials fix the model where the policy goes and cost 90 more, but the first round comes when the optimiser is already in a hole (real return 0.02 at generation 6), and in 2 of 12 logs the policy still ends below 0.5. Together they leave even the worst of twelve logs at 0.95, for 390 real trials against 46,080 real launches for the control. A log covering the whole actuator also avoids the failure (§3); the repairs are for when it does not.
The widget
What to try. Pick the simulator and drag the generations: real hits reach 100 % at generation 8 (K = 768; the data readout counts 46,080 launches). Pick the model: at 24 it reads 0.88 imagined, 0.03 real; the gap is 0.21 at generation 1 and 0.66 at generation 4, and the ticks under the bars end beyond 1.41 m/s, off the green band. Penalise it (λ = 1): 0.99 and 0.98, 18 % of nudges outside the log; at λ = 16 it turns timid (0.63 real). With real trials, step through 6, 12, 18, 24: real 0.02, 0.20, 0.43, 0.90, a purple line at each refit; both repairs end at 0.97. Set the operator to 0.8 m/s: 1.00 imagined, 0.96 real; to 3.0: 1.00 and 0.88. Finally select the simulator and raise the goal: its policy falls to 0 % hits, and the searched-afresh readout shows 83 % real (88 % imagined), its nudge the green path.
6 · What a practised policy is an answer to
Move the goal up 1.5 m, to (6.6, 4.0). A launch that was on target now needs an extra nudge of (4.0 − 2.5)/c = 0.56 m/s in y, far inside the actuator. The policy practised against the simulator and the one with both repairs hit 100 % at the old goal and 0 % and 0 % at the new one (return 0.013): θ is an answer to one question, goal included. The model never saw a goal. Search it afresh, as lesson 1's planner searched the simulator: for each of the 60 launches draw 128 random nudges, keep the one with the best penalised imagined return (λ = 1), run it. The model says 88 % hits and the world 83 % (twelve logs: median 87 %, worst 82 %). Without the penalty the model says 95 % and the world 78 %, so the penalty matters at decision time too. The model has become a tool that answers new questions: 60 × 128 = 7,680 queries, no practice.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
A policy can now practise on a model without being fooled by it: 0.97 real return on 390 trials instead of 46,080 launches. But the policy it produces answers one question. Move the goal and the same policy that hit 100 % hits 0 %, while the model, searched afresh at the moment of decision, hits 83 %. That search worked because one decision is two numbers and the model was trusted for the whole ending; with many decisions in a row, and a model that is right for only a few steps, it is not. What if the agent used the model at the moment of decision, searching over what to do next, and what does it do about a model that is trustworthy for only a few steps?
Interview prompts
- Why does a policy trained on a learned model often do worse in the world than the model predicts? (§3 — it selects the model's errors: with equal candidates the best of K estimates is optimistic by δ·E[maxK], and an adaptive search finds the worst hole.)
- How loose is the √(2 ln K) law, and what changes when candidates really differ? (§3 — E[max] is smaller at moderate K; only σJ²/(σJ² + δ²) of the apparent gain is real.)
- Derive the guarantee of a MOPO-style penalty. (§4 — if λU ≥ |Ĵ − J| then J ≥ Ĵ − λU, and the penalised maximiser is within 2λU(θ) of any θ.)
- Why not stop optimising early? (§3 — the imagined return of the best policy so far never falls, so it gives no signal.)
- What does Dyna add, and what does it cost? (§5 — real trials from where the policy goes; no protection before the first round.)
- Why does a practised policy fail when the goal changes while the model need not? (§6 — θ encodes one goal; the model is queried per decision.)
Companion reads: Reinforcement Learning · 16 Offline RL (pessimism for logged data, without a model) and Lesson 21 · From teacher forcing to rollout (why long rollouts drift).