Think coarser to see farther
Lesson 9's planner got past the wall by looking about 50 to 60 steps ahead. It can look past what its model is trustworthy for, but every step is paid for at every decision, and a better model lowers a step's error, not their number. Predict in jumps of s steps and the error is paid once per jump: at stride 8 a small learned model keeps a 5 cm tolerance for 24 steps if told nothing about what a jump skips and for 64 if it also forecasts the wall hits it holds; at stride 64, nothing against at least 192. But a plan that commits for 16 steps cannot steer inside them: the planner (it calls the simulator for its jumps; the learned models know the walls only) then docks in at best 12 of 20 episodes. Skills give the reaction back; their subgoals we supply.
New idea: predict the jump, not the step, and hand the jump a summary of the fast events it skips: the error is paid once per jump, so reach grows with the stride until the decision itself needs finer steps.
Forces next: The mechanism is complete in a toy where the state is four numbers and the observation is a noisy dot. Real observation is a video: a million numbers per frame, almost all of them irrelevant, and the same ball can be drawn a million ways. Each piece we built, state, dynamics, uncertainty and rollouts, has to survive that. What happens to a world model when the observation is pixels at internet scale?
1 · Price the horizon
Lesson 9's planner (60 plans, 5 rounds of the cross-entropy method, receding horizon) docks the ball in the maze of task T5 only if it looks past the wall. Re-run here on 20 starts within ±10 cm of (1, 1), it docks in 0, 0, 18 and 20 episodes at H = 12, 25, 50 and 60 steps, spending 1,159,200 model steps an episode on average at 60 (18,000 a decision). Lesson 9's model, with 15 % too much friction, was off by 5 cm along its own plans after 12 steps and the loop docked anyway, taking the far end of a plan as a direction. Accuracy is not what binds: all 60 steps are paid for at every decision, and a better model lowers the error of a step, not their number.
To believe a 60-step plan as a promise and not only as a direction, the model must be right to lesson 6's tolerance, τ = 5 cm, at its last step. If each application adds at most δ and errors add, a 60-step plan needs δ ≤ τ/H = 0.83 mm. If each application also multiplies the error already there by ρ, the error is δ(ρH − 1)/(ρ − 1), the shape of the bound of Asadi, Misra and Littman (2018) for an n-step model, and the budget shrinks to δ ≤ τ(ρ − 1)/(ρH − 1): 0.14 mm at ρ = 1.05, since 1.0560 = 18.7.
Our model is a small learner, so that every stride refits in a blink: 64 tanh features with fixed random weights and a ridge readout, fitted in closed form to 1,500 pairs (state, state after the jump) from launched balls in the plain Courtyard (four walls, no posts), where every jump has a ground truth. At stride 1 its median error is 6.3 mm a step, 7.5 times what a 60-step plan can afford, and its chained rollout (its own output fed back) stays within 5 cm for 6 steps, close to the 7 that the additive law allows. Three levers: a smaller δ (a bigger model, the road of §4), a smaller H (lesson 9 needs about 60), or fewer applications. If one application covers s steps, J of them cover sJ, and the budget Jes ≤ τ buys a reach of sτ/es: with the stride-1 error, s × 7 steps, so a 60-step plan needs s ≥ 9. Does es stay where it was?
2 · Predict the jump
A jump model fs maps a state (x, y, vx, vy) to the state s steps later (s is the stride); its training pairs are free (a launched ball at a random moment of its first 9.6 s, the simulator run s steps), but each stride is a separate fit. To act, a macro-action is a nudge at the first step and then s − 1 steps of coasting, so the model is applied to the state after the nudge, which may be as large as the sum of the s per-step limits, 0.6s m/s. We measure on balls in their first 4.8 s, when they move. The one-jump error es is the median distance between predicted and true position over 2,000 fresh jumps; the chained error feeds the model's output back J times from 100 balls; the reach, lesson 6's trustworthy horizon in base steps, is s times the largest J whose median chained error is within 5 cm. Told nothing about what happens inside a jump:
| stride s (steps) | 1 | 2 | 4 | 8 | 16 | 32 | 64 |
|---|---|---|---|---|---|---|---|
| one-jump error es (mm) | 6.3 | 6.2 | 6.4 | 9.1 | 21.7 | 58.3 | 118.0 |
| reach at 5 cm (steps) | 6 | 16 | 24 | 24 | 16 | 0 | 0 |
Had es stayed at its stride-1 value, 7 jumps of 8 steps would see 56. It did not: flat to stride 4, then 1.4, 3.5, 9.3 and 19 times larger at 8, 16, 32 and 64. The reach rises from 6 to 24 steps and falls to nothing, since at stride 32 even one jump is off by more than 5 cm. Twelve draws of the random features and the data (the page's is the sixth) agree: the best reach without a summary is between 16 and 32 steps, and at stride 64 12 of 12 have no jump within tolerance.
3 · Why the floor rises: the folds
A jump is simple until a wall. In free flight friction is linear: over T = 0.1s seconds the velocity is multiplied by d = e−γT (friction γ = 0.35 s−1) and the position moves by c times the velocity, c = (1 − e−γT)/γ, and a jump without a hit is exactly that map. A hit on the right wall at x = w reflects the unfolded end point u = x + cvx: x′ = w − ε(u − w), with ε = 0.9. So the map from a state to the state s steps later is piecewise affine, folded along every line where u meets a wall, and the slope of x′ in x changes from +1 to −ε: 1.00 and −0.90 on the simulator at stride 8, vx = 3 m/s. The formula holds to 8.1 mm over 551 single-bounce jumps (median 2.3; the simulator's walls reflect inside 0.02 s sub-steps).
A smooth fit cannot follow a fold, and folds are rare enough for the median to miss and dear enough to matter. The share of jumps with a wall hit is 1.6 % at stride 8, 3.3 % at 16, 10.9 % at 32 and 29.5 % at 64. Without a summary the model is off by 130 mm on those at 16 and 201 mm at 64, so at stride 8 the 1.6 % carry 55 % of the squared error. Least squares serves the largest residuals first, by bending the whole fit, and the jumps without a hit pay: 9.0 mm at stride 8, 21.2 at 16, 99.6 at 64 (6.3 at stride 1). Fit the learner's readout to the no-hit jumps alone and on such jumps it is back to 7.1 and 9.6 mm at 16 and 64: the folded jumps spoil the others. Tell the model which piece of the map it is on.
4 · Match each channel's rate to the decision's
The ball has two channels at two rates: its state changes slowly (friction's time constant is 1/γ = 2.9 s) and is read once per jump; the walls act in an instant and a jump samples them not at all. Keeping every s-th observation and nothing else, decimation, is the model of §2. The repair is a summary of the fast channel: the count of hits (nx, ny) in the jump, which names the piece of the map. Each count seen at least 12 times gets its own readout on the same 64 features, fitted to the jumps with that count. Three versions: none, as above; forecast, where a count head predicts the counts and the matching readout the state, which is what a planner can do, since it proposes a jump before it happens; observed, where a fast loop that watched the jump supplies the true counts, the most such a summary can give.
| stride s (steps) | 1 | 2 | 4 | 8 | 16 | 32 | 64 |
|---|---|---|---|---|---|---|---|
| es, forecast (mm) | 6.3 | 6.1 | 6.0 | 6.3 | 7.3 | 9.4 | 9.6 |
| reach, forecast (steps) | 7 | 14 | 32 | 64 | 144 | ≥ 192 | ≥ 192 |
With the summary the error per jump stays on a floor, 6.0 to 9.6 mm from stride 1 to 64 against 6.2 to 118.0 without, and reach follows the stride: 64 steps at 8, 144 at 16, the whole 192-step window from 32 (observed counts add little: 72 and 160). The chained error after 64 steps is 4.8 cm at stride 8, 2.6 at 16, 1.6 at 32 and 1.0 at 64, within 17 % of Jes at every stride: the budget of §1 holds once the floor is flat. In all twelve draws the forecast beats no summary at strides 8 to 64; at 8 the median over draws is 56 steps (24 to 96) against 24.
Three details. At stride 8 the count head says "no hit" every time, though 1.6 % of jumps have one; at 16, 32 and 64 it finds 38, 53 and 84 % of them. The gain there comes from the no-hit readout, no longer bent by the folds, serving the other 98.4 %; the jumps with a hit stay wrong (98 mm), and observed counts mend them from stride 16, where the data hold enough to fit a readout (19 mm against 130). Reach is a median: at stride 8 with the forecast 48 % of the 100 chains are beyond 5 cm after 64 steps and the 90th percentile is 33 cm; the 10 worst chains all contain a wall hit (27 % of the chains do), which the forecast never sees at this stride.
5 · Plan in jumps: the maze, and what a jump cannot do
Now lesson 9's planner chooses macro-actions: J nudges, each at most 0.6s m/s and followed by s − 1 steps of coasting; the cost is lesson 9's summed over jump ends (0.2s times the distance after each, 3 times the last); the cross-entropy method searches the 2J numbers with the same 60 plans and 5 rounds; the ball runs the first jump with no look inside it, we look, and plan again. Its jump model is the simulator itself, posts included: the learned models of §2 to §4 know the walls and not the posts. So the maze measures one side of a bargain, how many jumps the planner needs; the study of §2 to §4 measured the other, how many a learned model can be trusted for. Twenty episodes a cell:
| stride s | planner (simulator jumps): smallest J tried that docks 18 of 20 | learned model (walls only): Jmax, no summary | learned model: Jmax, forecast |
|---|---|---|---|
| 1 | 48 | 6 | 7 |
| 2 | 16 | 8 | 7 |
| 4 | 6 | 6 | 8 |
| 8 | 2 | 3 | 8 |
| 16, 32, 64 | none; best 12, 5 and 3 | 1, 0, 0 | 9, 6, 3 |
Down to stride 2 the planner needs more jumps than any model supplies; at 4 and 8 the two sides meet; past 8 no J docks. The planner finds the way (at its best J it enters the goal disc in 20, 20 and 19 of 20 episodes at strides 16, 32 and 64) and cannot stop in it: at the docking speed of 0.5 m/s the ball crosses the 0.9 m disc in 18 steps, so a planner that looks every 8 steps sees it twice, every 16 about once, every 64 almost never. Search is not the cause: ten times the evaluations dock 11 of 20 at stride 16, and eight times the evaluations leave the flat planner at H = 25 at 0 of 10. What a jump gives up is the next decision: with the summary, the price of jumping is paid in control and not in accuracy.
In this maze the summary is not what makes stride 8 feasible: the unsummarised model already supports 3 jumps and the plan needs 2; it buys margin (8 against 2). What coarse planning saves is calls, counting a jump as one call (the maze runs the simulator for it): 18,000 model steps a decision for the flat planner at H = 60, 900 jump calls at stride 8 with J = 3, 20 times fewer, and it docks 20 of 20 in a median of 49 steps against the flat planner's 64 (lesson 9's 66 is a mean).
6 · Skills and subgoals: reaction inside the jump
A jump is open loop; a skill is closed loop: "go to the waypoint" looks at the ball at every step, so it keeps what the jump gave up. Temporal abstraction in the planner then means choosing among skills, each aimed at a subgoal and run by the fine planner. In the maze: "go to (4.0, 4.45)", the middle of the gap, then "go to the goal". The low level is lesson 9's planner, aimed at the current subgoal with horizon Hlow; the high level, written by hand, switches when the ball is within 0.4 m of it. It docks in 20 of 20 episodes at every low-level horizon tried, Hlow = 6, 8, 12 and 25 (medians 32, 36, 46 and 51 steps); the flat planner docks 0 of 20 at 25, so the subgoal does the work of the long horizon. At Hlow = 8 an episode costs 88,080 model steps against 1,159,200 for the flat planner at 60, 13 times fewer.
The saving has a counting argument. With b choices a step, a flat search over H steps has bH leaves. Split into segments of s steps, search the H/s segment-level choices (bH/s) and inside each segment its own s-step tree (bs): bH/s + (H/s)bs in all, smallest near s = √H (the two terms balance there). With b = 5 nudges (stay, or one of four compass directions) and H = 64: 1044.7 flat, 3,515,625 for two levels of 8, a factor of 1038.2. Levels also multiply reach: if each keeps 8 jumps within tolerance, stride 8 reaches 64 steps, a second level whose step is that jump 512, a third 4,096, the thousands of steps that crossing a building or cooking a meal takes. That is arithmetic: it assumes the floor stays flat at every level, which needs the right summary at every level, and we did not build it.
Nor did we find the subgoal: the waypoint is where we know the gap is, and discovering skills and subgoals from experience is open. LeCun (2022) proposes H-JEPA, the hierarchical form of his predictive architecture, for hierarchical planning under uncertainty; it is a proposal, untested here. Two other levers shorten the horizon instead of lengthening the step: TD-MPC (Hansen, Wang and Su, 2022) plans 5 steps and lets a learned value stand for the rest, and MBPO (Janner et al., 2019) branches short model rollouts from real states because long ones were worse.
The widget
What to try. Told nothing, slide the stride from 1 to 64: the reach reads 6, 16, 24, 24, 16, 0 and 0 steps, and the red curve leaves the 5 cm line at the first jump from stride 32. Choose the forecast: 7, 14, 32, 64, 144, then the whole window. At stride 8 with a plan of 24 steps (3 jumps) episode 1 docks at step 33 on 4,500 simulator calls, and the learned model's error at the plan's end is 3.6 cm without a summary, inside the tolerance. At stride 64 with 64 steps (one jump) the ball enters the goal and stays 16 steps, never slower than 0.68 m/s: no dock. Stride 1 with 64 steps docks at step 54 on 1,036,800 calls. With the plan at 8 steps and the planner on skills it docks at step 36 on 86,400 model steps, at any stride.
7 · Where the toy ends
The state was four numbers and a jump a map of four numbers. A frame of video 256 × 256 × 3 is 196,608 numbers, 49,152 times the state; in 16 × 16 patches it is 256 tokens, at 8 frames a second 2,048 a second, 122,880 a minute and 7.37 million an hour (the arithmetic of lesson 17). In lesson 3's picture the ball carried 19 % of the pixel variance and 81 % was a band and a brightness with nothing to do with it. A summary of skipped events needs events to count and a jump needs a state to jump from; pixels offer neither directly.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
With the summary a learned jump model sees 64 steps at stride 8, and a planner that jumps by 8 (calling the simulator) docks in 20 of 20 episodes on 900 calls a decision where the flat planner needs 18,000; skills keep the reaction. But the toy is four numbers and a noisy dot, a jump needs a state to jump from and events to count (§7), and a frame is 196,608 numbers, and 81 % of the variance in lesson 3's picture was nuisance. Each piece we built, state, dynamics, uncertainty and rollouts, has to survive that. What happens to a world model when the observation is pixels at internet scale?
Interview prompts
- Why does predicting in jumps extend a model's reach, and what must stay true? (§1, §2 — the error is paid once per application, so reach is s·τ/es, if es does not rise.)
- Why does a jump model's error rise with the stride when the dynamics are simple? (§3 — the map folds wherever a jump skips a bounce, and a smooth fit bends everywhere to serve the few folded jumps.)
- State the rate-matching rule. (§4 — sample each channel at the rate its physics demands, summarise the skipped fast events for the slow loop, run the slow loop at the rate the decision needs.)
- What does a jump give up, and why is it not accuracy? (§5 — reaction: a plan committed for 16 steps cannot slow the ball inside them, so it enters the goal and does not dock.)
- Why is a hierarchical search so much cheaper, and what does it assume? (§6 — bH/s + (H/s)bs against bH, smallest near s = √H; the subgoals were written by hand.)
- Why is a toy with a four-number state not enough for the next step? (§7 — a jump needs a state to jump from and events to count; a frame is 196,608 numbers, most of them nuisance.)
Companion reads: Reinforcement Learning · Model and planning (search with a model), Lesson 17 · Two seats, one tuple (the token arithmetic of §7) and Lesson 25 · Horizon, memory, and the quadratic bill (what a long horizon costs).