all_lessons/World Models/10 · Temporal abstractionlesson 10 / 31

Think coarser to see farther

Lesson 9's planner got past the wall by looking about 50 to 60 steps ahead. It can look past what its model is trustworthy for, but every step is paid for at every decision, and a better model lowers a step's error, not their number. Predict in jumps of s steps and the error is paid once per jump: at stride 8 a small learned model keeps a 5 cm tolerance for 24 steps if told nothing about what a jump skips and for 64 if it also forecasts the wall hits it holds; at stride 64, nothing against at least 192. But a plan that commits for 16 steps cannot steer inside them: the planner (it calls the simulator for its jumps; the learned models know the walls only) then docks in at best 12 of 20 episodes. Skills give the reaction back; their subgoals we supply.

The thesis, here
A model's error is paid once per application, and a better model lowers the error of each, not their number; so a model that predicts s steps at a time needs s times fewer and reaches about s times as far, if its error per application stays flat. Unaided it does not: a jump skips fast events (here, wall bounces), the map from a state to the state s steps later folds where they happen, and the few folded jumps bend a smooth fit everywhere. A count of the skipped events restores the floor. The rule is to match each channel's rate to the decision's: sample every channel at the rate its physics demands, hand the slow loop a summary of what it skips, and run it no slower than the decision needs. That last clause is the price: a plan committed for s steps cannot react inside them. Skills, closed-loop sub-policies, buy the reaction back if someone supplies the subgoals.
Linear position
Forced by: A planner can look past the horizon its model is trustworthy for, because it re-plans and uses the far end of a plan as a direction. But the maze already needed about fifty steps of look-ahead, each paid for at every decision, and the tasks that matter, crossing a building or cooking a meal, take thousands. A better model lowers the error of each step, not the number of steps; the way to look farther at the same price is to take fewer, bigger steps. How can a model predict in jumps, and what does it lose?
New idea: predict the jump, not the step, and hand the jump a summary of the fast events it skips: the error is paid once per jump, so reach grows with the stride until the decision itself needs finer steps.
Forces next: The mechanism is complete in a toy where the state is four numbers and the observation is a noisy dot. Real observation is a video: a million numbers per frame, almost all of them irrelevant, and the same ball can be drawn a million ways. Each piece we built, state, dynamics, uncertainty and rollouts, has to survive that. What happens to a world model when the observation is pixels at internet scale?
The plan
Seven moves. (1) Price the horizon. (2) Predict s steps at once. (3) Find why the error rises with s. (4) Match each channel's rate to the decision's. (5) Plan in jumps in the maze. (6) Skills and hierarchy. (7) See where the toy ends.

1 · Price the horizon

Lesson 9's planner (60 plans, 5 rounds of the cross-entropy method, receding horizon) docks the ball in the maze of task T5 only if it looks past the wall. Re-run here on 20 starts within ±10 cm of (1, 1), it docks in 0, 0, 18 and 20 episodes at H = 12, 25, 50 and 60 steps, spending 1,159,200 model steps an episode on average at 60 (18,000 a decision). Lesson 9's model, with 15 % too much friction, was off by 5 cm along its own plans after 12 steps and the loop docked anyway, taking the far end of a plan as a direction. Accuracy is not what binds: all 60 steps are paid for at every decision, and a better model lowers the error of a step, not their number.

To believe a 60-step plan as a promise and not only as a direction, the model must be right to lesson 6's tolerance, τ = 5 cm, at its last step. If each application adds at most δ and errors add, a 60-step plan needs δ ≤ τ/H = 0.83 mm. If each application also multiplies the error already there by ρ, the error is δ(ρH − 1)/(ρ − 1), the shape of the bound of Asadi, Misra and Littman (2018) for an n-step model, and the budget shrinks to δ ≤ τ(ρ − 1)/(ρH − 1): 0.14 mm at ρ = 1.05, since 1.0560 = 18.7.

Our model is a small learner, so that every stride refits in a blink: 64 tanh features with fixed random weights and a ridge readout, fitted in closed form to 1,500 pairs (state, state after the jump) from launched balls in the plain Courtyard (four walls, no posts), where every jump has a ground truth. At stride 1 its median error is 6.3 mm a step, 7.5 times what a 60-step plan can afford, and its chained rollout (its own output fed back) stays within 5 cm for 6 steps, close to the 7 that the additive law allows. Three levers: a smaller δ (a bigger model, the road of §4), a smaller H (lesson 9 needs about 60), or fewer applications. If one application covers s steps, J of them cover sJ, and the budget Jes ≤ τ buys a reach of sτ/es: with the stride-1 error, s × 7 steps, so a 60-step plan needs s ≥ 9. Does es stay where it was?

2 · Predict the jump

A jump model fs maps a state (x, y, vx, vy) to the state s steps later (s is the stride); its training pairs are free (a launched ball at a random moment of its first 9.6 s, the simulator run s steps), but each stride is a separate fit. To act, a macro-action is a nudge at the first step and then s − 1 steps of coasting, so the model is applied to the state after the nudge, which may be as large as the sum of the s per-step limits, 0.6s m/s. We measure on balls in their first 4.8 s, when they move. The one-jump error es is the median distance between predicted and true position over 2,000 fresh jumps; the chained error feeds the model's output back J times from 100 balls; the reach, lesson 6's trustworthy horizon in base steps, is s times the largest J whose median chained error is within 5 cm. Told nothing about what happens inside a jump:

stride s (steps)1248163264
one-jump error es (mm)6.36.26.49.121.758.3118.0
reach at 5 cm (steps)61624241600

Had es stayed at its stride-1 value, 7 jumps of 8 steps would see 56. It did not: flat to stride 4, then 1.4, 3.5, 9.3 and 19 times larger at 8, 16, 32 and 64. The reach rises from 6 to 24 steps and falls to nothing, since at stride 32 even one jump is off by more than 5 cm. Twelve draws of the random features and the data (the page's is the sixth) agree: the best reach without a summary is between 16 and 32 steps, and at stride 64 12 of 12 have no jump within tolerance.

3 · Why the floor rises: the folds

A jump is simple until a wall. In free flight friction is linear: over T = 0.1s seconds the velocity is multiplied by d = e−γT (friction γ = 0.35 s−1) and the position moves by c times the velocity, c = (1 − e−γT)/γ, and a jump without a hit is exactly that map. A hit on the right wall at x = w reflects the unfolded end point u = x + cvx: x′ = w − ε(u − w), with ε = 0.9. So the map from a state to the state s steps later is piecewise affine, folded along every line where u meets a wall, and the slope of x′ in x changes from +1 to −ε: 1.00 and −0.90 on the simulator at stride 8, vx = 3 m/s. The formula holds to 8.1 mm over 551 single-bounce jumps (median 2.3; the simulator's walls reflect inside 0.02 s sub-steps).

A smooth fit cannot follow a fold, and folds are rare enough for the median to miss and dear enough to matter. The share of jumps with a wall hit is 1.6 % at stride 8, 3.3 % at 16, 10.9 % at 32 and 29.5 % at 64. Without a summary the model is off by 130 mm on those at 16 and 201 mm at 64, so at stride 8 the 1.6 % carry 55 % of the squared error. Least squares serves the largest residuals first, by bending the whole fit, and the jumps without a hit pay: 9.0 mm at stride 8, 21.2 at 16, 99.6 at 64 (6.3 at stride 1). Fit the learner's readout to the no-hit jumps alone and on such jumps it is back to 7.1 and 9.6 mm at 16 and 64: the folded jumps spoil the others. Tell the model which piece of the map it is on.

4 · Match each channel's rate to the decision's

The ball has two channels at two rates: its state changes slowly (friction's time constant is 1/γ = 2.9 s) and is read once per jump; the walls act in an instant and a jump samples them not at all. Keeping every s-th observation and nothing else, decimation, is the model of §2. The repair is a summary of the fast channel: the count of hits (nx, ny) in the jump, which names the piece of the map. Each count seen at least 12 times gets its own readout on the same 64 features, fitted to the jumps with that count. Three versions: none, as above; forecast, where a count head predicts the counts and the matching readout the state, which is what a planner can do, since it proposes a jump before it happens; observed, where a fast loop that watched the jump supplies the true counts, the most such a summary can give.

stride s (steps)1248163264
es, forecast (mm)6.36.16.06.37.39.49.6
reach, forecast (steps)7143264144≥ 192≥ 192

With the summary the error per jump stays on a floor, 6.0 to 9.6 mm from stride 1 to 64 against 6.2 to 118.0 without, and reach follows the stride: 64 steps at 8, 144 at 16, the whole 192-step window from 32 (observed counts add little: 72 and 160). The chained error after 64 steps is 4.8 cm at stride 8, 2.6 at 16, 1.6 at 32 and 1.0 at 64, within 17 % of Jes at every stride: the budget of §1 holds once the floor is flat. In all twelve draws the forecast beats no summary at strides 8 to 64; at 8 the median over draws is 56 steps (24 to 96) against 24.

Three details. At stride 8 the count head says "no hit" every time, though 1.6 % of jumps have one; at 16, 32 and 64 it finds 38, 53 and 84 % of them. The gain there comes from the no-hit readout, no longer bent by the folds, serving the other 98.4 %; the jumps with a hit stay wrong (98 mm), and observed counts mend them from stride 16, where the data hold enough to fit a readout (19 mm against 130). Reach is a median: at stride 8 with the forecast 48 % of the 100 chains are beyond 5 cm after 64 steps and the 90th percentile is 33 cm; the 10 worst chains all contain a wall hit (27 % of the chains do), which the forecast never sees at this stride.

The rule, in our words
Match each channel's rate to the decision's. Sample every channel at the rate its physics demands. Hand the slow loop a summary of the fast events it skips: dropping them does not remove them, it moves their cost into everything else (stride 16: 21.7 mm a jump against 7.3, reach 16 steps against 144). And run the slow loop at the rate the decision needs, no slower (§5 measures it).
Road not taken · a bigger model
Spend on the learner: 256 features and 6,000 jumps, four times each, 63 times the fitting arithmetic. The stride-1 error falls from 6.3 to 0.4 mm and the stride-1 reach from 6 to 21 steps (147 with the summary), but at stride 64 the error without a summary is still 54 mm and the reach 0. A better learner lowers every curve and leaves their shape alone; the stride multiplies whatever the floor is. For 60 steps a bigger model would do, in a toy with a four-number state; 2,400 steps (four minutes of Courtyard time) on its 147 need a stride of 17.

5 · Plan in jumps: the maze, and what a jump cannot do

Now lesson 9's planner chooses macro-actions: J nudges, each at most 0.6s m/s and followed by s − 1 steps of coasting; the cost is lesson 9's summed over jump ends (0.2s times the distance after each, 3 times the last); the cross-entropy method searches the 2J numbers with the same 60 plans and 5 rounds; the ball runs the first jump with no look inside it, we look, and plan again. Its jump model is the simulator itself, posts included: the learned models of §2 to §4 know the walls and not the posts. So the maze measures one side of a bargain, how many jumps the planner needs; the study of §2 to §4 measured the other, how many a learned model can be trusted for. Twenty episodes a cell:

stride splanner (simulator jumps): smallest J tried that docks 18 of 20learned model (walls only): Jmax, no summarylearned model: Jmax, forecast
14867
21687
4668
8238
16, 32, 64none; best 12, 5 and 31, 0, 09, 6, 3

Down to stride 2 the planner needs more jumps than any model supplies; at 4 and 8 the two sides meet; past 8 no J docks. The planner finds the way (at its best J it enters the goal disc in 20, 20 and 19 of 20 episodes at strides 16, 32 and 64) and cannot stop in it: at the docking speed of 0.5 m/s the ball crosses the 0.9 m disc in 18 steps, so a planner that looks every 8 steps sees it twice, every 16 about once, every 64 almost never. Search is not the cause: ten times the evaluations dock 11 of 20 at stride 16, and eight times the evaluations leave the flat planner at H = 25 at 0 of 10. What a jump gives up is the next decision: with the summary, the price of jumping is paid in control and not in accuracy.

In this maze the summary is not what makes stride 8 feasible: the unsummarised model already supports 3 jumps and the plan needs 2; it buys margin (8 against 2). What coarse planning saves is calls, counting a jump as one call (the maze runs the simulator for it): 18,000 model steps a decision for the flat planner at H = 60, 900 jump calls at stride 8 with J = 3, 20 times fewer, and it docks 20 of 20 in a median of 49 steps against the flat planner's 64 (lesson 9's 66 is a mean).

6 · Skills and subgoals: reaction inside the jump

A jump is open loop; a skill is closed loop: "go to the waypoint" looks at the ball at every step, so it keeps what the jump gave up. Temporal abstraction in the planner then means choosing among skills, each aimed at a subgoal and run by the fine planner. In the maze: "go to (4.0, 4.45)", the middle of the gap, then "go to the goal". The low level is lesson 9's planner, aimed at the current subgoal with horizon Hlow; the high level, written by hand, switches when the ball is within 0.4 m of it. It docks in 20 of 20 episodes at every low-level horizon tried, Hlow = 6, 8, 12 and 25 (medians 32, 36, 46 and 51 steps); the flat planner docks 0 of 20 at 25, so the subgoal does the work of the long horizon. At Hlow = 8 an episode costs 88,080 model steps against 1,159,200 for the flat planner at 60, 13 times fewer.

The saving has a counting argument. With b choices a step, a flat search over H steps has bH leaves. Split into segments of s steps, search the H/s segment-level choices (bH/s) and inside each segment its own s-step tree (bs): bH/s + (H/s)bs in all, smallest near s = √H (the two terms balance there). With b = 5 nudges (stay, or one of four compass directions) and H = 64: 1044.7 flat, 3,515,625 for two levels of 8, a factor of 1038.2. Levels also multiply reach: if each keeps 8 jumps within tolerance, stride 8 reaches 64 steps, a second level whose step is that jump 512, a third 4,096, the thousands of steps that crossing a building or cooking a meal takes. That is arithmetic: it assumes the floor stays flat at every level, which needs the right summary at every level, and we did not build it.

Nor did we find the subgoal: the waypoint is where we know the gap is, and discovering skills and subgoals from experience is open. LeCun (2022) proposes H-JEPA, the hierarchical form of his predictive architecture, for hierarchical planning under uncertainty; it is a proposal, untested here. Two other levers shorten the horizon instead of lengthening the step: TD-MPC (Hansen, Wang and Su, 2022) plans 5 steps and lets a learned value stand for the rest, and MBPO (Janner et al., 2019) branches short model rollouts from real states because long ones were worse.

The widget

Jump, plan and dock: how far a coarser step sees, and what it cannot steer
Top left: the maze (cyan: the ball; amber: ends of jumps; purple: the first plan; ×: the waypoint). Top right: median error of a chain of jumps against the steps it covers, with the 5 cm tolerance. Bottom: reach against stride; the dashed line is the 60 steps the flat planner needs. The planners call the simulator for their jumps; the error and reach readouts are the learned model's.
one-jump error
—
reach at 5 cm
—
model error at the plan's end
—
jumps with a wall hit
—
episode
—
in the goal disc
—
decisions, model calls
—
Show the core JS
L10.stepEv = function (w, s, ev) {
    if (st[0] < w.r) { st[0] = 2 * w.r - st[0]; st[2] = -w.e * st[2]; ev.nx++; }
...
L10.keyOf = function (nx, ny) { return Math.min(nx, 3) * 4 + Math.min(ny, 3); };
L10.Model.prototype.given = function (x, nx, ny, f) { return back(readout(this.Wg[L10.keyOf(nx, ny)] || this.Ws, f || this.phi(x), this.p, 4)); };
L10.Model.prototype.summ = function (x) { var f = this.phi(x), c = this.counts(x, f); return this.given(x, c[0], c[1], f); };
...
      a = L10.clipTo(v[2 * j], v[2 * j + 1], A); q = CY.step(WM, q, a);
      for (u = 1; u < S; u++) q = CY.step(WM, q, null);
      c += 0.2 * S * Math.hypot(q[0] - G.x, q[1] - G.y);

What to try. Told nothing, slide the stride from 1 to 64: the reach reads 6, 16, 24, 24, 16, 0 and 0 steps, and the red curve leaves the 5 cm line at the first jump from stride 32. Choose the forecast: 7, 14, 32, 64, 144, then the whole window. At stride 8 with a plan of 24 steps (3 jumps) episode 1 docks at step 33 on 4,500 simulator calls, and the learned model's error at the plan's end is 3.6 cm without a summary, inside the tolerance. At stride 64 with 64 steps (one jump) the ball enters the goal and stays 16 steps, never slower than 0.68 m/s: no dock. Stride 1 with 64 steps docks at step 54 on 1,036,800 calls. With the plan at 8 steps and the planner on skills it docks at step 36 on 86,400 model steps, at any stride.

What this lesson did not do
The maze planners call the simulator; the learned jump models know the walls and not the posts, and each post is another fast event needing its own count. One learner at one size, one tolerance, one draw of twelve; a wider learner moves every curve (§4). One summary, hit counts (a time since the last hit, or the mean of the skipped states, would be others), on a deterministic toy: with lesson 2's noisy observations the count itself would be uncertain. Macro-actions were a nudge then a coast. The subgoals were given and the levels are arithmetic. Twenty episodes a cell: 18 of 20 has a standard error of 0.07.

7 · Where the toy ends

The state was four numbers and a jump a map of four numbers. A frame of video 256 × 256 × 3 is 196,608 numbers, 49,152 times the state; in 16 × 16 patches it is 256 tokens, at 8 frames a second 2,048 a second, 122,880 a minute and 7.37 million an hour (the arithmetic of lesson 17). In lesson 3's picture the ball carried 19 % of the pixel variance and 81 % was a band and a brightness with nothing to do with it. A summary of skipped events needs events to count and a jump needs a state to jump from; pixels offer neither directly.

Common mistakes / failure modes

"a coarser step always sees farther"
Told nothing, the reach is 24 steps at stride 8 and 0 at 32: the error per jump rose 9.3 times (§2).
"a rare bounce only hurts its own jump"
At stride 8, 1.6 % of jumps carry 55 % of the squared error, and the others slip from 7.1 to 21.2 mm at stride 16 (§3).
"keep every s-th frame, it is the same thing"
That is the model with no summary: 16 steps at stride 16 against 144 (§4).
"plan at the coarsest stride the model allows"
At stride 64 the planner enters the goal in 19 of 20 episodes and docks in 3 (§5).

Checkpoint exercise

Try it
A jump model errs by 7 mm per jump and errors add; the tolerance is 5 cm. (a) How many jumps are within tolerance? (b) At stride 8, how many steps? (c) A task needs 1,000 steps and each level of stride 8 stays within tolerance: how many levels? Answer: (a) 50 mm / 7 mm = 7.1, so 7 jumps. (b) 7 × 8 = 56 steps. (c) One level covers 56 steps, the next, whose step is that jump, 7 × 56 = 392, the next 7 × 392 = 2,744: 3 levels.

Where this points next

With the summary a learned jump model sees 64 steps at stride 8, and a planner that jumps by 8 (calling the simulator) docks in 20 of 20 episodes on 900 calls a decision where the flat planner needs 18,000; skills keep the reaction. But the toy is four numbers and a noisy dot, a jump needs a state to jump from and events to count (§7), and a frame is 196,608 numbers, and 81 % of the variance in lesson 3's picture was nuisance. Each piece we built, state, dynamics, uncertainty and rollouts, has to survive that. What happens to a world model when the observation is pixels at internet scale?

Takeaway
A model's error is paid once per application, so predicting s steps at a time multiplies reach by s while the error per jump stays flat: J jumps within tolerance, reach sτ/es, and s times fewer applications to pay for at every decision. Unaided the floor rises because the map folds where a jump skips a bounce and the few folded jumps bend the fit for all the others; the reach peaks at 24 steps and falls to nothing. A count of the skipped events restores the floor (64 steps at 8, the whole window from 32): match each channel's rate to the decision's. The decision's own rate is the limit: a planner committed for 16 steps docks 12 of 20 at best, one committed for 8, 20. Skills give the reaction back at 13 times fewer model steps, if someone supplies the subgoals; learning them is open.

Interview prompts

Companion reads: Reinforcement Learning · Model and planning (search with a model), Lesson 17 · Two seats, one tuple (the token arithmetic of §7) and Lesson 25 · Horizon, memory, and the quadratic bill (what a long horizon costs).