all_lessons/ World Models/ 25 · horizon and memorylesson 25 / 31

Horizon, memory, and the quadratic bill

Object permanence is not a reasoning capability that emerges with scale. In a transformer world model it is a mechanical property of the context window: things that left the window longer ago than the window is long do not exist. And the window is priced quadratically.

Where we are
Lesson 18 separated two consistency failures — drift, which lesson 21's schedule addresses, and forgetting, which it cannot touch. Lesson 24 then added sampling steps that multiply with horizon. Now we price the remaining axis: how far back the model looks, what that buys in permanence, and what it costs in FLOPs.
Forced by 24Sampling cost multiplies with horizon, so budgets now bind. This stepTrade P(forget) against quadratic attention cost; find the context that maximises value per FLOP. Forces 26Four objectives now compete for one budget — they need a schedule.

1 · Forgetting has an exact mechanism, so it has an exact fix

Turn your head away from a chair for three seconds and turn back. If the chair is a different chair, people call it "the model lacks object permanence," which sounds like a cognitive deficiency. It is arithmetic. The model attends to C past frames. The chair left the frame k frames ago. If k > C, the chair is not in the input. There is nothing to remember it with.

So make the question quantitative: how long are the gaps? Model the interval between leaving and revisiting a region as roughly geometric with mean ḡ frames — a decent fit for the wandering attention of manipulation and navigation. Then

P(forget) = P(gap > C) ≈ exp( −C / ḡ )

Read the sensitivity: permanence improves exponentially in context length. That is the good news, and it is why modest context increases can feel transformative. Now the bad news.

2 · Context is quadratic, and tokens per frame is the multiplier

Attention over the context costs pair-operations in the square of the token count, and the token count is C frames times T tokens per frame:

cost ∝ (C · T)²

Both factors matter, and the second is the one people forget. At lesson 20's standard recipe T = 256, a 64-frame context is 16,384 tokens — and 268 million attention pairs per layer per step. Double the context and you pay 4×. But halve the patch size to sharpen contact and T becomes 1,024, so the same 64-frame context is 65,536 tokens and 4.3 billion pairs — a 16× increase from a decision that had nothing to do with memory.

The coupling that surprises people
Lesson 20's resolution choice and this lesson's memory choice multiply, then square. Tokenizer resolution is not a local decision about image quality; it is a global multiplier on your memory budget. If you need both fine contact detail and long permanence, uniform tokenization cannot deliver them — you need the non-uniform tokenization of lesson 20 §3, or retrieval.

Combine the two curves. Permanence rises exponentially and saturates; cost rises quadratically without bound. A ratio of a saturating benefit to an unbounded cost has an interior maximum, so there is a context length that maximises value per FLOP — and it is usually far smaller than the largest one you can fit.

The context sweet spot
Red is P(forget)=exp(−C/ḡ), dashed cyan is normalised attention cost (C·T)², green is value per FLOP with its maximum marked. Raise tokens-per-frame and watch the optimum slide left — sharper images buy shorter memory at fixed budget.
Interactive context-length budget.
tokens in context
—
attention pair-ops
—
P(forget) at optimum
—
optimal context
—
diagnosis
—

Push the revisit gap to 300 frames — a robot working across a whole bench, returning to a fixture a minute later — and the optimum pins to the right edge with the diagnosis "permanence needs retrieval, not more context." That is the honest reading. When the required memory span exceeds what a quadratic mechanism can afford, the answer is not a bigger window. It is a different mechanism.

3 · The three mechanisms, and what each is for

context · attention
Exact, recent, expensive
Every detail of the last C frames, quadratic. Right for dynamics — what is moving, in contact, about to collide. Wrong for anything a minute old.
recurrent state
Cheap, constant, lossy
A fixed-size summary carried forward. O(1) per step, unbounded span — but it must decide what to keep before knowing what will be asked. Compression without a query.
retrieval · external memory
Sparse, addressable, cheap
Write entities to a store, fetch the few relevant ones. Cost scales with what you retrieve, not with elapsed time. The only mechanism whose price is independent of span.

The decisive property is in that last cell. Attention costs the square of the span. A recurrent state costs nothing extra but cannot be queried about what it discarded. Retrieval alone breaks the link between how long ago something happened and how much it costs to recall — which is exactly why permanence and dynamics should be served by different mechanisms rather than one heroic context window.

lesson 12 and 13 derive what the retrieved units should be — places in a pose-indexed map and persistent object slots rather than raw frames — and that structure is what makes retrieval work here: a map grows with the area and an entity store with the number of things, while a frame store grows with the time covered.

4 · Horizon is a third, separate budget

Do not conflate context (how far back) with horizon (how far forward). They are independent, and they cost differently.

Context C — how far back the model attendscost ∝ (C·T)²
Horizon H — how many steps you roll forwardcost ∝ H (sequential)
Sampling N — steps per generated frame (lesson 24)cost ∝ N
Candidates K — plans the planner evaluatescost ∝ K
Total per control tick∝ (C·T)² · H · N · K

Four multiplicative factors, one 50 ms tick. This is the arithmetic that makes every later lesson necessary: staging (09) because you cannot train all of it at once, mixture design (10) because you cannot afford data for all of it, distillation (12) because N must come down, and receding-horizon control — imagine only far enough to choose the next action, then take a real observation — because H must come down too. lesson 6 and 09 make the case for the last one in full: re-observing restarts the error clock (06), and replanning from what was observed lets even a wrong model dock where its one-shot plan fails (09).

5 · How to measure permanence, since loss will not tell you

  1. Build a revisit test set. Clips where an object leaves the frame for a controlled number of frames k and returns. You need k as an explicit axis, not an average.
  2. Score identity, not pixels. Is it the same object — same instance, pose, and count? A pixel metric on a revisit frame is dominated by background and will look fine while the chair silently becomes a different chair.
  3. Plot accuracy against k and find the cliff. A context-limited model shows a sharp drop near k ≈ C. Seeing that cliff at the predicted location confirms the mechanism is the window and not something subtler — and tells you immediately whether more context or retrieval is the fix.
  4. Report object count stability too. Duplication — the model inventing a second chair on revisit — is the characteristic retrieval failure, distinct from forgetting, and it is invisible to an identity-only score.
Takeaway
Forgetting is mechanical: with mean revisit gap ḡ, P(forget) ≈ exp(−C/ḡ) — exponentially improving in context. But attention costs (C·T)², so tokens-per-frame from lesson 20 is a global multiplier on your memory budget, and a saturating benefit over an unbounded cost has an interior optimum well below the largest context you can fit. When the required span exceeds what quadratic attention affords, change mechanism: retrieval is the only one whose price is independent of elapsed time. And keep the four budgets distinct — the tick pays for (C·T)²·H·N·K.

Where this points next

We now have four things the model must be good at — reconstruct, predict one step, survive long rollouts, obey the action — and one budget. Training them simultaneously does not work, because their gradients disagree; training them in sequence does not work either, because the early ones decay while you train the later ones. Lesson 26 turns that into a decision rule with a crossing point you can compute: stage it when conflict beats forgetting.

Interview prompts