Horizon, memory, and the quadratic bill
Object permanence is not a reasoning capability that emerges with scale. In a transformer world model it is a mechanical property of the context window: things that left the window longer ago than the window is long do not exist. And the window is priced quadratically.
1 · Forgetting has an exact mechanism, so it has an exact fix
Turn your head away from a chair for three seconds and turn back. If the chair is a different chair, people call it "the model lacks object permanence," which sounds like a cognitive deficiency. It is arithmetic. The model attends to C past frames. The chair left the frame k frames ago. If k > C, the chair is not in the input. There is nothing to remember it with.
So make the question quantitative: how long are the gaps? Model the interval between leaving and revisiting a region as roughly geometric with mean ḡ frames — a decent fit for the wandering attention of manipulation and navigation. Then
P(forget) = P(gap > C) ≈ exp( −C / ḡ )Read the sensitivity: permanence improves exponentially in context length. That is the good news, and it is why modest context increases can feel transformative. Now the bad news.
2 · Context is quadratic, and tokens per frame is the multiplier
Attention over the context costs pair-operations in the square of the token count, and the token count is C frames times T tokens per frame:
cost ∝ (C · T)²Both factors matter, and the second is the one people forget. At lesson 20's standard recipe T = 256, a 64-frame context is 16,384 tokens — and 268 million attention pairs per layer per step. Double the context and you pay 4×. But halve the patch size to sharpen contact and T becomes 1,024, so the same 64-frame context is 65,536 tokens and 4.3 billion pairs — a 16× increase from a decision that had nothing to do with memory.
Combine the two curves. Permanence rises exponentially and saturates; cost rises quadratically without bound. A ratio of a saturating benefit to an unbounded cost has an interior maximum, so there is a context length that maximises value per FLOP — and it is usually far smaller than the largest one you can fit.
Push the revisit gap to 300 frames — a robot working across a whole bench, returning to a fixture a minute later — and the optimum pins to the right edge with the diagnosis "permanence needs retrieval, not more context." That is the honest reading. When the required memory span exceeds what a quadratic mechanism can afford, the answer is not a bigger window. It is a different mechanism.
3 · The three mechanisms, and what each is for
The decisive property is in that last cell. Attention costs the square of the span. A recurrent state costs nothing extra but cannot be queried about what it discarded. Retrieval alone breaks the link between how long ago something happened and how much it costs to recall — which is exactly why permanence and dynamics should be served by different mechanisms rather than one heroic context window.
lesson 12 and 13 derive what the retrieved units should be — places in a pose-indexed map and persistent object slots rather than raw frames — and that structure is what makes retrieval work here: a map grows with the area and an entity store with the number of things, while a frame store grows with the time covered.
4 · Horizon is a third, separate budget
Do not conflate context (how far back) with horizon (how far forward). They are independent, and they cost differently.
Four multiplicative factors, one 50 ms tick. This is the arithmetic that makes every later lesson necessary: staging (09) because you cannot train all of it at once, mixture design (10) because you cannot afford data for all of it, distillation (12) because N must come down, and receding-horizon control — imagine only far enough to choose the next action, then take a real observation — because H must come down too. lesson 6 and 09 make the case for the last one in full: re-observing restarts the error clock (06), and replanning from what was observed lets even a wrong model dock where its one-shot plan fails (09).
5 · How to measure permanence, since loss will not tell you
- Build a revisit test set. Clips where an object leaves the frame for a controlled number of frames k and returns. You need k as an explicit axis, not an average.
- Score identity, not pixels. Is it the same object — same instance, pose, and count? A pixel metric on a revisit frame is dominated by background and will look fine while the chair silently becomes a different chair.
- Plot accuracy against k and find the cliff. A context-limited model shows a sharp drop near k ≈ C. Seeing that cliff at the predicted location confirms the mechanism is the window and not something subtler — and tells you immediately whether more context or retrieval is the fix.
- Report object count stability too. Duplication — the model inventing a second chair on revisit — is the characteristic retrieval failure, distinct from forgetting, and it is invisible to an identity-only score.
Where this points next
We now have four things the model must be good at — reconstruct, predict one step, survive long rollouts, obey the action — and one budget. Training them simultaneously does not work, because their gradients disagree; training them in sequence does not work either, because the early ones decay while you train the later ones. Lesson 26 turns that into a decision rule with a crossing point you can compute: stage it when conflict beats forgetting.
Interview prompts
- Give the mechanical account of failed object permanence in a transformer world model. The model attends to C past frames; an object absent for more than C frames is simply not in the input, so with geometric revisit gaps P(forget) ≈ exp(−C/ḡ). (§1)
- Why does halving the tokenizer's patch size raise the memory bill 16×? Tokens per frame scale as (res/p)², so halving p quadruples T, and attention costs (C·T)² — squaring the quadrupling. (§2)
- Why does a value-per-FLOP optimum in context length exist at all? Permanence saturates exponentially while cost grows quadratically without bound, so their ratio peaks at a finite context. (§2)
- Which memory mechanism's cost is independent of elapsed time, and why does that matter? Retrieval — cost scales with what you fetch, not how long ago it happened — which is why permanence and dynamics should use different mechanisms. (§3)
- Distinguish context from horizon and give each one's cost scaling. Context is how far back, costing (C·T)² per step; horizon is how far forward, costing linearly in H but sequentially. They are independent budgets. (§4)
- Write the total per-tick cost and name the lesson that reduces each factor. (C·T)²·H·N·K — tokenization and retrieval for C·T (03, 08), receding-horizon control for H (lessons 6, 9), distillation for N (12), and the exploitation-aware search limit for K (13). (§4)
- How would you confirm that a permanence failure is a context-window failure? Plot revisit identity accuracy against the absence length k and look for a sharp cliff near k ≈ C; also report object count stability, since duplication is a distinct retrieval failure. (§5)