Capstone: the whole bill
One specification, put through all thirteen constraints, producing numbers. The point of the exercise is not the numbers — it is that the ordering of the decisions is forced, and that two of the thirteen constraints turn out to bind before any of the expensive ones matter.
1 · The brief
2 · Check the two constraints that can kill the project (before spending anything)
Before a single GPU hour, run the two arithmetic checks that no amount of later work can undo.
Check 1 — the representational ceiling (lesson 20). The scene camera spans roughly 600 mm across 640 px, so 0.94 mm per pixel; at patch 16 the blur floor is 15 mm against a 2 mm clearance. Fails by 7.5×. The wrist camera spans about 120 mm across 480 px, giving 0.25 mm per pixel and a patch-16 floor of 4 mm. Still fails, by 2×. At patch 8 on the wrist camera the floor is 2 mm — marginal, and marginal is not a plan.
Check 2 — the latency budget (lessons 18, 24, 29). Per tick the planner needs H steps × N samples × K plans. At a 100 ms world-model step, a 1.5 s horizon is 15 steps:
15 steps × 8 plans = 120 model evaluations in 50 ms ⟹ 0.42 ms per evaluationTwo arithmetic checks, and the architecture is largely determined. Both were available on day zero. This is the actual thesis of the capstone: the cheapest decisions are the most binding, and they are the ones people defer.
3 · The data plan
Lesson 27's mixture, against real availability, staged per lesson 26.
| stage | mixture | drawn | epochs on scarcest |
|---|---|---|---|
| 1 · tokenizer | 70% broad video, 30% in-domain | 8,000 h | 0.24 on in-domain |
| 2 · short-horizon dynamics | 50% broad, 25% in-domain, 25% sim | 30,000 h | 0.75 on in-domain |
| 3 · long-horizon | 30% broad, 30% sim, 40% in-domain | 12,000 h | 0.48 on in-domain |
| 4 · action-conditioned post-training | 60% grounded real, 40% sim | 500 h | 0.86 on grounded |
| 5 · distillation | teacher rollouts only | — | — |
Stage 4's row is the one to read carefully. It asks for 300 hours of grounded interaction against roughly 350 available — 0.86 epochs, just inside one pass. Push the post-training budget to 1,000 hours and you are at 1.7 epochs and lesson 27's logarithmic discount begins. So the grounded corpus size caps the post-training stage, which caps controllability, which is the acceptance test most at risk. The correct response is not a mixture weight. It is to collect more, which is the subject of lessons 16 to 24 of the robot series.
Now use it as a sensitivity tool, which is its real purpose. Move resolution from 256 to 512: tokens per frame quadruples and so does the entire bill. Move frame rate from 8 to 16: it doubles. Move parameters from 109.5 to 1010: it roughly triples. The sampling plan — resolution and frame rate — has more leverage on cost than the model size does, because it multiplies the token count that the parameter count then multiplies again. Yet resolution and frame rate are usually chosen casually, in an afternoon, by whoever set up the cameras.
4 · The complete decision sheet
5 · What this seat cannot do
We have a simulator that answers counterfactual questions accurately, obediently, honestly and fast. It has never moved anything. It has no goals, no notion of success, and — crucially — it does not change the distribution of data it is evaluated on. A world model is trained on a fixed corpus and tested on a fixed corpus. Whatever else is hard about it, that much is stationary.
A policy is not. The moment you deploy a policy it starts visiting states, and which states it visits depends on the policy, which depends on training, which means the policy authors its own test set. That single asymmetry generates a different discipline with a different central theorem, a different scarcest resource, and a different set of tricks. It is the actor's seat, and it is where we go next.
Where this points next
One series follows. Training a Robot Model takes the actor's seat: the inverse map, the policy that authors its own test set, and why behaviour cloning fails as O(εT²). Its last nine lessons take the ledger: what all of this actually eats, ranked two ways that disagree, with the arithmetic that answers "which corpus is the real consumer" honestly.
A trained world model answers what would happen if the robot did something, and it has never chosen anything: it has no goal, and the data it is tested on do not change when it is used. A robot has to choose, about twenty times a second, from what it has just sensed, and the cheapest teacher for a chooser is a person who already is one. What does a robot learn by copying a person doing the task, and how would we know it had learned it?
Interview prompts
- Which two checks should be run before any training, and why those two? The blur floor against the task's clearance and the per-tick evaluation budget — both are pure arithmetic available on day zero, and both determine the architecture in ways no later work can undo. (§2)
- Given 640 px spanning 600 mm at patch 16, compute the floor and state the consequence. 0.94 mm per pixel × 16 = 15 mm against a 2 mm clearance, a 7.5× shortfall, so contact state must come from the force channel rather than pixels. (§2)
- Show why a 20 Hz planner with 8 plans over 1.5 s forces a hierarchy. 15 steps × 8 plans is 120 evaluations in 50 ms, i.e. 0.42 ms each, which no sampling head achieves — so branching prediction moves to 2 Hz with a cheap tracker at 20 Hz. (§2)
- Why does batching the candidate plans change the analysis? The K rollouts are independent, so batching moves K out of the latency term into a memory term, leaving only the sequential horizon steps to fit the tick. (§2)
- Which has more leverage on total cost: the sampling plan or the parameter count? The sampling plan — resolution and frame rate set the token count, which the parameter count then multiplies, so quadrupling resolution quadruples the whole bill. (§3)
- What caps controllability in this build, and what is the right response? The grounded corpus size caps the post-training stage at under one epoch; the fix is collecting more grounded data, not a mixture weight. (§3)
- State the structural asymmetry between training a world model and training a policy. A world model is trained and tested on fixed distributions; a policy determines which states it visits, so it authors its own test set — which generates a different discipline. (§5)