all_lessons/ World Models/ 31 · capstonelesson 31 / 31

Capstone: the whole bill

One specification, put through all thirteen constraints, producing numbers. The point of the exercise is not the numbers — it is that the ordering of the decisions is forced, and that two of the thirteen constraints turn out to bind before any of the expensive ones matter.

Where we are
The last lesson of the predictor's seat. Everything derived so far is now applied to a single concrete brief, in the order the dependencies require, with the arithmetic shown at each step.
Forced by 30An evaluation protocol that survives optimiser pressure. This stepAssemble the full specification and cost it end to end. Forces robot_model_training/01A trained simulator does not act. The inverse map is a different problem.

1 · The brief

Specification
Train a world model for a bimanual manipulation cell that assembles small electronic parts. Deployment: a receding-horizon planner at 20 Hz (50 ms tick), evaluating K = 8 candidate plans over a 1.5 s horizon. Success is dominated by connector insertions with about 2 mm clearance. Two 640×480 scene cameras plus a wrist camera; joint encoders and wrist force-torque available at 1 kHz. Budget: real money, not unlimited.

2 · Check the two constraints that can kill the project (before spending anything)

Before a single GPU hour, run the two arithmetic checks that no amount of later work can undo.

Check 1 — the representational ceiling (lesson 20). The scene camera spans roughly 600 mm across 640 px, so 0.94 mm per pixel; at patch 16 the blur floor is 15 mm against a 2 mm clearance. Fails by 7.5×. The wrist camera spans about 120 mm across 480 px, giving 0.25 mm per pixel and a patch-16 floor of 4 mm. Still fails, by 2×. At patch 8 on the wrist camera the floor is 2 mm — marginal, and marginal is not a plan.

verdictVision alone cannot resolve the decisive distinction at any affordable tokenization. The insertion clearance must come from the force-torque channel, not from pixels. Decision: coarse visual tokens (patch 16, wrist camera at patch 8) for the approach, plus an explicit low-dimensional head on proprioception and force for contact state. This is lesson 20 §3's option 3 and lesson 28 §3's option 3 arriving at the same answer — and it was free to discover.

Check 2 — the latency budget (lessons 18, 24, 29). Per tick the planner needs H steps × N samples × K plans. At a 100 ms world-model step, a 1.5 s horizon is 15 steps:

15 steps × 8 plans = 120 model evaluations in 50 ms ⟹ 0.42 ms per evaluation
verdictNo sampling-based head reaches 0.42 ms. So the architecture is forced to be hierarchical before any training happens: branching prediction at 2 Hz for plan selection, a one-pass deterministic tracker at 20 Hz. Batching the 8 plans (lesson 29 §3) removes K from the latency term, leaving 15 sequential steps in 500 ms at 2 Hz — 33 ms per step, which a 3-step distilled student meets at 11 ms per step.

Two arithmetic checks, and the architecture is largely determined. Both were available on day zero. This is the actual thesis of the capstone: the cheapest decisions are the most binding, and they are the ones people defer.

3 · The data plan

Lesson 27's mixture, against real availability, staged per lesson 26.

stagemixturedrawnepochs on scarcest
1 · tokenizer70% broad video, 30% in-domain8,000 h0.24 on in-domain
2 · short-horizon dynamics50% broad, 25% in-domain, 25% sim30,000 h0.75 on in-domain
3 · long-horizon30% broad, 30% sim, 40% in-domain12,000 h0.48 on in-domain
4 · action-conditioned post-training60% grounded real, 40% sim500 h0.86 on grounded
5 · distillationteacher rollouts only——

Stage 4's row is the one to read carefully. It asks for 300 hours of grounded interaction against roughly 350 available — 0.86 epochs, just inside one pass. Push the post-training budget to 1,000 hours and you are at 1.7 epochs and lesson 27's logarithmic discount begins. So the grounded corpus size caps the post-training stage, which caps controllability, which is the acceptance test most at risk. The correct response is not a mixture weight. It is to collect more, which is the subject of lessons 16 to 24 of the robot series.

The whole bill
Token count from the sampling plan, training FLOPs as 6·Nparams·Ntokens, then GPU-hours at 4×10¹⁴ FLOP/s and 40% utilisation, dollars at $2.50 per GPU-hour, wall-clock on 512 accelerators. Every step is arithmetic you can redo by hand.
Interactive end-to-end training cost.
tokens / frame
—
total tokens
—
training FLOPs
—
GPU-hours
—
cost
—
wall-clock
—

Now use it as a sensitivity tool, which is its real purpose. Move resolution from 256 to 512: tokens per frame quadruples and so does the entire bill. Move frame rate from 8 to 16: it doubles. Move parameters from 109.5 to 1010: it roughly triples. The sampling plan — resolution and frame rate — has more leverage on cost than the model size does, because it multiplies the token count that the parameter count then multiplies again. Yet resolution and frame rate are usually chosen casually, in an afternoon, by whoever set up the cameras.

4 · The complete decision sheet

Acceptance gate — controllability, consistency, calibration, latencylesson 18
Prediction space — latent rollout, decode for inspection, explicit contact headlesson 19
Tokenizer — patch 16 scene / patch 8 wrist; blur floor checked firstlesson 20
Schedule — diffusion forcing (rollout-grade J at ≈1× cost)lesson 21
Conditioning — adaptive normalisation, 10% action dropout, guidance re-swept per stagelesson 22
Action labels — IDM on the in-domain corpus; buy labels while dG/dL > 1lesson 23
Head — categorical over tokens for the fast loop; flow head at 2 Hz for branchinglesson 24
Memory — context at the value-per-FLOP optimum; entity retrieval for permanencelesson 25
Curriculum — five stages with replay; evaluate jointly at every boundarylesson 26
Mixture — per stage, per the table in §3; never selected on validation losslesson 27
Post-training — action-diverse curated set; stop at ~90% of headroomlesson 28
Latency — cache, batch, then consistency-distil to 3 steps; re-measure ECElesson 29
Evaluation — planner-selected-action error; pre-registered gate; K = 8, derivedlesson 30
Two checks that determined the architecture, run on day zero, for free03 and 01/07/12

5 · What this seat cannot do

We have a simulator that answers counterfactual questions accurately, obediently, honestly and fast. It has never moved anything. It has no goals, no notion of success, and — crucially — it does not change the distribution of data it is evaluated on. A world model is trained on a fixed corpus and tested on a fixed corpus. Whatever else is hard about it, that much is stationary.

A policy is not. The moment you deploy a policy it starts visiting states, and which states it visits depends on the policy, which depends on training, which means the policy authors its own test set. That single asymmetry generates a different discipline with a different central theorem, a different scarcest resource, and a different set of tricks. It is the actor's seat, and it is where we go next.

Takeaway
Run the two free checks first. The blur floor against the task clearance (15 mm scene / 4 mm wrist against 2 mm needed — vision loses, so contact state must come from force) and the per-tick evaluation budget (120 model evaluations in 50 ms is 0.42 ms each — impossible, so the architecture is forced hierarchical). Both were available on day zero and both determined the architecture. Then note where cost actually lives: the sampling plan (resolution, frame rate) multiplies the token count that parameters then multiply again, so it outweighs model size — and the scarce grounded corpus caps the post-training stage, which caps controllability. Which is why the next-largest lever is not in this part at all; it is the data.

Where this points next

One series follows. Training a Robot Model takes the actor's seat: the inverse map, the policy that authors its own test set, and why behaviour cloning fails as O(εT²). Its last nine lessons take the ledger: what all of this actually eats, ranked two ways that disagree, with the arithmetic that answers "which corpus is the real consumer" honestly.

A trained world model answers what would happen if the robot did something, and it has never chosen anything: it has no goal, and the data it is tested on do not change when it is used. A robot has to choose, about twenty times a second, from what it has just sensed, and the cheapest teacher for a chooser is a person who already is one. What does a robot learn by copying a person doing the task, and how would we know it had learned it?

Interview prompts