Two seats, one tuple
Before any architecture, one observation reorganises everything: a world model and a robot policy are the forward and the inverse of the same object. That is why they compete for one resource, why that resource is the whole story, and why this series and the robot series that follows it are one arc.
1 · One stream, read two ways
Everything an embodied learner ever sees arrives as one interleaved stream: something observed, something done, something observed next.
o1, a1, o2, a2, o3, a3, …Now notice there are exactly two useful ways to slice that stream, and they are the two seats in this whole field.
They are not two fields that happen to be adjacent. They are two conditionals of one joint distribution over the same triple. Write the joint and the point is unavoidable:
p(o, a, o′) — condition on (o,a) → a world model; condition on (o,o′) or (o,g) → a policySo both seats need the same rows of the same table. And here is the consequence that organises everything from here on: the table has three columns, and they are wildly unequal in price.
2 · Pixels are free, actions are dear, contact is dearest
Count what it costs to acquire one row.
- The o column is nearly free. Cameras have been pointed at the world for a century and the internet stores the result. Meta reports pretraining V-JEPA 2 on over 1,000,000 hours of internet video. Nobody had to be paid per hour to make that.
- The a column is not free. A video of a hand closing on a mug does not record the joint commands that closed it. To get a you must instrument the actor — a teleoperation rig, a simulator, a game client that logs keypresses. Somebody must be in the loop.
- The forces are dearest of all. A camera cannot see 4 newtons. Torque, slip, and compliance require sensors bolted to a body that is doing the task for real, and the body wears out.
Put a number on the gap using one experiment. V-JEPA 2's action-conditioned model, V-JEPA 2-AC, was pretrained on that million-plus hours of ordinary video, then given an action-grounded stage trained on roughly 62 hours of DROID robot video carrying a 7-dimensional end-effector state. That is the ratio the whole field lives inside:
free video : action-grounded video ≈ 1,000,000 : 62 ≈ 16,000 : 1Press the V-JEPA 2-AC preset and then drag the free-video slider from 10³ up to 10⁶ hours. The top bar grows a thousandfold; the bottom two do not move at all, because no quantity of unlabelled video adds a single action-grounded hour. That is the shape of the whole problem: the axis you can scale is not the axis that binds.
3 · Why video is simultaneously the biggest and the least valuable input
Both things are true at once, and holding both is the beginning of competence here. Take the standard recipe: a 256×256 frame cut into 16×16 patches gives
(256 / 16)² = 16 × 16 = 256 tokens per frameAt 8 frames per second that is 2,048 tokens per second, so
1 minute ≈ 122,880 tokens · 1 hour ≈ 7.37 × 10⁶ tokensCompare a minute of speech: about 150 words, call it 200 tokens. Video is roughly 600× denser per wall-clock minute than language. So if you ask "what does this training consume the most," the answer by sheer token count is video, and it is not close. One million hours of video is about 7.4 × 10¹² tokens — the same order as an entire modern text pretraining corpus.
Now run the same arithmetic on the grounded column. DROID's 350 hours of real robot interaction come to about 2.6 × 10⁹ tokens — roughly 5,700× smaller than a 15-trillion-token text corpus. The column that is trivially the largest is the column that cannot, by itself, teach a model what its own actions do.
4 · What this part will and will not do
Lessons 1 to 16 derive the objects: belief under partial observation, latent state, the learned filter and its KL term, calibrated branch distributions, the compounding-error recurrence, imagination, planning, temporal abstraction, video, places, objects, latent actions, bodies and contact. The lessons from here on never re-derive those. They cite them and get on with the training craft.
Concretely: which space you predict in and what a fixed bit budget buys you there; why the tokenizer, not the transformer, sets the ceiling on expressible physics; how to stop teacher forcing from hiding the only error that matters; how the action gets in and why models learn to ignore it; where action labels come from when they do not exist; why a mean is not a future; what a context window actually costs; the staging that keeps four objectives from fighting; mixture design against real availability; post-training; distillation into a real-time loop; and the evaluation that predicts downstream usefulness rather than flattering it.
5 · The route
Where this points next
We now know what we are training and what it is made of. What we do not have is a definition of success. "Low validation loss" is not one: a model can minimise next-frame error beautifully and still be useless as a simulator, because the thing a planner needs from it is not resemblance. Lesson 18 writes down the acceptance tests — controllability, consistency, calibration — and shows that no pixel metric implies any of them.
Interview prompts
- Why are world-model training and policy training resource-coupled? They are the forward and inverse conditionals of the same joint p(o,a,o′), so both require rows that contain the action column. (§1)
- Order the three columns of a transition tuple by acquisition cost, with a reason. Observations (passively recorded at internet scale), then actions (require instrumenting the actor), then forces and contact (require sensors on a body actually performing the task). (§2)
- Roughly how many tokens is an hour of video under a standard recipe, and how does that compare to speech? 256²/patch 16 = 256 tokens per frame; at 8 fps ≈ 7.37 M tokens per hour, about 600× denser per minute than the ~200 tokens/min of speech. (§3)
- "Robot data is the bottleneck" — restate this as a price rather than a complaint. Roughly four orders of magnitude separate freely available observation hours from action-grounded hours, so the field's methods are arbitrage across that spread. (§2)
- Which corpus does embodied training consume most, and is that the same as which matters most? By token volume, video, overwhelmingly; by marginal value per dollar, action-grounded contact-rich on-policy data. The rankings are inverted. (§3)
- Name three results this part deliberately does not re-derive, and where they live. The transition objective and its KL term (lesson 4), the compounding recurrence e(k+1)=J·e(k)+δ (lesson 6), and model exploitation δ·√(2 ln K) (lesson 8). (§4)