all_lessons/ World Models/ 17 · two seats, one tuplelesson 17 / 31

Two seats, one tuple

Before any architecture, one observation reorganises everything: a world model and a robot policy are the forward and the inverse of the same object. That is why they compete for one resource, why that resource is the whole story, and why this series and the robot series that follows it are one arc.

Where we are
Lesson 16 ends on this hand-off: Everything in this series assumed that a trained world model exists. How is one trained: which loss, which schedule, which data, and what does it cost? That is the question the rest of this series answers. The first sixteen lessons answered what a world model is and what you do with one — belief, latent state, rollouts, causality, imagination, planning, and the exam that says whether a model is good — and left the model itself assumed. From here the series discharges the assumption: how you actually train one. We start by locating the training problem precisely, because the naming of the bottleneck decides every later choice.
lessons 1–16What a world model isThe forward map p(o′|o,a) as an object: belief, state, rollouts, planning, and the exam that grades it. lessons 17–31 · this partTraining a world modelWhich loss, which schedule, which conditioning, which mixture, which deadline. next seriesTraining a robot modelThe inverse map p(a|o,g), then the shared bill: every source of data priced, and two rankings that disagree.
Forced by 16A world model is useful, evaluable, and assumed to exist; none of that says how one is trained. This stepName the object being trained: the forward half of the tuple (o,a,o′), and the one resource it needs. Forces 18If we are training a simulator, "trained" needs a definition no pixel metric supplies.

1 · One stream, read two ways

Everything an embodied learner ever sees arrives as one interleaved stream: something observed, something done, something observed next.

o1, a1, o2, a2, o3, a3, …

Now notice there are exactly two useful ways to slice that stream, and they are the two seats in this whole field.

forward · the predictor
p(o′ | o, a)
Given where you are and what you do, what happens? This is a world model. It answers consequence given handle.
inverse · the actor
p(a | o, g)
Given where you are and what you want, what should you do? This is a policy. It answers handle given goal.

They are not two fields that happen to be adjacent. They are two conditionals of one joint distribution over the same triple. Write the joint and the point is unavoidable:

p(o, a, o′) — condition on (o,a) → a world model; condition on (o,o′) or (o,g) → a policy

So both seats need the same rows of the same table. And here is the consequence that organises everything from here on: the table has three columns, and they are wildly unequal in price.

2 · Pixels are free, actions are dear, contact is dearest

Count what it costs to acquire one row.

Put a number on the gap using one experiment. V-JEPA 2's action-conditioned model, V-JEPA 2-AC, was pretrained on that million-plus hours of ordinary video, then given an action-grounded stage trained on roughly 62 hours of DROID robot video carrying a 7-dimensional end-effector state. That is the ratio the whole field lives inside:

free video : action-grounded video ≈ 1,000,000 : 62 ≈ 16,000 : 1
Read the ratio correctly
It is not "robot data is small, sad." It is a price signal. Four orders of magnitude separate the column you can have and the column you need. Every technique in these two series — latent action models, inverse dynamics pseudo-labeling, teleoperation rigs, cross-embodiment pooling, simulation, training inside a learned model — is an attempt to arbitrage that gap. Once you see the trade, the techniques stop looking like a grab bag and start looking like a market.
The handle ratio: what each seat is allowed to consume
Slide the two corpora. Unlabeled video trains an unconditioned predictor; only the action-grounded hours can train p(o′|o,a) or a policy. Token counts use the arithmetic we derive in §3: 7.37 M tokens per hour at 256², patch 16, 8 fps.
Interactive comparison of free versus action-grounded corpora.
free video
—
action-grounded
—
ratio
—
grounded tokens
—
diagnosis
—

Press the V-JEPA 2-AC preset and then drag the free-video slider from 10³ up to 10⁶ hours. The top bar grows a thousandfold; the bottom two do not move at all, because no quantity of unlabelled video adds a single action-grounded hour. That is the shape of the whole problem: the axis you can scale is not the axis that binds.

3 · Why video is simultaneously the biggest and the least valuable input

Both things are true at once, and holding both is the beginning of competence here. Take the standard recipe: a 256×256 frame cut into 16×16 patches gives

(256 / 16)² = 16 × 16 = 256 tokens per frame

At 8 frames per second that is 2,048 tokens per second, so

1 minute ≈ 122,880 tokens   ·   1 hour ≈ 7.37 × 10⁶ tokens

Compare a minute of speech: about 150 words, call it 200 tokens. Video is roughly 600× denser per wall-clock minute than language. So if you ask "what does this training consume the most," the answer by sheer token count is video, and it is not close. One million hours of video is about 7.4 × 10¹² tokens — the same order as an entire modern text pretraining corpus.

Now run the same arithmetic on the grounded column. DROID's 350 hours of real robot interaction come to about 2.6 × 10⁹ tokens — roughly 5,700× smaller than a 15-trillion-token text corpus. The column that is trivially the largest is the column that cannot, by itself, teach a model what its own actions do.

The question sharpened
"What does the training eat most?" has two answers. By volume: video. By marginal value per dollar: action-grounded, contact-rich, on-policy data. The two rankings are almost exactly inverted, and that inversion is the field's central engineering problem. Lessons 16 to 24 of the robot series make the ledger explicit; training a world model and training a robot model are the two ways of spending against it.

4 · What this part will and will not do

Lessons 1 to 16 derive the objects: belief under partial observation, latent state, the learned filter and its KL term, calibrated branch distributions, the compounding-error recurrence, imagination, planning, temporal abstraction, video, places, objects, latent actions, bodies and contact. The lessons from here on never re-derive those. They cite them and get on with the training craft.

The transition objective and its KL termlesson 4
Calibrated stochastic futureslesson 5
The recurrence e(k+1)=J·e(k)+δlesson 6
Model exploitation, δ·√(2 ln K)lesson 8
Latent actions as an interfacelesson 14
Sim as a designed data-generating processsynthetic_vision/01–10
Closed-loop trajectory evaluationsynthetic_vision/11–12
What is left, and what this part isthe training itself

Concretely: which space you predict in and what a fixed bit budget buys you there; why the tokenizer, not the transformer, sets the ceiling on expressible physics; how to stop teacher forcing from hiding the only error that matters; how the action gets in and why models learn to ignore it; where action labels come from when they do not exist; why a mean is not a future; what a context window actually costs; the staging that keeps four objectives from fighting; mixture design against real availability; post-training; distillation into a real-time loop; and the evaluation that predicts downstream usefulness rather than flattering it.

5 · The route

01 what "trained" means three acceptance tests no pixel metric implies 19 the prediction space a fixed bit budget: pixels vs tokens vs latents 20 the tokenizer it caps the physics you can express, forever 21 teacher forcing → rollout the exam with the answer key open 22 action conditioning how the handle gets in, and why it gets ignored 23 inferring the actions IDM and latent actions: manufacturing the handle 24 spreads, not averages the mean of two futures is neither of them 25 horizon and memory the quadratic bill for remembering 26 the curriculum stage when conflict beats forgetting 27 the mixture the highest-leverage hyperparameter you have 28 post-training what it fixes, and the ceiling it cannot cross 29 distillation 50 steps per frame is not a control loop 30 training-side evaluation metrics that predict downstream value 31 capstone the whole bill, end to end
Takeaway
A world model and a policy are the forward and inverse conditionals of one triple (o,a,o′), so they draw on one supply. In that supply the observation column is nearly free, the action column is expensive, and the force column is scarcer still — the V-JEPA 2 recipe puts the free-to-grounded ratio near 16,000 : 1. Video is by far the largest consumer by token count (7.37 M tokens per hour, ~600× denser than speech) and by far the weakest per token for learning consequences. Every method ahead is an arbitrage between those two facts.

Where this points next

We now know what we are training and what it is made of. What we do not have is a definition of success. "Low validation loss" is not one: a model can minimise next-frame error beautifully and still be useless as a simulator, because the thing a planner needs from it is not resemblance. Lesson 18 writes down the acceptance tests — controllability, consistency, calibration — and shows that no pixel metric implies any of them.

Interview prompts

Sources for the figures in this lessonV-JEPA 2 and V-JEPA 2-AC scale (>1 M hours internet video; ~62 h of DROID with 7-D end-effector state): arXiv:2506.09985. DROID scale (76 k trajectories ≈ 350 h): arXiv:2403.12945. Token and density arithmetic is derived in §3 from the stated resolution, patch size and frame rate.