The actions do not exist — infer them
The corpus with the scale has no action column. The corpus with the action column has no scale. The resolution is to manufacture the missing column — and there is a principled reason it works, plus arithmetic that tells you exactly when to stop paying for labels.
1 · Why the action is recoverable in the first place
This should feel suspicious. If nobody recorded what the hand did, how can it be recovered from the pixels? The answer is a statement about what an action is.
An action is precisely the information needed to make the next observation predictable from the current one. Formally, the residual uncertainty the action removes:
at ≈ argmina H(ot+1 | ot, a) — the minimal extra bits that explain the transitionAnd that quantity is visible in the data. The transition ot → ot+1 is right there in the video. If the gripper moved 3 cm left between frames, the video contains the evidence that a leftward command was issued. The action is not hidden; it is encoded in its own consequences.
Which gives the key asymmetry the whole method rests on:
2 · Inverse dynamics: buy a few labels, spend them everywhere
The recipe has three steps and one crucial detail.
- Collect a small labeled set. Record observations and actions from instrumented actors — teleoperators, a game client logging keypresses, a simulator.
- Train an inverse dynamics model. Fit g: (ot−k, …, ot+k) → at. Note the window is two-sided: the IDM is allowed to see the future, because it is not a policy and will never be deployed. This is the crucial detail — it makes the task dramatically easier than the policy's job and is why so few labels suffice.
- Pseudo-label the ocean. Run g over the unlabeled corpus to synthesise the action column, then train the world model on the result.
The canonical demonstration is OpenAI's Video PreTraining. A contractor set on the order of two thousand hours of Minecraft play, with keypresses and mouse movement recorded, trained an IDM; the IDM then pseudo-labeled 70,000 hours of ordinary online Minecraft video, which became the pretraining corpus for a behaviour-cloning foundation model that could craft diamond tools. OpenAI notes the contractor data cost about $2,000 — the arbitrage is roughly 35× in hours, and vastly more in dollars.
3 · Latent actions: when no label exists at all
Sometimes there is no instrumented version of the activity. Nobody has a joint-command log for a person chopping onions. Then you cannot pseudo-label a known action space — so invent one.
Train an encoder that looks at (ot, ot+1) and emits a small discrete code zt, together with a decoder that predicts ot+1 from (ot, zt), and put a hard bottleneck on z:
max I(ot+1 ; zt | ot) subject to |z| ≤ log₂ K bitsThe bottleneck is doing all the work. With only a few bits available, the code cannot smuggle the whole next frame — it can only carry the most transition-explanatory few bits, which is the definition of an action from §1. So the codes converge on the discrete controls of the domain: move, turn, grasp, release. Genie learned a codebook on this order of magnitude, and the codes correspond to interpretable controls without a single action label.
Then bind the invented vocabulary to real controls with a tiny labeled set — dozens of examples, since you are learning a mapping between two small discrete sets, not a perception system. lesson 14 treats latent actions as the interface of a world you can play in real time; here they are a data acquisition technique. Same mechanism, different job.
4 · Pseudo-labels are a noisy channel — price them properly
An IDM with 92% accuracy does not give you 92% of a labeled corpus. Treat the pseudo-label as transmission of the true action through a noisy channel and compute what actually arrives. For an action vocabulary of size |A| and accuracy α, the information that survives per label is
I = log₂|A| − H(ε), H(ε) = −α log₂α − ε log₂(ε/(|A|−1)), ε = 1 − αand the useful fraction is η = I / log₂|A|. So U pseudo-labeled hours are worth U·η grounded hours. Two consequences that matter:
- Accuracy enters non-linearly. Going from 80% to 90% buys much more than 90% to 95% costs you in labels — the entropy term is steep near the middle and flat near the top.
- Vocabulary size matters. A large continuous-ish action space is harder to hit exactly, but each correct label carries more bits. There is a real design choice in how coarsely to discretise, and it is not obvious in advance.
Now the decision rule. Total grounded hours as a function of labeled hours purchased:
G(L) = L + U · η(α(L))Differentiate. Labeled hours are worth buying while dG/dL > 1 — that is, while one purchased hour yields more than one grounded hour. Past that crossing, you should stop improving the IDM and collect grounded data directly instead, because direct collection dominates. That crossing is a computable number, not a matter of taste.
Drag the unlabeled corpus down to 10³ hours and watch the knee move sharply left: with a small ocean to label, the IDM is barely worth building. Drag it up and the knee moves right — a bigger ocean justifies a better IDM. That relationship is the actual planning tool. How good your IDM should be is a function of how much unlabeled data you have.
5 · The three failure modes, and how each betrays itself
Where this points next
We now have observations, a representation, a schedule, a live action pathway, and an action column even where none was recorded. One thing remains broken in the objective itself. Knowing the action exactly does not make the future unique: the mug can still tip either way, and a model trained to minimise squared error will answer with the average of both — which is a state the world never occupies. Lesson 24 shows why that average is not merely imprecise but invalid, and prices the fix against the latency budget from lesson 18.
Interview prompts
- Why is recovering the action from video possible at all? An action is the minimal extra information that makes o_{t+1} predictable from o_t, and that information is encoded in the observed consequence, which the video already contains. (§1)
- State the asymmetry that makes inverse dynamics cheap. Forward prediction is high-dimensional, physics-dependent and genuinely uncertain; inverse prediction is low-dimensional, has the consequence given, and may use a two-sided window including the future. (§1)
- Why is the IDM allowed to see future frames when a policy is not? It is a labeling tool, never deployed for control, so non-causal access is legitimate and makes its task far easier. (§2)
- How does a latent action model produce an action space without labels? An information bottleneck of a few bits between (o_t,o_{t+1}) and the reconstruction of o_{t+1} can only carry the most transition-explanatory bits, which is the definition of an action. (§3)
- Why isn't a 92%-accurate pseudo-label worth 92% of a real label? Value is the information surviving a noisy channel, η = 1 − H(ε)/log₂|A|, which is non-linear in accuracy and depends on vocabulary size. (§4)
- Give the stopping rule for buying action labels. Buy while dG/dL > 1 for G(L) = L + U·η(α(L)); past that knee, direct collection of grounded data dominates further IDM improvement. (§4)
- Which IDM failure mode is most dangerous and why? Systematic bias — the world model learns the wrong physics confidently and cleanly, with no loss anomaly; only situation-stratified IDM evaluation against a measured holdout reveals it. (§5)