all_lessons/ World Models/ 23 · inferring the actionslesson 23 / 31

The actions do not exist — infer them

The corpus with the scale has no action column. The corpus with the action column has no scale. The resolution is to manufacture the missing column — and there is a principled reason it works, plus arithmetic that tells you exactly when to stop paying for labels.

Where we are
Lesson 22 made the action load-bearing, assuming it was recorded. Lesson 17 priced that assumption: roughly 16,000 : 1 separates freely available video hours from action-grounded ones. This lesson is the arbitrage. It is the highest-leverage move in the whole part, and its logic recurs in robot_model_training/09 for bodies.
Forced by 22Conditioning only works if an action column exists. This stepInfer actions: inverse dynamics for pseudo-labels, latent actions when no label exists at all. Forces 24Even with the action known, the future still branches — a mean is not a future.

1 · Why the action is recoverable in the first place

This should feel suspicious. If nobody recorded what the hand did, how can it be recovered from the pixels? The answer is a statement about what an action is.

An action is precisely the information needed to make the next observation predictable from the current one. Formally, the residual uncertainty the action removes:

at ≈ argmina H(ot+1 | ot, a)   —   the minimal extra bits that explain the transition

And that quantity is visible in the data. The transition ot → ot+1 is right there in the video. If the gripper moved 3 cm left between frames, the video contains the evidence that a leftward command was issued. The action is not hidden; it is encoded in its own consequences.

Which gives the key asymmetry the whole method rests on:

The asymmetry
Predicting ot+1 from (ot, at) is hard — it requires knowing physics, and the answer is high-dimensional and genuinely uncertain. Predicting at from (ot, ot+1) is easy — the consequence is already given, the answer is low-dimensional, and you may look at the future. Because the inverse problem is the easy direction, a small labeled set suffices to solve it, and its solution then labels everything else.

2 · Inverse dynamics: buy a few labels, spend them everywhere

The recipe has three steps and one crucial detail.

  1. Collect a small labeled set. Record observations and actions from instrumented actors — teleoperators, a game client logging keypresses, a simulator.
  2. Train an inverse dynamics model. Fit g: (ot−k, …, ot+k) → at. Note the window is two-sided: the IDM is allowed to see the future, because it is not a policy and will never be deployed. This is the crucial detail — it makes the task dramatically easier than the policy's job and is why so few labels suffice.
  3. Pseudo-label the ocean. Run g over the unlabeled corpus to synthesise the action column, then train the world model on the result.

The canonical demonstration is OpenAI's Video PreTraining. A contractor set on the order of two thousand hours of Minecraft play, with keypresses and mouse movement recorded, trained an IDM; the IDM then pseudo-labeled 70,000 hours of ordinary online Minecraft video, which became the pretraining corpus for a behaviour-cloning foundation model that could craft diamond tools. OpenAI notes the contractor data cost about $2,000 — the arbitrage is roughly 35× in hours, and vastly more in dollars.

3 · Latent actions: when no label exists at all

Sometimes there is no instrumented version of the activity. Nobody has a joint-command log for a person chopping onions. Then you cannot pseudo-label a known action space — so invent one.

Train an encoder that looks at (ot, ot+1) and emits a small discrete code zt, together with a decoder that predicts ot+1 from (ot, zt), and put a hard bottleneck on z:

max I(ot+1 ; zt | ot)   subject to   |z| ≤ log₂ K bits

The bottleneck is doing all the work. With only a few bits available, the code cannot smuggle the whole next frame — it can only carry the most transition-explanatory few bits, which is the definition of an action from §1. So the codes converge on the discrete controls of the domain: move, turn, grasp, release. Genie learned a codebook on this order of magnitude, and the codes correspond to interpretable controls without a single action label.

Then bind the invented vocabulary to real controls with a tiny labeled set — dozens of examples, since you are learning a mapping between two small discrete sets, not a perception system. lesson 14 treats latent actions as the interface of a world you can play in real time; here they are a data acquisition technique. Same mechanism, different job.

4 · Pseudo-labels are a noisy channel — price them properly

An IDM with 92% accuracy does not give you 92% of a labeled corpus. Treat the pseudo-label as transmission of the true action through a noisy channel and compute what actually arrives. For an action vocabulary of size |A| and accuracy α, the information that survives per label is

I = log₂|A| − H(ε),    H(ε) = −α log₂α − ε log₂(ε/(|A|−1)),    ε = 1 − α

and the useful fraction is η = I / log₂|A|. So U pseudo-labeled hours are worth U·η grounded hours. Two consequences that matter:

Now the decision rule. Total grounded hours as a function of labeled hours purchased:

G(L) = L + U · η(α(L))

Differentiate. Labeled hours are worth buying while dG/dL > 1 — that is, while one purchased hour yields more than one grounded hour. Past that crossing, you should stop improving the IDM and collect grounded data directly instead, because direct collection dominates. That crossing is a computable number, not a matter of taste.

The exchange rate, and where to stop buying labels
Violet is total grounded hours G(L); the dashed line is direct collection, where one labeled hour buys exactly one grounded hour. The red marker is the knee where dG/dL falls below 1 — beyond it, extra labels are worse than just collecting.
Interactive label exchange-rate calculator.
IDM accuracy at knee
—
channel efficiency η
—
grounded hours
—
amplification
—
stop buying at
—

Drag the unlabeled corpus down to 10³ hours and watch the knee move sharply left: with a small ocean to label, the IDM is barely worth building. Drag it up and the knee moves right — a bigger ocean justifies a better IDM. That relationship is the actual planning tool. How good your IDM should be is a function of how much unlabeled data you have.

5 · The three failure modes, and how each betrays itself

systematic bias
Confidently wrong, consistently
The IDM mislabels one situation the same way every time. The world model learns that wrong physics cleanly — with high confidence and no visible loss anomaly. Catch it by evaluating the IDM stratified by situation, never in aggregate.
the embodiment gap
Right label, wrong units
The IDM was trained on your robot and applied to human video. It reports joint commands a human never issued. Detect it by checking pseudo-label distribution shift against the labeled set. This gap is robot_model_training/08's entire subject.
unobservable actions
No consequence, no label
Grip force under a static hold changes nothing visible, so the IDM cannot recover it — and the very actions that are invisible are often the decisive ones in contact. Argues, again, for non-visual channels.
The circularity to avoid
Do not train the IDM on data the world model generated, then train the world model on the IDM's labels. The two will agree beautifully and both be wrong, with no signal anywhere to tell you. Keep a locked, human-labeled holdout whose actions were measured, and evaluate both models against it only. lesson 8 shows the same failure for a policy trained against a model, and robot_model_training/22 for a model trained on its own generations.
Takeaway
An action is the minimal information that makes the next observation predictable, and that information is encoded in its own consequences — so the inverse problem is the easy direction, and a two-sided window makes it easier still. Hence: a small labeled set trains an inverse dynamics model that pseudo-labels an ocean (VPT: ~2k contractor hours unlocked 70,000 hours, ≈35×), and where no action space exists at all, an information bottleneck invents one (latent actions). Price pseudo-labels as a noisy channel, η = 1 − H(ε)/log₂|A|, and buy labels only while dG/dL > 1 — how good your IDM should be depends on how much unlabeled data you have.

Where this points next

We now have observations, a representation, a schedule, a live action pathway, and an action column even where none was recorded. One thing remains broken in the objective itself. Knowing the action exactly does not make the future unique: the mug can still tip either way, and a model trained to minimise squared error will answer with the average of both — which is a state the world never occupies. Lesson 24 shows why that average is not merely imprecise but invalid, and prices the fix against the latency budget from lesson 18.

Interview prompts

SourcesVPT contractor-to-online scale and the 70,000 hours of IDM-labeled video: arXiv:2206.11795 and OpenAI's accompanying write-up (which states the contractor data cost about $2,000). Genie's small discrete latent-action codebook: DeepMind's Genie line of work. The channel-capacity treatment of pseudo-labels in §4 is derived here from standard information theory.