all_lessons/ World Models/ 22 · how the handle gets inlesson 22 / 31

How the handle gets in

A model that produces gorgeous, stable, physically plausible video and completely ignores your controller is not a rare pathology. It is the default outcome, it is favoured by the objective, and getting the action to matter is a quantitative design problem rather than a matter of wiring it in somewhere.

Where we are
Lesson 21 gave us rollouts that survive their own errors. Lesson 18 flagged the failure this lesson fixes: controllability is the acceptance test that fidelity metrics cannot see, and the loss actively prefers a model that drops the action. Now we compute how much it prefers that, and what to do about it.
Forced by 21Stable long rollouts — of a movie that may ignore the controller. This stepPrice the action's likelihood contribution; inject, drop out, and guide to make it load-bearing. Forces 23All of this assumes action labels exist. For almost all video, they do not.

1 · Price the action in bits

Picture the failure first, because it is unnervingly convincing. You hold a controller. You push forward; the video continues, smooth and physically plausible. You push left; the video continues, smooth and physically plausible, and identical. Nothing you do changes anything, and yet every frame looks right. To see why gradient descent would build this on purpose, we have to price what the action is worth to it.

Ask the only question that determines whether gradient descent will bother: how much likelihood does knowing the action buy? That is a mutual information, conditioned on what the model already knows:

I(ot+1 ; at | o≤t) = H(ot+1 | o≤t) − H(ot+1 | o≤t, at)

Now estimate the two terms for real video. Backgrounds persist. Motion is smooth and largely continues. Lighting is stable. Empirically, the overwhelming majority of the next frame is predictable from the history alone. Call the predictable share r, and take r ≈ 0.8 as a conservative figure for third-person manipulation video.

Then the action is competing for the remaining 20% of the explainable variance — and it is competing against a much easier alternative hypothesis, "keep doing what was happening." From the optimiser's point of view the deal is:

Likelihood available from history alone≈ 80%
Additional likelihood available from the action≈ 20%
Optimisation difficulty of using the action pathwayhigh (new, narrow, low-signal)
Rational early-training behaviourignore the action

And here is the part that makes it sticky rather than a phase: once the model is good at history-only prediction, the residual the action could explain has shrunk, so the gradient reaching the conditioning pathway shrinks with it. Ignoring the action is a self-reinforcing local optimum. You have to design against it.

2 · Where to inject, ranked by how hard it is to ignore

All the usual mechanisms work; they differ in how easy it is for the network to route around them.

mechanismhow it entershow easily ignored
channel concatenationaction tiled and stacked onto the inputvery easily — one layer can zero the weights
prefix / tokenaction embedded as extra sequence tokenseasily — attention simply does not look
cross-attentionevery block queries the action sequencemoderately — must be ignored at every block
adaptive normalisationaction predicts per-layer scale and shifthard — it modulates the whole activation path

The ranking has a simple logic: the harder it is to express "ignore this" in the parameters, the more likely the pathway survives training. Adaptive normalisation wins because there is no weight setting that cleanly removes it without also breaking the layer it modulates. Concatenation loses because "ignore" is a single zero matrix, and a zero matrix is easy to find.

A diagnostic worth ten hypotheses
Before tuning anything, run the paired intervention from lesson 18: same history, two different action sequences, measure the divergence. If it is near zero, your conditioning is not weak, it is absent, and no amount of loss reweighting will help — change the injection mechanism. Practitioners lose weeks tuning the strength of a pathway that carries no gradient at all.

3 · Action dropout, and the branch it creates

During training, replace the action with a learned null token some fraction p of the time. This does two quite different things, and it is worth separating them.

It regularises. The model cannot become brittle in the action's absence, so the unconditional pathway stays healthy.

It creates a second model for free. Now one set of weights implements both p(o′|o,a) and p(o′|o). Having both lets you amplify the difference between them at inference — classifier-free guidance:

ô = f(o, ∅) + s · [ f(o, a) − f(o, ∅) ]

Read the bracket: it is exactly "the part of the prediction the action is responsible for." Scale it by s > 1 and the action's influence grows superlinearly. This is the standard lever for a model that technically listens but too quietly.

It is not free, though. The dropout rate itself costs you conditioned gradient — at rate p, a fraction of your batches teach nothing about the action:

effective conditioning capacity   κeff = κ · (1 − p)

So p has an interior optimum too. Zero dropout means no unconditional branch and no guidance available at all. High dropout means a great unconditional branch and a starved conditional one. In practice a modest rate — around 10% — is the usual landing spot, and the arithmetic above is why.

4 · Guidance has an optimum, and it is not "as high as possible"

Turning s up increases controllability and degrades fidelity, because extrapolating past the conditional prediction pushes the sample off the data manifold. Two mechanisms with different curvature means a product with an interior maximum:

controllability(s) ≈ min(1, β·s / (1 + c₁(s−1)²))   ·   fidelity(s) ≈ 1 / (1 + c₂(s−1)³)

Neither functional form is sacred — the shapes are. Controllability grows roughly linearly and then saturates; artifacts grow superlinearly. The product therefore peaks at a finite s, and that peak is your operating point.

The guidance optimum, and what starves it
Violet is controllability, dashed cyan is fidelity, green is their product with its maximum marked. Then attack it from behind: raise history redundancy toward 1 and the action has nothing left to explain — no guidance scale can rescue a pathway with no signal.
Interactive guidance-scale trade-off.
variance the action explains
—
κ effective
—
optimal scale
—
value at optimum
—
diagnosis
—

Set dropout to 0 and read the diagnosis: no unconditional branch exists, so there is nothing to guide with. Set redundancy to 97% and note that the optimum survives but the value at the optimum collapses — the lever still works, it is just moving almost nothing. That distinction, broken lever versus nothing to lift, is the one to make before choosing a fix.

5 · Fixing the data, not just the architecture

If redundancy is the binding constraint, the answer is upstream. Three moves, in order of leverage:

  1. Train on segments where the action matters. Most frames of most videos are actions doing nothing interesting — a gripper travelling through free space. Contact events, direction changes, and grasp transitions carry nearly all of I(o′;a|o). Sampling toward them raises the action's share of the gradient without touching the model. This is lesson 27's mixture question in miniature, and it is usually the biggest single win.
  2. Increase action diversity at collection time. If the demonstrator always approaches from the same angle, the action column is nearly constant and carries no information by construction — you cannot recover a signal that was never varied. Deliberate perturbation during collection is cheap and directly raises the mutual information. lesson 7 makes the causal version of this argument.
  3. Shorten the conditioning gap. Predicting 2 seconds ahead from one action leaves the action explaining little of a long, mostly-inertial interval. Predicting the next 100 ms makes it dominant. Then chunk the horizon rather than the influence.
Takeaway
Whether the model uses the action is decided by I(ot+1;at|o≤t), and for ordinary video that quantity is small — roughly 80% of the next frame follows from history alone — so ignoring the handle is a self-reinforcing local optimum, not a phase. Counter it on three fronts: inject where "ignore" is hard to express (adaptive normalisation over cross-attention over prefix over concatenation); use ~10% action dropout, since κeff=κ(1−p) makes both zero and large dropout bad; and guide at the finite s that maximises controllability × fidelity. Always run the paired-intervention test first, to tell a broken lever from nothing to lift.

Where this points next

Everything in this lesson assumed the action column exists. For the corpus that actually has the scale — a million hours of internet video — it does not. There is no keypress log for a YouTube cooking video, no joint command stream for someone folding laundry. Lesson 23 shows how to manufacture the missing column, why it works at all, and — with the arithmetic — exactly when to stop paying for labels and start collecting directly.

Interview prompts