How the handle gets in
A model that produces gorgeous, stable, physically plausible video and completely ignores your controller is not a rare pathology. It is the default outcome, it is favoured by the objective, and getting the action to matter is a quantitative design problem rather than a matter of wiring it in somewhere.
1 · Price the action in bits
Picture the failure first, because it is unnervingly convincing. You hold a controller. You push forward; the video continues, smooth and physically plausible. You push left; the video continues, smooth and physically plausible, and identical. Nothing you do changes anything, and yet every frame looks right. To see why gradient descent would build this on purpose, we have to price what the action is worth to it.
Ask the only question that determines whether gradient descent will bother: how much likelihood does knowing the action buy? That is a mutual information, conditioned on what the model already knows:
I(ot+1 ; at | o≤t) = H(ot+1 | o≤t) − H(ot+1 | o≤t, at)Now estimate the two terms for real video. Backgrounds persist. Motion is smooth and largely continues. Lighting is stable. Empirically, the overwhelming majority of the next frame is predictable from the history alone. Call the predictable share r, and take r ≈ 0.8 as a conservative figure for third-person manipulation video.
Then the action is competing for the remaining 20% of the explainable variance — and it is competing against a much easier alternative hypothesis, "keep doing what was happening." From the optimiser's point of view the deal is:
And here is the part that makes it sticky rather than a phase: once the model is good at history-only prediction, the residual the action could explain has shrunk, so the gradient reaching the conditioning pathway shrinks with it. Ignoring the action is a self-reinforcing local optimum. You have to design against it.
2 · Where to inject, ranked by how hard it is to ignore
All the usual mechanisms work; they differ in how easy it is for the network to route around them.
| mechanism | how it enters | how easily ignored |
|---|---|---|
| channel concatenation | action tiled and stacked onto the input | very easily — one layer can zero the weights |
| prefix / token | action embedded as extra sequence tokens | easily — attention simply does not look |
| cross-attention | every block queries the action sequence | moderately — must be ignored at every block |
| adaptive normalisation | action predicts per-layer scale and shift | hard — it modulates the whole activation path |
The ranking has a simple logic: the harder it is to express "ignore this" in the parameters, the more likely the pathway survives training. Adaptive normalisation wins because there is no weight setting that cleanly removes it without also breaking the layer it modulates. Concatenation loses because "ignore" is a single zero matrix, and a zero matrix is easy to find.
3 · Action dropout, and the branch it creates
During training, replace the action with a learned null token some fraction p of the time. This does two quite different things, and it is worth separating them.
It regularises. The model cannot become brittle in the action's absence, so the unconditional pathway stays healthy.
It creates a second model for free. Now one set of weights implements both p(o′|o,a) and p(o′|o). Having both lets you amplify the difference between them at inference — classifier-free guidance:
ô = f(o, ∅) + s · [ f(o, a) − f(o, ∅) ]Read the bracket: it is exactly "the part of the prediction the action is responsible for." Scale it by s > 1 and the action's influence grows superlinearly. This is the standard lever for a model that technically listens but too quietly.
It is not free, though. The dropout rate itself costs you conditioned gradient — at rate p, a fraction of your batches teach nothing about the action:
effective conditioning capacity κeff = κ · (1 − p)So p has an interior optimum too. Zero dropout means no unconditional branch and no guidance available at all. High dropout means a great unconditional branch and a starved conditional one. In practice a modest rate — around 10% — is the usual landing spot, and the arithmetic above is why.
4 · Guidance has an optimum, and it is not "as high as possible"
Turning s up increases controllability and degrades fidelity, because extrapolating past the conditional prediction pushes the sample off the data manifold. Two mechanisms with different curvature means a product with an interior maximum:
controllability(s) ≈ min(1, β·s / (1 + c₁(s−1)²)) · fidelity(s) ≈ 1 / (1 + c₂(s−1)³)Neither functional form is sacred — the shapes are. Controllability grows roughly linearly and then saturates; artifacts grow superlinearly. The product therefore peaks at a finite s, and that peak is your operating point.
Set dropout to 0 and read the diagnosis: no unconditional branch exists, so there is nothing to guide with. Set redundancy to 97% and note that the optimum survives but the value at the optimum collapses — the lever still works, it is just moving almost nothing. That distinction, broken lever versus nothing to lift, is the one to make before choosing a fix.
5 · Fixing the data, not just the architecture
If redundancy is the binding constraint, the answer is upstream. Three moves, in order of leverage:
- Train on segments where the action matters. Most frames of most videos are actions doing nothing interesting — a gripper travelling through free space. Contact events, direction changes, and grasp transitions carry nearly all of I(o′;a|o). Sampling toward them raises the action's share of the gradient without touching the model. This is lesson 27's mixture question in miniature, and it is usually the biggest single win.
- Increase action diversity at collection time. If the demonstrator always approaches from the same angle, the action column is nearly constant and carries no information by construction — you cannot recover a signal that was never varied. Deliberate perturbation during collection is cheap and directly raises the mutual information. lesson 7 makes the causal version of this argument.
- Shorten the conditioning gap. Predicting 2 seconds ahead from one action leaves the action explaining little of a long, mostly-inertial interval. Predicting the next 100 ms makes it dominant. Then chunk the horizon rather than the influence.
Where this points next
Everything in this lesson assumed the action column exists. For the corpus that actually has the scale — a million hours of internet video — it does not. There is no keypress log for a YouTube cooking video, no joint command stream for someone folding laundry. Lesson 23 shows how to manufacture the missing column, why it works at all, and — with the arithmetic — exactly when to stop paying for labels and start collecting directly.
Interview prompts
- Write the quantity that decides whether a model will use its action input. The conditional mutual information I(o_{t+1}; a_t | o_{≤t}) — the likelihood the action buys beyond what history already explains. (§1)
- Why is ignoring the action self-reinforcing rather than a transient phase? As history-only prediction improves, the residual the action could explain shrinks, so the gradient reaching the conditioning pathway shrinks too. (§1)
- Rank four injection mechanisms and give the criterion behind the ranking. Adaptive normalisation, cross-attention, prefix tokens, channel concatenation — ordered by how hard it is to express "ignore this" in the parameters. (§2)
- What two distinct things does action dropout buy, and what does it cost? It regularises the unconditional pathway and creates an unconditional branch that makes guidance possible; it costs conditioned gradient, since κ_eff = κ(1−p). (§3)
- Why does classifier-free guidance have an interior optimum? Controllability grows roughly linearly then saturates while off-manifold artifacts grow superlinearly, so their product peaks at a finite scale. (§4)
- Distinguish "broken lever" from "nothing to lift," and give the test. Broken lever means the conditioning pathway carries no gradient — caught by a near-zero paired-intervention divergence; nothing to lift means history redundancy leaves the action almost no variance to explain, so the optimum exists but its value is tiny. (§2, §4)
- Give the highest-leverage data-side fix for low controllability, and why it works. Sample training segments toward contact events and direction changes, which carry nearly all of I(o′;a|o), raising the action's share of the gradient without touching the model. (§5)