Actions without labels, worlds in real time
Lesson 13 gave the model a state made of things, and structure fixes what a model can represent, not what it can learn from. The video it would learn from shows what happened, not what anyone did. Here the push is read out of the frames: it is what the ball's passive law cannot explain between two frames. An inverse dynamics model reads it after a few labelled pairs; a latent action model, a code of nine symbols trained to explain the next frame, finds all nine buttons of a player with no labels, and a handful of labels names them. Then the lesson prices a button press in passes per frame. It cannot handle contact: a smooth model fitted to these clips is 134 times worse at the steps that touch a wall than at the others.
New idea: an action is the unexplained change between two frames, and a code narrow enough to carry only that learns it from video alone. Because the code carries the player's choice, what the model still has to generate is nearly one blob, and that is what a few passes per frame can afford.
Forces next: Actions can be inferred from video and the model can run in real time, for a world seen through a screen. A robot's world is touched, not only seen: contact is discontinuous, force is invisible to a camera, and the goal arrives in words. What changes when the agent has a body, and the model must predict what it will feel?
1 · The video, and two things a model fitted to it cannot do
The video is the plain Courtyard seen from above: 8 m by 5 m, x to the right and y up, steps of 0.1 s, friction taking 3.4 % of the velocity each step. A player pushes the ball. At every step they press one of nine buttons: neutral, half the time, or one of the eight compass directions, one sixteenth of the time each, which adds 0.6 m/s to the ball's velocity. Nobody recorded the presses, as in lesson 13's video of nudged balls. A frame is a reading of the ball's state, position to 1 cm and velocity to 0.06 m/s, independent from frame to frame; it stands for what lessons 2 to 4 showed how to obtain. The video is 3,600 pairs of consecutive frames in 136 clips of about 26 steps. A clip ends when the ball comes within 0.4 m of a wall, so until §6 no pair touches a wall (0 of them do), and the ball moves at about 1.0 m/s.
Fit the passive law first. With nothing pushing, friction gives v′ = d·v and x′ = x + g·v on each axis, with d = e−γΔt = 0.9656 (γ is the friction rate, Δt = 0.1 s) and g = (1 − d)/γ = 0.0983 s. Least squares on the unlabelled pairs finds d̂ = 0.9604 and ĝ = 0.0981 s: the pushes average out, so the law can be learned from video that contains them. (The 0.5 % shortfall in d̂ is the noise in the readings of v, which shrinks a regression slope.) What the law cannot explain is the next velocity itself: the unexplained part has a root mean square of 0.447 m/s over the two axes together. A model that has only the past is wrong by that much about the one quantity the next frame shows.
It cannot obey. Tell even the best such model, one that knows the whole distribution of what the player does, that you pressed east, and then north. It has nowhere to put the information, and its draws for the two buttons are as far apart as two draws for the same button. Lesson 11 defined controllability as the spread of draws between actions over the spread between two draws of one action; here it is 1.00. Its answer is the average player.
It cannot answer in time. An interactive world has to show the effect of a press in the next frame, which at 24 frames per second leaves 41.7 ms: lesson 13's second bill. These are the two halves of the baton. They look separate, and §5 computes why they are not: the code that answers the first is what lets one pass do the work of many.
2 · An action leaves a trace
In free flight a push a at the start of a step enters the passive law as if the velocity had been v + a: v′ = d(v + a), x′ = x + g(v + a). Solve for the push:
a = v′/d − v
Call the right side, computed with the fitted d̂, the unexplained push u: two numbers in m/s, the coordinates of the nudge plane that the widget draws. With readings it carries noise: v′/d has variance σv²/d² and v has σv², so each axis of u is off by σv√(1 + d²)/d = 0.0864 m/s (0.0863 in the video), which is 0.122 m/s over both axes: the floor for a push read from the velocities alone.
An inverse dynamics model (IDM) learns the same map from labelled examples. Here it is a ridge regression (least squares with a small penalty on the weights) from the 8 numbers of the two frames and a constant to the push. It also reads the positions, since x′ − x = g(v + a) carries the push a second time. Trained on n random labelled pairs and scored on fresh video, its error over both axes, averaged over 300 random choices of the pairs, is 0.75 m/s for n = 3, 0.27 for 9, 0.15 for 20, and 0.113 with every label. Nine numbers per output must be pinned down, and with three pairs the fit is worse than saying "no push" (0.424 m/s).
That last number matters. A model that sees only the frame before the push is a policy, a prediction of the player. This player presses at random, so the past says nothing and the best it can do is the spread of the pushes themselves: 0.424 m/s, right 50.0 % of the time as a classifier, by answering neutral. The IDM reads the frame after as well, which is why it is 3.8 times better (right 99.7 % of the time). It is non-causal, and that is what makes labelling easier than acting. Video PreTraining (Baker et al., 2022), whose bill lesson 13 counted, uses exactly this: an IDM trained on a small labelled set labels web video of Minecraft, and a policy is cloned from the result.
3 · Take the labels away: a latent action model
Give a model two frames and ask it to predict the second from the first and a code c, one of K symbols that an encoder chooses by looking at both frames; the decoder predicts frame t + 1 from frame t and c. Train both to get the prediction right. A code of K symbols costs log2K bits. What is it worth spending them on? The decoder already holds the passive law, so only the unexplained push is left to say, and when K is small the code that cuts the error most is the best K-symbol summary of u. The code learns the action because nothing else is left to learn. Genie (Bruce et al., 2024) is built this way: its latent action model, trained on video without labels with a VQ-VAE objective, has 8 latent actions and is discarded at inference, where the user supplies them.
Here the encoder is "the nearest of K centres in the nudge plane" and the decoder is "the passive law plus the code's centre as the push". Choosing the centres and the assignment to minimise the squared error of the decoded velocity is k-means, run from twelve seeded starts with the best kept; a single start can split one button over two codes and dissolve another (purity 93.0 % at K = 9, against 99.6 %). In the table, bits is log2K; decoder error is the error of the decoded next velocity over both axes; purity is the share of pairs that fall in a code whose majority button is theirs; the last two columns are the bits the code uses (its entropy) and how many of those are about the real button (mutual information).
| codes K | bits | decoder error, m/s | purity, % | bits used | about the button |
|---|---|---|---|---|---|
| 1 | 0.00 | 0.447 | 48.8 | 0.00 | 0.00 |
| 8 | 3.00 | 0.144 | 93.3 | 2.43 | 2.32 |
| 9 | 3.17 | 0.121 | 99.6 | 2.54 | 2.50 |
| 64 | 6.00 | 0.050 | 99.3 | 5.71 | 2.50 |
Read it from the top. With one code the decoder is wrong by 0.447 m/s: nothing is explained. At K = 9 it is wrong by 0.121 m/s, the floor of §2 (0.122). The code uses 2.54 bits, 2.50 of them about the button out of the 2.53 bits that the player's buttons contain, and every code is one button. K = 8, the number Genie used, merges two neighbours (purity 93.3 %): the codebook must be as large as the alphabet, and without labels it is found as the smallest K whose decoder error reaches the floor. Past nine the error goes below the floor, which no push can explain. At K = 64 the decoder is 59 % better than the floor because the code carries 5.71 bits, only 2.50 of them about the button; the other 3.21 bits are the particular noise of that pair of frames. A code that wide has stopped being an action: it would carry the frame itself, so the bottleneck has to be narrow.
What the code does not have is a name. Swap the numbers of any two codes and swap their centres with them, and every prediction of the decoder is unchanged: over all 3,600 pairs the largest difference is exactly 0. The 9! = 362,880 relabellings are equally good codebooks, so nothing in unlabelled video says which symbol is "east". The video can discover the alphabet. It cannot name it.
4 · Name the codes, then ask whether they steer
Naming. Label n pairs, and call each code by the button most of its labelled pairs show; a code with no labelled pair is read as neutral. With random labels the cost is a coupon collector's: every code must turn up once. Half the labels are neutral and carry nothing; the other half choose among the eight compass codes with equal chances, so meeting all eight takes the collector's 8·H8 such labels, where H8 = 1 + 1/2 + … + 1/8 = 2.72, and twice as many labels in all: 16 × 2.72 = 43.5 (the exact count agrees).
The unlabelled video says where to look. Label the pair nearest the centre of each code, one per code: 9 labels. Named that way every button is right: 99.7 % of pairs are read as the right button, 8 of 8 buttons steer where they should, and the push is read to 0.029 m/s. Nine random labels name only 3.6 of the 8 on average (72.2 %); over 300 random orders the share is 86.3 % at 20 labels, 96.0 % at 40 and 98.8 % at 60.
On the same random labels the IDM of §2 does better: 76.6 % at 9 labels and 96.9 % at 20, against 72.2 and 86.3. In this toy that is expected: the physics is linear and nine numbers pin the IDM down. Three things favour the codes: at zero labels they already work as buttons; they presuppose less (a symbol per labelled example, not a measured push); and a named code returns the exact vector of its button, its error of 0.029 m/s being only the pairs that fell in the wrong code, where the IDM's estimate carries the floor of 0.113. Real systems take this step with a small labelled set. Genie's imitation test maps its latent actions to real ones; LAPA (Ye et al., 2025) quantises the action between frames, pretrains a policy to predict the codes and finetunes on a little robot data to map them to the robot's own; Dreamer 4 (Hafner, Yan and Lillicrap, 2025) keeps more than 80 % of the accuracy of using every action from the small paired share that lesson 13 counted.
Do they steer? The exam of an interactive model is not purity but control. Lesson 11's controllability is the spread of draws between two different buttons over the spread between two draws of the same one. With the nine codes, draws for different buttons are 0.81 m/s apart on average and two draws for one button 0.15 m/s: controllability 5.4. The real buttons are 0.80 m/s apart, so the model separates them by as much as the world does. The same video with no code is at 1.00. A codebook that merges buttons steers less (4.4 at K = 8). The ratio itself needs no label.
5 · A button press in 41.7 ms
A press is a control only if its effect arrives in the next frame. At 24 frames per second a frame has 1000/24 = 41.7 ms, and a frame costs one network pass per denoising step (lesson 11). With a pass of t ms, ⌊41.7/t⌋ passes fit: 4 at 10 ms, which is 40 ms and 25 frames per second, and a fifth pass drops the rate to 20.
Why a frame needs more than one pass. One pass is a regression, and a regression returns the mean (lesson 5). A model with no code has to produce the player's choice itself, and that choice has nine modes in the nudge plane. We can measure what n passes buy with the best denoiser there could be, so that the only error is the number of steps. For a mixture of Gaussians the denoiser at noise level σ is known in closed form: the mean of the push given a noisy value x is Σc wc(x)[μc + sc²/(sc² + σ²)·(x − μc)], with weights wc ∝ πc N(x; μc, (sc² + σ²)I) (μc, sc and πc are the centre, width and frequency of code c). The sampler starts from noise of size 1.5 m/s and takes n Euler steps down the probability-flow ODE, dx/dσ = (x − D(x, σ))/σ, with σ1/7 evenly spaced from 1.5 down to 0.02 and then to 0: n passes, one evaluation of D each. The nine codes of §3 are the mixture; 2,000 draws per row. Distance is half the sum of absolute differences between how often the draws land on each code and how often the video presses it.
| passes n | codes reached | distance to the press rates |
|---|---|---|
| 1 | 1 of 9 | 0.51 |
| 3 | 9 of 9 | 0.26 |
| 4 | 9 of 9 | 0.189 |
| 8 | 9 of 9 | 0.077 |
| 16 | 9 of 9 | 0.048 |
One pass puts every draw within 0.09 m/s of the mean of the mixture, which lies 0.01 m/s from zero: on the neutral code, the player's most frequent press, whatever the player does. Three passes reach all nine codes; sixteen get the press rates to within 0.048, against 0.028 for drawing the mixture directly, which is sampling noise.
What the code changes. Here the two questions meet. Give the model the button: the distribution of the push is now one Gaussian of width 0.086 m/s, the denoiser is linear at every σ, and a single pass lands on the button's push to within 0.002 m/s. The player's choice no longer has to be generated; it is an input, and what is left for passes to spend on is the jitter around the mean and, in pictures, detail. Real systems in real time use 3 or 4 passes per frame (lesson 11), and GameNGen (Valevski et al., 2025) reaches about 50 frames per second with a one-step distilled model at a cost in quality. Distillation trains a student to reproduce the many-step result in fewer passes; it is described here, not run. With the button its target is one blob that the teacher already reaches in one pass; without it, nine modes whose one-pass answer is their mean.
Causal streaming. Make the model causal, so that frame t + 1 depends only on frames up to t and the current button. A model that denoises a clip of 16 frames together lets every frame see every other and can take a button only at a clip boundary, up to 16/24 = 0.67 s later (0.33 s on average); the causal model answers in the next frame. It can also keep the keys and values of past frames: a new frame of T tokens against a window of n frames then asks T·nT pairs of attention instead of recomputing (nT)², n times fewer. For the Courtyard's 40-token pictures and a window of 16 that is 25,600 pairs against 409,600, 16 times fewer. Diffusion Forcing (Chen et al., 2024) trains such a model on tokens at independent noise levels; Self Forcing (Huang et al., 2025) trains a few-step one on its own rollouts with cached keys and values, and streams in real time on a single GPU.
The widget
What to try. Start at the defaults: nine codes, 20 labelled pairs at random, 4 passes of 10 ms. With 20 labels 8 of the 9 codes have a name and 6 of the 8 compass buttons are named right, so 87.3 % of pairs are read as the right button, against 50.0 % at zero labels and 99.7 % at 60. Switch which examples to one per code: nine labels give 99.7 % and 8 of 8, where nine random labels in this order give 74.5 % and 4. Move codes: at 8 the purity is 93.3 %; at 9 it is 99.6 % with a decoder error of 0.121 m/s against the floor of 0.122; at 12 the error is 0.102, below the floor, and the code uses 3.36 bits for 2.53 bits of buttons. Move the passes: at 1 the amber rings sit on one code (1 of 9, distance 0.51), at 3 on all nine (0.26), at 16 the distance is 0.048. Move time per pass: at 10 ms four passes take 40 ms (25 fps) and five take 50 (20 fps), while at 8 ms five fit in 40. Finally choose the wall (§6): a smooth model fitted to clean clips is wrong by 0.020 m/s on free steps and 2.73 at the wall steps, 134 times more; fitted to a video with the bounces left in, 0.158 and 2.45, still 15.5 times more.
6 · How long a session lasts
A session feeds the model's own frames back as its next input: lesson 6's recursion. GameNGen reports that without noise added to the context frames in training, quality degrades fast after 20 to 30 steps; that noise, per-frame noise levels (Diffusion Forcing) and training on the model's own rollouts (Self Forcing) are three repairs of the same gap. In free flight the Courtyard shows how mild the problem is. The smooth model of this section is a linear map plus 48 random tanh features, with output weights fitted by ridge regression, from state and push to the change of state. Fitted to the clean clips and driven by a random player through 4 s sessions, its position error after 10, 20 and 40 steps is 0.040, 0.121 and 0.354 m (medians over 158 sessions that never touch a wall).
Now put the walls back. In the Courtyard 2.1 % of the steps of this player's video touch a wall, and 60.5 % of 4 s sessions do. Score the model step by step on fresh video with walls, the true push given so that every error is the model's. On steps that touch nothing the velocity error is 0.020 m/s. On steps that touch a wall it is 2.73 m/s, 134 times more (88 of its steps touch a wall). The natural repair is to leave the bounces in the training video. Fitted that way the model is wrong by 2.45 m/s at wall steps and 0.158 elsewhere: still 15.5 times worse at the walls, and now 7.8 times worse everywhere else than before, because a smooth function that cannot make a corner spreads the damage. It is not a matter of capacity or data: a 32 × 32 network fitted to four times the video is at 2.48 m/s at the walls and 0.121 elsewhere, 20 times worse.
A session is over at its first wall. Right after the wall step the position error is 0.20 m (median), and four steps later 1.14 m, which is 9.4 times what 20 steps of free flight cost. The model's ball goes through the wall, or leaves it the wrong way, and every later frame inherits the error.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
Actions can be inferred from video. A code of nine symbols found the player's nine buttons without a label (99.6 % pure), nine well-chosen labels named them, and the model steers, with a controllability of 5.4 against 1.00 for the same video without the code. It also runs in real time: given a button, one pass lands on its push, and four passes of 10 ms fit a 24 frames per second budget. All of this is a world seen through a screen, where a button changes the picture smoothly. The Courtyard's own walls break it. At a wall the velocity reverses within one step. The smooth model fitted to clean clips, wrong by 0.020 m/s in free flight, is wrong by 2.73 m/s at a wall step, 134 times more; the one fitted with the bounces in its video is still 15.5 times worse there than elsewhere. (Both are random-feature models scored by wall steps; lesson 15 trains a tanh network and scores it by distance to the wall, so its ratios differ.) A robot's world is touched, not only seen, and a camera cannot see the force. What changes when the agent has a body, and the model must predict what it will feel?
Interview prompts
- How can a model learn actions from video that has none? (§3 — give the decoder the passive law and force the rest through a code of K symbols; the code that cuts the error most quantises the unexplained change.)
- Why is inverse dynamics easier than a policy? (§2 — it sees the frame after the action; from the frame before, the push is as unpredictable as the player.)
- Why can nobody tell you which learned code is "jump"? (§3 — relabelling the codes leaves every prediction unchanged; only labelled examples name them.)
- How would you test that a world model obeys a button? (§4 — controllability: the spread of draws between two buttons over the spread between two draws of one; about 1 means the button is ignored.)
- Why does a diffusion world model need several passes, and what does conditioning on the action change? (§5 — one pass is the mean; without the action the player's choice makes the modes, with it one blob that one pass lands on.)
- Why does a smooth model fail at contact even when it has seen bounces? (§6 — a bounce is a corner in the state; a smooth function spreads the error, so wall steps stay 15 times worse and other steps get worse too.)
Companion reads: Lesson 23 · The actions do not exist: infer them (latent actions at corpus scale), Lesson 29 · Distilling into a real-time loop (few-step students) and Robot Model Training · 09 Watching: learning from video without actions (latent actions for robots).