all_lessons/World Models/14 · Latent actions, real timelesson 14 / 31

Actions without labels, worlds in real time

Lesson 13 gave the model a state made of things, and structure fixes what a model can represent, not what it can learn from. The video it would learn from shows what happened, not what anyone did. Here the push is read out of the frames: it is what the ball's passive law cannot explain between two frames. An inverse dynamics model reads it after a few labelled pairs; a latent action model, a code of nine symbols trained to explain the next frame, finds all nine buttons of a player with no labels, and a handful of labels names them. Then the lesson prices a button press in passes per frame. It cannot handle contact: a smooth model fitted to these clips is 134 times worse at the steps that touch a wall than at the others.

The thesis, here
An action is the part of what changes between two frames that the past does not predict. A model can learn it from video alone if it must explain the next frame through a code too narrow to carry anything else, and the code it finds is the player's own alphabet, up to the names of the buttons. The same code makes real time affordable: given the button, what is left to generate is one blob that a single pass lands on, while without it the passes go to the player's choice and the model still cannot be steered.
Linear position
Forced by: Structure inside the model fixes what it can represent, not what it can learn from. The video on the internet shows what happened and not what anyone did to make it happen, and an interactive world must answer a button press now, not after a minute of sampling. Where do actions come from when the data has none, and how fast can a model run?
New idea: an action is the unexplained change between two frames, and a code narrow enough to carry only that learns it from video alone. Because the code carries the player's choice, what the model still has to generate is nearly one blob, and that is what a few passes per frame can afford.
Forces next: Actions can be inferred from video and the model can run in real time, for a world seen through a screen. A robot's world is touched, not only seen: contact is discontinuous, force is invisible to a camera, and the goal arrives in words. What changes when the agent has a body, and the model must predict what it will feel?
The plan
Six moves. (1) Fit a model to unlabelled video and see what it cannot do: obey a button, and answer within a frame time. (2) Find where an action hides in two frames and read it with an inverse dynamics model. (3) Take the labels away: a code too narrow to carry anything but the action. (4) Name the codes with a few labels and test whether they steer. (5) Price a button press in passes per frame. (6) Run a long session and find where it breaks.

1 · The video, and two things a model fitted to it cannot do

The video is the plain Courtyard seen from above: 8 m by 5 m, x to the right and y up, steps of 0.1 s, friction taking 3.4 % of the velocity each step. A player pushes the ball. At every step they press one of nine buttons: neutral, half the time, or one of the eight compass directions, one sixteenth of the time each, which adds 0.6 m/s to the ball's velocity. Nobody recorded the presses, as in lesson 13's video of nudged balls. A frame is a reading of the ball's state, position to 1 cm and velocity to 0.06 m/s, independent from frame to frame; it stands for what lessons 2 to 4 showed how to obtain. The video is 3,600 pairs of consecutive frames in 136 clips of about 26 steps. A clip ends when the ball comes within 0.4 m of a wall, so until §6 no pair touches a wall (0 of them do), and the ball moves at about 1.0 m/s.

Fit the passive law first. With nothing pushing, friction gives v′ = d·v and x′ = x + g·v on each axis, with d = e−γΔt = 0.9656 (γ is the friction rate, Δt = 0.1 s) and g = (1 − d)/γ = 0.0983 s. Least squares on the unlabelled pairs finds d̂ = 0.9604 and ĝ = 0.0981 s: the pushes average out, so the law can be learned from video that contains them. (The 0.5 % shortfall in d̂ is the noise in the readings of v, which shrinks a regression slope.) What the law cannot explain is the next velocity itself: the unexplained part has a root mean square of 0.447 m/s over the two axes together. A model that has only the past is wrong by that much about the one quantity the next frame shows.

It cannot obey. Tell even the best such model, one that knows the whole distribution of what the player does, that you pressed east, and then north. It has nowhere to put the information, and its draws for the two buttons are as far apart as two draws for the same button. Lesson 11 defined controllability as the spread of draws between actions over the spread between two draws of one action; here it is 1.00. Its answer is the average player.

It cannot answer in time. An interactive world has to show the effect of a press in the next frame, which at 24 frames per second leaves 41.7 ms: lesson 13's second bill. These are the two halves of the baton. They look separate, and §5 computes why they are not: the code that answers the first is what lets one pass do the work of many.

2 · An action leaves a trace

In free flight a push a at the start of a step enters the passive law as if the velocity had been v + a: v′ = d(v + a), x′ = x + g(v + a). Solve for the push:

a = v′/d − v

Call the right side, computed with the fitted d̂, the unexplained push u: two numbers in m/s, the coordinates of the nudge plane that the widget draws. With readings it carries noise: v′/d has variance σv²/d² and v has σv², so each axis of u is off by σv√(1 + d²)/d = 0.0864 m/s (0.0863 in the video), which is 0.122 m/s over both axes: the floor for a push read from the velocities alone.

An inverse dynamics model (IDM) learns the same map from labelled examples. Here it is a ridge regression (least squares with a small penalty on the weights) from the 8 numbers of the two frames and a constant to the push. It also reads the positions, since x′ − x = g(v + a) carries the push a second time. Trained on n random labelled pairs and scored on fresh video, its error over both axes, averaged over 300 random choices of the pairs, is 0.75 m/s for n = 3, 0.27 for 9, 0.15 for 20, and 0.113 with every label. Nine numbers per output must be pinned down, and with three pairs the fit is worse than saying "no push" (0.424 m/s).

That last number matters. A model that sees only the frame before the push is a policy, a prediction of the player. This player presses at random, so the past says nothing and the best it can do is the spread of the pushes themselves: 0.424 m/s, right 50.0 % of the time as a classifier, by answering neutral. The IDM reads the frame after as well, which is why it is 3.8 times better (right 99.7 % of the time). It is non-causal, and that is what makes labelling easier than acting. Video PreTraining (Baker et al., 2022), whose bill lesson 13 counted, uses exactly this: an IDM trained on a small labelled set labels web video of Minecraft, and a policy is cloned from the result.

Road not taken · label everything
Label every hour with an IDM and be done. The IDM presupposes the answer to the question: which buttons exist, and how a push is written down. It speaks the schema of the hours labelled for it; for video whose controller was never recorded, or of people, nobody holds that schema, and labelling by hand is the bill lesson 13 counted. The IDM needs labels in the action space. The next move needs none.

3 · Take the labels away: a latent action model

Give a model two frames and ask it to predict the second from the first and a code c, one of K symbols that an encoder chooses by looking at both frames; the decoder predicts frame t + 1 from frame t and c. Train both to get the prediction right. A code of K symbols costs log2K bits. What is it worth spending them on? The decoder already holds the passive law, so only the unexplained push is left to say, and when K is small the code that cuts the error most is the best K-symbol summary of u. The code learns the action because nothing else is left to learn. Genie (Bruce et al., 2024) is built this way: its latent action model, trained on video without labels with a VQ-VAE objective, has 8 latent actions and is discarded at inference, where the user supplies them.

Here the encoder is "the nearest of K centres in the nudge plane" and the decoder is "the passive law plus the code's centre as the push". Choosing the centres and the assignment to minimise the squared error of the decoded velocity is k-means, run from twelve seeded starts with the best kept; a single start can split one button over two codes and dissolve another (purity 93.0 % at K = 9, against 99.6 %). In the table, bits is log2K; decoder error is the error of the decoded next velocity over both axes; purity is the share of pairs that fall in a code whose majority button is theirs; the last two columns are the bits the code uses (its entropy) and how many of those are about the real button (mutual information).

codes Kbitsdecoder error, m/spurity, %bits usedabout the button
10.000.44748.80.000.00
83.000.14493.32.432.32
93.170.12199.62.542.50
646.000.05099.35.712.50

Read it from the top. With one code the decoder is wrong by 0.447 m/s: nothing is explained. At K = 9 it is wrong by 0.121 m/s, the floor of §2 (0.122). The code uses 2.54 bits, 2.50 of them about the button out of the 2.53 bits that the player's buttons contain, and every code is one button. K = 8, the number Genie used, merges two neighbours (purity 93.3 %): the codebook must be as large as the alphabet, and without labels it is found as the smallest K whose decoder error reaches the floor. Past nine the error goes below the floor, which no push can explain. At K = 64 the decoder is 59 % better than the floor because the code carries 5.71 bits, only 2.50 of them about the button; the other 3.21 bits are the particular noise of that pair of frames. A code that wide has stopped being an action: it would carry the frame itself, so the bottleneck has to be narrow.

What the code does not have is a name. Swap the numbers of any two codes and swap their centres with them, and every prediction of the decoder is unchanged: over all 3,600 pairs the largest difference is exactly 0. The 9! = 362,880 relabellings are equally good codebooks, so nothing in unlabelled video says which symbol is "east". The video can discover the alphabet. It cannot name it.

4 · Name the codes, then ask whether they steer

Naming. Label n pairs, and call each code by the button most of its labelled pairs show; a code with no labelled pair is read as neutral. With random labels the cost is a coupon collector's: every code must turn up once. Half the labels are neutral and carry nothing; the other half choose among the eight compass codes with equal chances, so meeting all eight takes the collector's 8·H8 such labels, where H8 = 1 + 1/2 + … + 1/8 = 2.72, and twice as many labels in all: 16 × 2.72 = 43.5 (the exact count agrees).

The unlabelled video says where to look. Label the pair nearest the centre of each code, one per code: 9 labels. Named that way every button is right: 99.7 % of pairs are read as the right button, 8 of 8 buttons steer where they should, and the push is read to 0.029 m/s. Nine random labels name only 3.6 of the 8 on average (72.2 %); over 300 random orders the share is 86.3 % at 20 labels, 96.0 % at 40 and 98.8 % at 60.

On the same random labels the IDM of §2 does better: 76.6 % at 9 labels and 96.9 % at 20, against 72.2 and 86.3. In this toy that is expected: the physics is linear and nine numbers pin the IDM down. Three things favour the codes: at zero labels they already work as buttons; they presuppose less (a symbol per labelled example, not a measured push); and a named code returns the exact vector of its button, its error of 0.029 m/s being only the pairs that fell in the wrong code, where the IDM's estimate carries the floor of 0.113. Real systems take this step with a small labelled set. Genie's imitation test maps its latent actions to real ones; LAPA (Ye et al., 2025) quantises the action between frames, pretrains a policy to predict the codes and finetunes on a little robot data to map them to the robot's own; Dreamer 4 (Hafner, Yan and Lillicrap, 2025) keeps more than 80 % of the accuracy of using every action from the small paired share that lesson 13 counted.

Do they steer? The exam of an interactive model is not purity but control. Lesson 11's controllability is the spread of draws between two different buttons over the spread between two draws of the same one. With the nine codes, draws for different buttons are 0.81 m/s apart on average and two draws for one button 0.15 m/s: controllability 5.4. The real buttons are 0.80 m/s apart, so the model separates them by as much as the world does. The same video with no code is at 1.00. A codebook that merges buttons steers less (4.4 at K = 8). The ratio itself needs no label.

5 · A button press in 41.7 ms

A press is a control only if its effect arrives in the next frame. At 24 frames per second a frame has 1000/24 = 41.7 ms, and a frame costs one network pass per denoising step (lesson 11). With a pass of t ms, ⌊41.7/t⌋ passes fit: 4 at 10 ms, which is 40 ms and 25 frames per second, and a fifth pass drops the rate to 20.

Why a frame needs more than one pass. One pass is a regression, and a regression returns the mean (lesson 5). A model with no code has to produce the player's choice itself, and that choice has nine modes in the nudge plane. We can measure what n passes buy with the best denoiser there could be, so that the only error is the number of steps. For a mixture of Gaussians the denoiser at noise level σ is known in closed form: the mean of the push given a noisy value x is Σc wc(x)[μc + sc²/(sc² + σ²)·(x − μc)], with weights wc ∝ πc N(x; μc, (sc² + σ²)I) (μc, sc and πc are the centre, width and frequency of code c). The sampler starts from noise of size 1.5 m/s and takes n Euler steps down the probability-flow ODE, dx/dσ = (x − D(x, σ))/σ, with σ1/7 evenly spaced from 1.5 down to 0.02 and then to 0: n passes, one evaluation of D each. The nine codes of §3 are the mixture; 2,000 draws per row. Distance is half the sum of absolute differences between how often the draws land on each code and how often the video presses it.

passes ncodes reacheddistance to the press rates
11 of 90.51
39 of 90.26
49 of 90.189
89 of 90.077
169 of 90.048

One pass puts every draw within 0.09 m/s of the mean of the mixture, which lies 0.01 m/s from zero: on the neutral code, the player's most frequent press, whatever the player does. Three passes reach all nine codes; sixteen get the press rates to within 0.048, against 0.028 for drawing the mixture directly, which is sampling noise.

What the code changes. Here the two questions meet. Give the model the button: the distribution of the push is now one Gaussian of width 0.086 m/s, the denoiser is linear at every σ, and a single pass lands on the button's push to within 0.002 m/s. The player's choice no longer has to be generated; it is an input, and what is left for passes to spend on is the jitter around the mean and, in pictures, detail. Real systems in real time use 3 or 4 passes per frame (lesson 11), and GameNGen (Valevski et al., 2025) reaches about 50 frames per second with a one-step distilled model at a cost in quality. Distillation trains a student to reproduce the many-step result in fewer passes; it is described here, not run. With the button its target is one blob that the teacher already reaches in one pass; without it, nine modes whose one-pass answer is their mean.

Road not taken · buy control with passes
If the model must generate the player's choice, why not spend enough passes to get it right? Because the choice it generates is its own: at 16 passes the draws reproduce the press rates to within 0.048, yet the controllability of a model that gets no button is 1.00. Passes buy fidelity to the data; only an input buys a handle on it.

Causal streaming. Make the model causal, so that frame t + 1 depends only on frames up to t and the current button. A model that denoises a clip of 16 frames together lets every frame see every other and can take a button only at a clip boundary, up to 16/24 = 0.67 s later (0.33 s on average); the causal model answers in the next frame. It can also keep the keys and values of past frames: a new frame of T tokens against a window of n frames then asks T·nT pairs of attention instead of recomputing (nT)², n times fewer. For the Courtyard's 40-token pictures and a window of 16 that is 25,600 pairs against 409,600, 16 times fewer. Diffusion Forcing (Chen et al., 2024) trains such a model on tokens at independent noise levels; Self Forcing (Huang et al., 2025) trains a few-step one on its own rollouts with cached keys and values, and streams in real time on a single GPU.

The widget

Name the codes, steer, and spend the passes
The nudge plane: each dot is the push the passive law cannot explain in one frame pair, coloured by its code; squares are the nine real buttons, circles the codes. The curves: pairs read as the right button against the labelled examples. The sampler panel: the draws of the passes when no button is given. The last panel: the passes of a frame against the frame time, or, in the wall view, one session of the smooth model of §6 (its two tiles read clean clips → clips with the bounces).
purity · bits in the code
—
decoder error (floor)
—
pairs read as the right button
—
error of the named push
—
buttons that steer right
—
inverse dynamics, same labels
—
controllability, with codes
—
controllability, no codes
—
codes the draws reach
—
distance to the press rates
—
time per frame
—
passes that fit 24 fps
—
smooth model, free steps
—
smooth model, wall steps
—
Show the core JS
// the push the passive law leaves unexplained
L.unexplained = function (rows, pas) { return rows.map(function (r) { return [r.z2[2] / pas.d - r.z[2], r.z2[3] / pas.d - r.z[3]]; }); };
// k-means with restarts; the encoder is "nearest code"
  for (sd = 1; sd <= (restarts || 12); sd++) { var km = CY.kmeans(U, K, { seed: sd, iters: 10 }); if (!best || km.inertia < best.inertia) best = km; }
L.encode = function (lam, u) { return CY.kmeans.nearest(lam.centers, u); };
// naming: count the labelled buttons in each code (the name is the majority)
  for (i = 0; i < n; i++) cnt[lam.assign[order[i]]][rows[order[i]].k + 1]++;
// n Euler passes, exact mixture denoiser
...
    var w = Math.exp(lw[k] - mx), v2 = mix.sd[k] * mix.sd[k], sh = v2 / (v2 + sg * sg); Z += w;
    o0 += w * (mix.mu[k][0] + sh * (x[0] - mix.mu[k][0])); o1 += w * (mix.mu[k][1] + sh * (x[1] - mix.mu[k][1]));
...
    for (i = 0; i < n; i++) { var d = L.denoise(mix, x, sg[i]), h = (sg[i + 1] - sg[i]) / sg[i]; x = [x[0] + h * (x[0] - d[0]), x[1] + h * (x[1] - d[1])]; }

What to try. Start at the defaults: nine codes, 20 labelled pairs at random, 4 passes of 10 ms. With 20 labels 8 of the 9 codes have a name and 6 of the 8 compass buttons are named right, so 87.3 % of pairs are read as the right button, against 50.0 % at zero labels and 99.7 % at 60. Switch which examples to one per code: nine labels give 99.7 % and 8 of 8, where nine random labels in this order give 74.5 % and 4. Move codes: at 8 the purity is 93.3 %; at 9 it is 99.6 % with a decoder error of 0.121 m/s against the floor of 0.122; at 12 the error is 0.102, below the floor, and the code uses 3.36 bits for 2.53 bits of buttons. Move the passes: at 1 the amber rings sit on one code (1 of 9, distance 0.51), at 3 on all nine (0.26), at 16 the distance is 0.048. Move time per pass: at 10 ms four passes take 40 ms (25 fps) and five take 50 (20 fps), while at 8 ms five fit in 40. Finally choose the wall (§6): a smooth model fitted to clean clips is wrong by 0.020 m/s on free steps and 2.73 at the wall steps, 134 times more; fitted to a video with the bounces left in, 0.158 and 2.45, still 15.5 times more.

6 · How long a session lasts

A session feeds the model's own frames back as its next input: lesson 6's recursion. GameNGen reports that without noise added to the context frames in training, quality degrades fast after 20 to 30 steps; that noise, per-frame noise levels (Diffusion Forcing) and training on the model's own rollouts (Self Forcing) are three repairs of the same gap. In free flight the Courtyard shows how mild the problem is. The smooth model of this section is a linear map plus 48 random tanh features, with output weights fitted by ridge regression, from state and push to the change of state. Fitted to the clean clips and driven by a random player through 4 s sessions, its position error after 10, 20 and 40 steps is 0.040, 0.121 and 0.354 m (medians over 158 sessions that never touch a wall).

Now put the walls back. In the Courtyard 2.1 % of the steps of this player's video touch a wall, and 60.5 % of 4 s sessions do. Score the model step by step on fresh video with walls, the true push given so that every error is the model's. On steps that touch nothing the velocity error is 0.020 m/s. On steps that touch a wall it is 2.73 m/s, 134 times more (88 of its steps touch a wall). The natural repair is to leave the bounces in the training video. Fitted that way the model is wrong by 2.45 m/s at wall steps and 0.158 elsewhere: still 15.5 times worse at the walls, and now 7.8 times worse everywhere else than before, because a smooth function that cannot make a corner spreads the damage. It is not a matter of capacity or data: a 32 × 32 network fitted to four times the video is at 2.48 m/s at the walls and 0.121 elsewhere, 20 times worse.

A session is over at its first wall. Right after the wall step the position error is 0.20 m (median), and four steps later 1.14 m, which is 9.4 times what 20 steps of free flight cost. The model's ball goes through the wall, or leaves it the wrong way, and every later frame inherits the error.

What this lesson did not do
It built a latent action model out of the smallest parts that still compute. The decoder is the passive law plus a push, the quantiser is k-means, the frames are readings and not pictures, the clips stop 0.4 m before a wall until §6, and the sampler used the exact denoiser, so its only error is the step count; a real latent action model is a VQ-VAE over pixels and a real denoiser is trained. Its codes capture whatever the past cannot predict, so a second moving object, a camera cut or a bounce would claim codes. The player pressed at random; one who chooses the button from the state brings lesson 7's confounding into what the model learns about the button. Few-step distillation was described, not run, and every statement about a real system is the paper's. Contact, force and words are lesson 15's.

Common mistakes / failure modes

"a latent action is what the player pressed"
It is whatever the past cannot predict, in K symbols. Here only the buttons are unexplained, so nine codes are nine buttons; a bounce or a second ball would claim codes too (§3, §6).
"code 3 means north"
Codes have no names: the 9! = 362,880 relabellings leave every prediction unchanged. Naming takes labels (§3, §4).
"a wider codebook learns more of the action"
Past K = 9 the error falls below the floor and the extra bits carry noise: K = 64 uses 5.71 bits, 2.50 about the button (§3).
"an inverse dynamics model can be used as the policy"
It reads the frame after the push. A predictor that sees only the frame before is off by 0.424 m/s, the IDM by 0.113 (§2).
"more denoising passes make a model controllable"
At 16 passes the draws match the press rates and the controllability is 1.00. Control comes from the input, plausibility from the passes (§5).
"a model that is good in free flight is good in the Courtyard"
At wall steps a smooth model fitted to clean clips is 134 times worse, and still 15.5 times worse when the bounces are in its training video (§6).

Checkpoint exercise

Try it
A world model must run at 30 frames per second and one pass of its network takes 8 ms. (a) How many passes fit in a frame? (b) Its player presses neutral 70 % of the time and each of eight other buttons 3.75 % of the time. About how many random labelled pairs are needed to meet every button, using 1/p times H8 for the eight rare ones? (c) With 16 frames of context of 256 tokens each, how many attention pairs does a causal model with cached keys and values need for one new frame, and how many does recomputing the window need? Answer: (a) 1000/30 = 33.3 ms, and ⌊33.3/8⌋ = 4 passes (32 ms). (b) (1/0.0375) × 2.72 = 72.5 labels (the exact count is the same to this precision). (c) 256 × 4096 = 1,048,576 against 4096² = 16,777,216: 16 times fewer.

Where this points next

Actions can be inferred from video. A code of nine symbols found the player's nine buttons without a label (99.6 % pure), nine well-chosen labels named them, and the model steers, with a controllability of 5.4 against 1.00 for the same video without the code. It also runs in real time: given a button, one pass lands on its push, and four passes of 10 ms fit a 24 frames per second budget. All of this is a world seen through a screen, where a button changes the picture smoothly. The Courtyard's own walls break it. At a wall the velocity reverses within one step. The smooth model fitted to clean clips, wrong by 0.020 m/s in free flight, is wrong by 2.73 m/s at a wall step, 134 times more; the one fitted with the bounces in its video is still 15.5 times worse there than elsewhere. (Both are random-feature models scored by wall steps; lesson 15 trains a tanh network and scores it by distance to the wall, so its ratios differ.) A robot's world is touched, not only seen, and a camera cannot see the force. What changes when the agent has a body, and the model must predict what it will feel?

Takeaway
An action is the part of the change between two frames that the past does not predict. In free flight the Courtyard's law reads it off, a = v′/d − v, to a floor of 0.122 m/s set by the readings; an inverse dynamics model learns that map from labelled pairs and is non-causal (it reads the next frame), which is why labelling is easier than acting. A latent action model needs no labels: a code of K symbols that must explain the next frame recovers the player's alphabet when K matches it (nine codes, 99.6 % pure, error at the floor) and swallows noise when it is larger. Its codes have no names; a handful of labels, one per code if well chosen, supplies them, and the exam is controllability: 5.4 with the codes, 1.00 without. A press must be answered within a frame time, 41.7 ms at 24 frames per second: without a code a model must draw the player's choice itself and needs several passes to reach all nine modes, while given the code one pass lands on the push, and a causal model that caches keys and values answers in the next frame. All of it treats the world as smooth: a smooth model fitted to clean clips is 134 times worse at a wall step than in free flight.

Interview prompts

Companion reads: Lesson 23 · The actions do not exist: infer them (latent actions at corpus scale), Lesson 29 · Distilling into a real-time loop (few-step students) and Robot Model Training · 09 Watching: learning from video without actions (latent actions for robots).