all_lessons/World Models/11 · Video modelslesson 11 / 31

Pixels: the world model becomes a video model

Ten lessons ran on a state of four numbers and a sensor that is a noisy dot. Make the observation a picture and every piece is rebuilt. The state becomes the last W pictures, each a short list of discrete tokens whose bits cap everything above it. The future becomes a distribution over those tokens, and it gives a picture only when it is drawn token by token. The action becomes an input whose effect we measure, and the bill grows with the square of the window. What the model cannot do is remember: a ball hidden for longer than the window comes back somewhere plausible and wrong.

The thesis, here
A video model keeps the four parts of the earlier lessons and changes what they are made of. Each picture is a code whose bits cap everything above it; the next picture is drawn token by token, so that it is one coherent picture; an action is an input whose effect we measure. The state is the last W pictures, so what left the window is gone, and a wider window costs the square.
Linear position
Forced by: The mechanism is complete in a toy where the state is four numbers and the observation is a noisy dot. Real observation is a video: a million numbers per frame, almost all of them irrelevant, and the same ball can be drawn a million ways. Each piece we built, state, dynamics, uncertainty and rollouts, has to survive that. What happens to a world model when the observation is pixels at internet scale?
New idea: a world model of pictures is a tokenizer plus a sequence model: code each picture as a short list of discrete tokens, whose bits are a ceiling, and draw the next picture token by token, given the last W pictures and the action. Its state is that window, and what leaves the window is gone.
Forces next: A video model predicts the next frames convincingly, but its state is the last few hundred frames it is allowed to look at. Walk away from the sofa and come back and the sofa has changed. A finite window is a finite memory. What would let a world model stay consistent about places it has stopped looking at?
The plan
Six moves. (1) Ask the four questions of a picture. (2) Code the picture and measure the ceiling. (3) Draw the next picture so that it is a picture. (4) Feed the draws back and test the action. (5) Count the bill. (6) Ask what the window forgets.

1 · The four questions, asked of a picture

Lesson 3 gave the Courtyard pictures: 24 × 15 = 360 numbers, the ball a Gaussian blob and the curtain a dim band. They say where the ball is (two numbers) and otherwise repeat a floor, a goal disc and a curtain that never change. A real frame, 256 × 256 pixels in three colours, is 196,608 numbers (lesson 10), and what matters in it can be drawn in many ways: in lesson 3's pictures with the band on, 81 % of the pixel variance is nuisance and 19 % is the ball. Each of four questions the earlier lessons answered has to be answered again.

questionwith four numberswith a picture
the statea belief over (x, y, vx, vy): lessons 2–4a short list of discrete tokens per picture, and the last few pictures (§2)
the next step, many futuress′ = f(s, a), a Gaussian mixture: lesson 5a distribution over the next picture's tokens, given the window and the action; one Gaussian over 360 numbers averages the modes, a categorical over tokens can keep each (§3)
a long runfeed each prediction back in: lesson 6draw a picture, append it, drop the oldest (§4)
what an action doesan input of f: lesson 7a conditioning input the model may ignore (§4)

Three constraints decide the design. The code must be short, because window, sampling and attention are paid per token. It must give a distribution that can be sampled exactly and have many modes, since the mean of the modes is not an outcome (lesson 5). And it must keep what matters, because what the code drops is gone for every later step. Discrete tokens meet all three; continuous latents drawn by a denoiser do too, and §3 and §5 price that road. We build the first.

2 · The tokenizer is the ceiling

Cut the picture into 3 × 3 patches (a square metre of floor): 8 × 5 = 40 patches of 9 numbers. A codebook of C codes is a list of C patches. Replace each patch by the index of its nearest code, a token, and the picture is 40 tokens. A token costs log2C bits and a picture 40 log2C: 200 bits at C = 32, against 360 numbers. Decoding pastes each code's patch back. We choose the codes by k-means on 400 pictures: assign every patch to its nearest code, move each code to the mean of its patches, repeat; no step raises the squared error.

Two scores. PSNR, 10 log10(1/MSE) for brightness in [0, 1], scores the whole picture. The ball's floor scores the ball: tokenise a picture with and without the ball, decode the tokens that differ, read a position off them (the centroid of what they add; a picture whose tokens do not differ has lost the ball and counts as 3 m); the median error is how well these tokens locate the ball.

codes Cbits per picturePSNR (dB)ball floor (cm)
24023.2110, and the ball is lost in 36 % of pictures
812029.617.5
3220041.47.0
12828048.34.3

The floor is the ceiling: a model that predicted the tokens perfectly would still place the ball only to 7.0 cm at 200 bits and 17.5 cm at 120 by this reading, like lesson 1's 10 cm sensor (9.6 cm costs 160 bits). The ball does not vanish gracefully: gone at C = 1, lost in 36 % of pictures at 2, found in every visible picture from 4 codes up, and the floor is not monotone at the small end (14.3 cm at 4, 17.5 at 8): the codebook minimises squared error over all patches and is not asked to find the ball, so an extra code need not go to it. PSNR keeps rising where the floor stops: from 64 to 128 codes it gains 4.1 dB while the floor moves from 4.6 to 4.3 cm. A fixed-length code spends its bits on every token alike: at C = 32 only 6.0 of the 40 tokens carry the ball, and the other 34, 170 of the 200 bits, are the same in every picture.

Lesson 3's nuisance spends bits too. Add its band (0.4 m wide, as bright as the ball, at a random place in each picture) and retrain: at C = 16 the PSNR falls 11.4 dB, from 37.5 to 26.1, while the ball is still found about as well (10.0 cm against 9.6). But one ball position with the band at 200 places comes out, at C = 64, as 121 different lists of tokens, which the sequence model must learn are one ball.

3 · Dynamics on tokens: draw the picture token by token

The model must say how likely each next picture is, given the last W pictures and the action. There are C40 pictures (about 1060 at C = 32), too many to list, but the chain rule factors the distribution exactly:

p(x1, …, x40 | window, action) = p(x1 | ·) · p(x2 | x1, ·) ⋯ p(x40 | x1, …, x39, ·)

Each factor is a categorical distribution over the C codes, given the window, the action and the tokens already drawn. A transformer learns these 40 conditionals (IRIS, Micheli et al., 2023; GAIA-1, Wayve, 2023). Here no network is trained: the Courtyard is small enough for a look-up of recorded windows to stand in for the transformer. Among the 19,200 windows of 600 recorded clips, take the three whose last W pictures are closest to the context (squared distance between decoded pictures, among windows that show the ball in the same pictures; ties are kept, up to 40; if the model is told the action, a window with a different action is pushed away) and use what came next in them.

The uncertainty: the operator may nudge the ball once while it is behind the curtain, up, down or not at all (the nudge of lesson 1), and the model sees nothing of it, so the picture where the ball comes out has three versions, three modes as at lesson 5's fork, now in a picture. For 90 launches nudged while the ball is hidden (W = 16, C = 32, the model not told), draw that picture four ways from the same recorded futures:

draw byhowone clean ballno ballpieces or smearnot a recorded picture
pixel meanaverage the candidates36054always
greedy tokenper token, the most common code4941057
independent tokenper token, its own distribution5533285
token by tokenthe chain rule85500

The pixel mean is lesson 5's answer: three balls at three heights average to faint balls or a smear. Greedy keeps the most common code per token; where the futures disagree about the ball that code is the empty floor, so the ball is deleted. Independent draws give each token its own distribution: every token is right and the joint is wrong, token 7 from the future with the ball high and token 12 from the one with it low. Token by token is the chain rule: draw token 1, keep the candidates that agree with it, draw token 2 from those, and on. The probability of a picture is the product of the conditionals, the share of candidates that show it, so every draw is a recorded future (with a look-up, a random pick among the candidates; a transformer has no list and learns the conditionals).

Road not taken · regress the pixels
Predict the 360 numbers by squared error: no tokenizer, no sampling, one loss, and it trains. It is lesson 5's one-Gaussian model with 360 outputs, and answers like the first row of the table: a clean ball in 36 of 90 pictures, a ghost in 54, and the best PSNR of the four.

What being a picture costs. If the truth and the m candidates are draws from one distribution with variance σ² per pixel, one draw is off by 2σ² and the mean of m draws by σ²(1 + 1/m): a ratio of 2m/(m + 1), 1.5 for m = 3. Measured where the ball comes out, token by token has 1.44 times the squared error of the mean: 18.3 dB against 19.8. The ghost wins the metric and fails the exam.

A second way to draw one coherent picture has no chain rule: start from noise and remove it over several passes, each adjusting the whole picture. DIAMOND (Alonso et al., 2024) draws a frame in 3 denoising steps where IRIS needed 16 network evaluations; GameNGen (Valevski et al., 2025) uses 4. Oasis (Decart and Etched, 2024), UniSim (Yang et al., 2024), the Cosmos diffusion models (NVIDIA, 2025) and Sora (OpenAI, 2024) are diffusion models too. It trades 40 sequential draws for a few passes; §5 prices both.

4 · Roll it out, and tell it what you did

Rolling out is lesson 6's free-running loop: draw a picture, append it, drop the oldest, draw again, so that after one step the model's inputs are its own draws. Ranzato et al. (2016) called the gap between training on real inputs and sampling from one's own outputs exposure bias; GameNGen sees quality degrade within 20 to 30 steps unless its context frames are noised in training, and Diffusion Forcing (Chen et al., 2024) and Self Forcing (Huang et al., 2025) also work on it.

The action is one more input; in the look-up it enters the distance. Does the model obey? Let the operator nudge up, down or not at all while the ball is hidden, draw the picture where the ball comes out and read the ball's height. Controllability is the mean gap between the drawn heights for two different actions, against the mean gap between two draws of one action; the ratio is about 1 if the action is ignored. Over 60 launches (W = 16, C = 32), told the action, drawn heights for different actions are 1.88 m apart, where the true heights are 1.89 m apart, and two draws of one action 0.17 m: a ratio of 11. Not told: 0.83 m between and 0.89 m within, a ratio of 0.93, the same whatever you did. Told, with a window of 4: 0.66 m between and 1.01 m within. An action is obeyed only while the model still holds the ball it steers; §6 is about why it stops holding it.

5 · The bill: tokens, passes and pairs

Tokens and pairs. A 256 × 256 frame in 16 × 16 patches is 256 tokens of 768 numbers each: 2,048 tokens a second at 8 frames a second, 7.37 million an hour (lesson 10). A transformer scores every token against every other in its window, so n pictures of T tokens make (nT)² pairs (a causal mask halves the count and changes nothing below): 640 tokens and 409,600 pairs for our window of 16, and four times the window is 16 times the pairs. At lesson 10's frame size a second at 24 frames per second is 6,144 tokens and 37.7 million pairs, a minute 368,640 tokens and 136 billion pairs: 3,600 times as many for 60 times the window.

Time. At 24 frames per second a frame gets 41.7 ms. Token by token, a 256-token frame takes 256 sequential passes, 0.16 ms each if they are to fit. A denoiser makes K passes over the whole frame, 10.4 ms each at K = 4, 64 times as long a pass. Reported figures: Genie (Bruce et al., 2024) looks back 16 frames and runs at about 1 frame per second. Oasis (Decart and Etched, 2024) uses latent diffusion with 1.6 s of context (the 500M model) at 20 frames per second. DIAMOND (Alonso et al., 2024) takes 3 denoising steps per frame on Atari, and its CS:GO model runs at 10 Hz on an RTX 3090. GameNGen (Valevski et al., 2025) takes 4 steps at 20 frames per second on one TPU-v5, a budget of 12.5 ms a step. Dreamer 4 (Hafner et al., 2025) makes K = 4 passes with 9.6 s of context at 21 frames per second on one H100, a budget of 11.9 ms a pass.

A video model of the Courtyard: tokens, four ways to draw, a window
One clip, every second picture from the last before the ball goes behind the curtain: the truth, what the tokens keep (amber squares carry the ball) and what the model draws free-running with a window of W pictures. Then PSNR and the ball position the tokens keep, against bits; the picture where the ball comes out, drawn four ways, and under the three nudges; and, after Measure, how often the ball is drawn where it is, against W. The model is §3's look-up of recorded windows.
bits per picture
—
PSNR
—
ball position kept
—
tokens that carry the ball
—
the four draws
—
drawn height: between · within actions
—
free-run ball vs truth
—
drawn right, this window
—
same: hidden 9 · hidden 10
—
cost of this window
—
Show the core JS
  for (j = 0; j < N; j++) {
    var id = ids[j], m = (id / T) | 0, t = id % T, d = 0;
    for (i = 0; i < W; i++) {
      var tc = t - i;
      if (!blank) d += tc >= 0 ? frameDist(cb, qfr[i], co, m * (T + 1) + tc) : qfr[i].e;
      if (useAct && (tc >= 0 ? co.act[m * T + tc] : 0) !== qac[i]) d += 1e3;
    }
    dist[j] = d;
  }
...
  var live = []; for (i = 0; i < n; i++) live.push(i);                   // 'chain': token by token, each drawn from the candidates consistent with the tokens drawn so far
  for (p = 0; p < NP; p++) {
    var pick = live[Math.min(live.length - 1, Math.floor(rng() * live.length))], tok = grids[pick][p], keep = [];
    for (q = 0; q < live.length; q++) if (grids[live[q]][p] === tok) keep.push(live[q]);
    out[p] = tok; live = keep;
  }
  return { tok: out };

What to try. At C = 32 a picture costs 200 bits, the PSNR is 41.4 dB and 6.0 of the 40 tokens (the amber squares) place the ball to 7.0 cm. Drag the codebook down: at C = 8, 29.6 dB and 17.5 cm; at C = 4, 27.5 dB and a smaller error, 14.3 cm; at C = 2 the median error is 110 cm with the ball lost in 36 % of pictures. Back at 32, the middle panel draws clip 3 (a nudge up) four ways, model not told: the mean and the independent draw come out in pieces, greedy deletes the ball, token by token draws one. Set the model told and all four draw a ball; the three nudges are drawn 2.59 m apart and two draws of one nudge 0.38 m apart, against 1.46 m and 2.29 m not told. The free-running row puts the ball 0.53 m from the true one when told and 3.06 m when not; at W = 4, not told, it draws no ball where the ball comes out. Then press Measure forgetting (§6).

6 · What the window forgets

The state of this model is its window: §3 conditions the next picture on the last W pictures and the action, and on nothing else. The ball behind the curtain is the Courtyard's walk away from the sofa, and the exam is the walk back. Launch the ball with no nudge, at 2.6 m/s, behind a curtain 1.6 m wide (twice lesson 2's, so that the hiding lasts about a second). It goes behind the curtain at picture 8 and comes out at 17 or 18, hidden for 9 or 10 pictures, 0.9 or 1.0 s. Give the model the true pictures up to picture 7, the last that shows the ball, and let it run free, told nobody acts. At the picture where the ball comes out, draw it token by token; it is right if a ball is within 0.5 m of the true one. The widget runs 120 launches per window.

window W (pictures)246810121620
ball right where it comes out (%)58111651728386

A staircase. A window too short to reach the ball does not draw nothing: at W = 8 a ball is drawn in 84 % of launches and is in the right place in 16 %, a plausible ball from recorded windows that look like a ball that went behind the curtain a while ago. That is the exit of the lesson with a number: past its window the model re-draws the ball somewhere plausible, not where it is.

Time away. Split the launches by how long the curtain hides the ball. At W = 10 the ball hidden for 9 pictures is drawn right in 69 % of launches and the ball hidden for 10 in 33 %: the window reaches the last picture that shows the ball when the hiding lasts 9 and falls one short when it lasts 10. Reaching the last sighting takes W = hidden + 1; the velocity needs two sightings (lesson 4's window), W = hidden + 2, and by then the curve has risen (72 % at 12).

Two things the step is not. It is not compounding error: while the ball is hidden the true pictures are empty, and with the true pictures in the window the step is where it was (8 % at W = 8, 63 % at 12, 85 % at 16). And its height is not memory: above the step the model is right in 83 %, not 100, because the look-up knows the physics only as well as its 600 clips; at W = 16 it reaches 54 % with 150 clips and 94 % with 2,400. The window moves the step, data raises the top.

Why not widen it? Ten times the window is 100 times the attention (409,600 pairs become 41.0 million here), and it only moves the step; a finer picture (4 times the tokens) costs 16 times the pairs at one window. Real windows are seconds: WHAM (Microsoft, 2025) looks back 1 s, Oasis 1.6 s and Dreamer 4 9.6 s. Genie 3 (DeepMind, 2025) reports visual memory back to about a minute, while DIAMOND on CS:GO and Oasis list limited memory as a weakness.

What this lesson did not do
It replaced the learned sequence model by a look-up of recorded windows and the learned tokenizer by k-means on patches; real systems train both and generalise to windows no clip contains. It priced the denoiser without running one. Actions were given; in real video nobody labels them (lesson 14). One ball and one fixed camera: no second object (lesson 13), no moving viewpoint (lesson 12). And it graded pictures by the ball's position and by PSNR; grading by the decisions a model supports is lesson 16.

Common mistakes / failure modes

"better PSNR means a better world model"
The pixel mean has the best PSNR (19.8 dB against 18.3) and one clean ball in 36 of 90 pictures where token by token has 85 (§3).
"each token has a distribution, so sample each token"
Every marginal is right and the joint is not: 85 of 90 pictures were ones no recorded future showed (§3).
"an action input makes the model controllable"
Not told, a ratio of 0.93; told, with a window of 4, 0.66 m against 1.01. The ball must be in the window (§4).
"it forgot because errors compounded"
With the true pictures in the window the step is where it was (8 % at W = 8, 85 % at 16). The information left the window (§6).

Checkpoint exercise

Try it
A video model sees 128 × 128 RGB pictures cut into 8 × 8 patches, with a codebook of 1,024 codes, keeps a window of 24 pictures at 20 frames per second, and draws each frame with 4 denoising passes. (a) How many tokens and bits is one picture? (b) How many tokens and attention pairs in the window? (c) How long may one pass take? (d) A ball goes behind a wall for 2 s: what does the model know about it when it comes out? Answer: (a) (128/8)² = 256 tokens and 256 × log21024 = 2,560 bits, for 49,152 numbers. (b) 24 × 256 = 6,144 tokens and 6,144² = 37.7 million pairs. (c) 1000/20/4 = 12.5 ms. (d) The window is 24/20 = 1.2 s, shorter than 2 s: the last picture of the ball has left it, and the model draws a ball from what balls usually do, not from this one.

Where this points next

A video model of the Courtyard keeps all four pieces: a code that caps what it can say, a next picture drawn token by token, an action obeyed while the model still holds the ball, and a window that costs the square of its length. That window is its state. With the ball hidden for about a second, a window of 8 pictures draws it in the right place in 16 % of launches and a window of 16 in 83 %, and four times the window costs 16 times the pairs. A model can only guess at what lies outside what it holds: walk away from the sofa and come back and it has changed. A finite window is a finite memory. What would let a world model stay consistent about places it has stopped looking at?

Takeaway
A world model of pictures is a tokenizer plus a sequence model. The tokenizer fixes a ceiling (7.0 cm for the ball at 200 bits), and a fixed-length code spends most of its bits on what never changes. The sequence model is a distribution over a picture's 40 tokens, and only the chain rule returns a picture: 85 clean balls in 90 against 36 for the pixel mean, which has the better PSNR. An action is obeyed when the gap across actions exceeds the gap within one (11 told, 0.93 not). The bill is tokens, passes and (nT)² pairs. The state is the window: the step of the forgetting curve sits at the hidden time plus one, and more clips raise the top, not the step.

Interview prompts

Companion reads: Lesson 20 · The tokenizer is the ceiling (tokenizers as a recipe), Lesson 25 · Horizon, memory, and the quadratic bill (the window's cost), Generative Models · 11 VQ-VAE · VQGAN · FSQ (vector quantisation), and Computer Vision · 16 Video and temporal models.