Pixels: the world model becomes a video model
Ten lessons ran on a state of four numbers and a sensor that is a noisy dot. Make the observation a picture and every piece is rebuilt. The state becomes the last W pictures, each a short list of discrete tokens whose bits cap everything above it. The future becomes a distribution over those tokens, and it gives a picture only when it is drawn token by token. The action becomes an input whose effect we measure, and the bill grows with the square of the window. What the model cannot do is remember: a ball hidden for longer than the window comes back somewhere plausible and wrong.
New idea: a world model of pictures is a tokenizer plus a sequence model: code each picture as a short list of discrete tokens, whose bits are a ceiling, and draw the next picture token by token, given the last W pictures and the action. Its state is that window, and what leaves the window is gone.
Forces next: A video model predicts the next frames convincingly, but its state is the last few hundred frames it is allowed to look at. Walk away from the sofa and come back and the sofa has changed. A finite window is a finite memory. What would let a world model stay consistent about places it has stopped looking at?
1 · The four questions, asked of a picture
Lesson 3 gave the Courtyard pictures: 24 × 15 = 360 numbers, the ball a Gaussian blob and the curtain a dim band. They say where the ball is (two numbers) and otherwise repeat a floor, a goal disc and a curtain that never change. A real frame, 256 × 256 pixels in three colours, is 196,608 numbers (lesson 10), and what matters in it can be drawn in many ways: in lesson 3's pictures with the band on, 81 % of the pixel variance is nuisance and 19 % is the ball. Each of four questions the earlier lessons answered has to be answered again.
| question | with four numbers | with a picture |
|---|---|---|
| the state | a belief over (x, y, vx, vy): lessons 2–4 | a short list of discrete tokens per picture, and the last few pictures (§2) |
| the next step, many futures | s′ = f(s, a), a Gaussian mixture: lesson 5 | a distribution over the next picture's tokens, given the window and the action; one Gaussian over 360 numbers averages the modes, a categorical over tokens can keep each (§3) |
| a long run | feed each prediction back in: lesson 6 | draw a picture, append it, drop the oldest (§4) |
| what an action does | an input of f: lesson 7 | a conditioning input the model may ignore (§4) |
Three constraints decide the design. The code must be short, because window, sampling and attention are paid per token. It must give a distribution that can be sampled exactly and have many modes, since the mean of the modes is not an outcome (lesson 5). And it must keep what matters, because what the code drops is gone for every later step. Discrete tokens meet all three; continuous latents drawn by a denoiser do too, and §3 and §5 price that road. We build the first.
2 · The tokenizer is the ceiling
Cut the picture into 3 × 3 patches (a square metre of floor): 8 × 5 = 40 patches of 9 numbers. A codebook of C codes is a list of C patches. Replace each patch by the index of its nearest code, a token, and the picture is 40 tokens. A token costs log2C bits and a picture 40 log2C: 200 bits at C = 32, against 360 numbers. Decoding pastes each code's patch back. We choose the codes by k-means on 400 pictures: assign every patch to its nearest code, move each code to the mean of its patches, repeat; no step raises the squared error.
Two scores. PSNR, 10 log10(1/MSE) for brightness in [0, 1], scores the whole picture. The ball's floor scores the ball: tokenise a picture with and without the ball, decode the tokens that differ, read a position off them (the centroid of what they add; a picture whose tokens do not differ has lost the ball and counts as 3 m); the median error is how well these tokens locate the ball.
| codes C | bits per picture | PSNR (dB) | ball floor (cm) |
|---|---|---|---|
| 2 | 40 | 23.2 | 110, and the ball is lost in 36 % of pictures |
| 8 | 120 | 29.6 | 17.5 |
| 32 | 200 | 41.4 | 7.0 |
| 128 | 280 | 48.3 | 4.3 |
The floor is the ceiling: a model that predicted the tokens perfectly would still place the ball only to 7.0 cm at 200 bits and 17.5 cm at 120 by this reading, like lesson 1's 10 cm sensor (9.6 cm costs 160 bits). The ball does not vanish gracefully: gone at C = 1, lost in 36 % of pictures at 2, found in every visible picture from 4 codes up, and the floor is not monotone at the small end (14.3 cm at 4, 17.5 at 8): the codebook minimises squared error over all patches and is not asked to find the ball, so an extra code need not go to it. PSNR keeps rising where the floor stops: from 64 to 128 codes it gains 4.1 dB while the floor moves from 4.6 to 4.3 cm. A fixed-length code spends its bits on every token alike: at C = 32 only 6.0 of the 40 tokens carry the ball, and the other 34, 170 of the 200 bits, are the same in every picture.
Lesson 3's nuisance spends bits too. Add its band (0.4 m wide, as bright as the ball, at a random place in each picture) and retrain: at C = 16 the PSNR falls 11.4 dB, from 37.5 to 26.1, while the ball is still found about as well (10.0 cm against 9.6). But one ball position with the band at 200 places comes out, at C = 64, as 121 different lists of tokens, which the sequence model must learn are one ball.
3 · Dynamics on tokens: draw the picture token by token
The model must say how likely each next picture is, given the last W pictures and the action. There are C40 pictures (about 1060 at C = 32), too many to list, but the chain rule factors the distribution exactly:
p(x1, …, x40 | window, action) = p(x1 | ·) · p(x2 | x1, ·) ⋯ p(x40 | x1, …, x39, ·)
Each factor is a categorical distribution over the C codes, given the window, the action and the tokens already drawn. A transformer learns these 40 conditionals (IRIS, Micheli et al., 2023; GAIA-1, Wayve, 2023). Here no network is trained: the Courtyard is small enough for a look-up of recorded windows to stand in for the transformer. Among the 19,200 windows of 600 recorded clips, take the three whose last W pictures are closest to the context (squared distance between decoded pictures, among windows that show the ball in the same pictures; ties are kept, up to 40; if the model is told the action, a window with a different action is pushed away) and use what came next in them.
The uncertainty: the operator may nudge the ball once while it is behind the curtain, up, down or not at all (the nudge of lesson 1), and the model sees nothing of it, so the picture where the ball comes out has three versions, three modes as at lesson 5's fork, now in a picture. For 90 launches nudged while the ball is hidden (W = 16, C = 32, the model not told), draw that picture four ways from the same recorded futures:
| draw by | how | one clean ball | no ball | pieces or smear | not a recorded picture |
|---|---|---|---|---|---|
| pixel mean | average the candidates | 36 | 0 | 54 | always |
| greedy token | per token, the most common code | 49 | 41 | 0 | 57 |
| independent token | per token, its own distribution | 55 | 3 | 32 | 85 |
| token by token | the chain rule | 85 | 5 | 0 | 0 |
The pixel mean is lesson 5's answer: three balls at three heights average to faint balls or a smear. Greedy keeps the most common code per token; where the futures disagree about the ball that code is the empty floor, so the ball is deleted. Independent draws give each token its own distribution: every token is right and the joint is wrong, token 7 from the future with the ball high and token 12 from the one with it low. Token by token is the chain rule: draw token 1, keep the candidates that agree with it, draw token 2 from those, and on. The probability of a picture is the product of the conditionals, the share of candidates that show it, so every draw is a recorded future (with a look-up, a random pick among the candidates; a transformer has no list and learns the conditionals).
What being a picture costs. If the truth and the m candidates are draws from one distribution with variance σ² per pixel, one draw is off by 2σ² and the mean of m draws by σ²(1 + 1/m): a ratio of 2m/(m + 1), 1.5 for m = 3. Measured where the ball comes out, token by token has 1.44 times the squared error of the mean: 18.3 dB against 19.8. The ghost wins the metric and fails the exam.
A second way to draw one coherent picture has no chain rule: start from noise and remove it over several passes, each adjusting the whole picture. DIAMOND (Alonso et al., 2024) draws a frame in 3 denoising steps where IRIS needed 16 network evaluations; GameNGen (Valevski et al., 2025) uses 4. Oasis (Decart and Etched, 2024), UniSim (Yang et al., 2024), the Cosmos diffusion models (NVIDIA, 2025) and Sora (OpenAI, 2024) are diffusion models too. It trades 40 sequential draws for a few passes; §5 prices both.
4 · Roll it out, and tell it what you did
Rolling out is lesson 6's free-running loop: draw a picture, append it, drop the oldest, draw again, so that after one step the model's inputs are its own draws. Ranzato et al. (2016) called the gap between training on real inputs and sampling from one's own outputs exposure bias; GameNGen sees quality degrade within 20 to 30 steps unless its context frames are noised in training, and Diffusion Forcing (Chen et al., 2024) and Self Forcing (Huang et al., 2025) also work on it.
The action is one more input; in the look-up it enters the distance. Does the model obey? Let the operator nudge up, down or not at all while the ball is hidden, draw the picture where the ball comes out and read the ball's height. Controllability is the mean gap between the drawn heights for two different actions, against the mean gap between two draws of one action; the ratio is about 1 if the action is ignored. Over 60 launches (W = 16, C = 32), told the action, drawn heights for different actions are 1.88 m apart, where the true heights are 1.89 m apart, and two draws of one action 0.17 m: a ratio of 11. Not told: 0.83 m between and 0.89 m within, a ratio of 0.93, the same whatever you did. Told, with a window of 4: 0.66 m between and 1.01 m within. An action is obeyed only while the model still holds the ball it steers; §6 is about why it stops holding it.
5 · The bill: tokens, passes and pairs
Tokens and pairs. A 256 × 256 frame in 16 × 16 patches is 256 tokens of 768 numbers each: 2,048 tokens a second at 8 frames a second, 7.37 million an hour (lesson 10). A transformer scores every token against every other in its window, so n pictures of T tokens make (nT)² pairs (a causal mask halves the count and changes nothing below): 640 tokens and 409,600 pairs for our window of 16, and four times the window is 16 times the pairs. At lesson 10's frame size a second at 24 frames per second is 6,144 tokens and 37.7 million pairs, a minute 368,640 tokens and 136 billion pairs: 3,600 times as many for 60 times the window.
Time. At 24 frames per second a frame gets 41.7 ms. Token by token, a 256-token frame takes 256 sequential passes, 0.16 ms each if they are to fit. A denoiser makes K passes over the whole frame, 10.4 ms each at K = 4, 64 times as long a pass. Reported figures: Genie (Bruce et al., 2024) looks back 16 frames and runs at about 1 frame per second. Oasis (Decart and Etched, 2024) uses latent diffusion with 1.6 s of context (the 500M model) at 20 frames per second. DIAMOND (Alonso et al., 2024) takes 3 denoising steps per frame on Atari, and its CS:GO model runs at 10 Hz on an RTX 3090. GameNGen (Valevski et al., 2025) takes 4 steps at 20 frames per second on one TPU-v5, a budget of 12.5 ms a step. Dreamer 4 (Hafner et al., 2025) makes K = 4 passes with 9.6 s of context at 21 frames per second on one H100, a budget of 11.9 ms a pass.
What to try. At C = 32 a picture costs 200 bits, the PSNR is 41.4 dB and 6.0 of the 40 tokens (the amber squares) place the ball to 7.0 cm. Drag the codebook down: at C = 8, 29.6 dB and 17.5 cm; at C = 4, 27.5 dB and a smaller error, 14.3 cm; at C = 2 the median error is 110 cm with the ball lost in 36 % of pictures. Back at 32, the middle panel draws clip 3 (a nudge up) four ways, model not told: the mean and the independent draw come out in pieces, greedy deletes the ball, token by token draws one. Set the model told and all four draw a ball; the three nudges are drawn 2.59 m apart and two draws of one nudge 0.38 m apart, against 1.46 m and 2.29 m not told. The free-running row puts the ball 0.53 m from the true one when told and 3.06 m when not; at W = 4, not told, it draws no ball where the ball comes out. Then press Measure forgetting (§6).
6 · What the window forgets
The state of this model is its window: §3 conditions the next picture on the last W pictures and the action, and on nothing else. The ball behind the curtain is the Courtyard's walk away from the sofa, and the exam is the walk back. Launch the ball with no nudge, at 2.6 m/s, behind a curtain 1.6 m wide (twice lesson 2's, so that the hiding lasts about a second). It goes behind the curtain at picture 8 and comes out at 17 or 18, hidden for 9 or 10 pictures, 0.9 or 1.0 s. Give the model the true pictures up to picture 7, the last that shows the ball, and let it run free, told nobody acts. At the picture where the ball comes out, draw it token by token; it is right if a ball is within 0.5 m of the true one. The widget runs 120 launches per window.
| window W (pictures) | 2 | 4 | 6 | 8 | 10 | 12 | 16 | 20 |
|---|---|---|---|---|---|---|---|---|
| ball right where it comes out (%) | 5 | 8 | 11 | 16 | 51 | 72 | 83 | 86 |
A staircase. A window too short to reach the ball does not draw nothing: at W = 8 a ball is drawn in 84 % of launches and is in the right place in 16 %, a plausible ball from recorded windows that look like a ball that went behind the curtain a while ago. That is the exit of the lesson with a number: past its window the model re-draws the ball somewhere plausible, not where it is.
Time away. Split the launches by how long the curtain hides the ball. At W = 10 the ball hidden for 9 pictures is drawn right in 69 % of launches and the ball hidden for 10 in 33 %: the window reaches the last picture that shows the ball when the hiding lasts 9 and falls one short when it lasts 10. Reaching the last sighting takes W = hidden + 1; the velocity needs two sightings (lesson 4's window), W = hidden + 2, and by then the curve has risen (72 % at 12).
Two things the step is not. It is not compounding error: while the ball is hidden the true pictures are empty, and with the true pictures in the window the step is where it was (8 % at W = 8, 63 % at 12, 85 % at 16). And its height is not memory: above the step the model is right in 83 %, not 100, because the look-up knows the physics only as well as its 600 clips; at W = 16 it reaches 54 % with 150 clips and 94 % with 2,400. The window moves the step, data raises the top.
Why not widen it? Ten times the window is 100 times the attention (409,600 pairs become 41.0 million here), and it only moves the step; a finer picture (4 times the tokens) costs 16 times the pairs at one window. Real windows are seconds: WHAM (Microsoft, 2025) looks back 1 s, Oasis 1.6 s and Dreamer 4 9.6 s. Genie 3 (DeepMind, 2025) reports visual memory back to about a minute, while DIAMOND on CS:GO and Oasis list limited memory as a weakness.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
A video model of the Courtyard keeps all four pieces: a code that caps what it can say, a next picture drawn token by token, an action obeyed while the model still holds the ball, and a window that costs the square of its length. That window is its state. With the ball hidden for about a second, a window of 8 pictures draws it in the right place in 16 % of launches and a window of 16 in 83 %, and four times the window costs 16 times the pairs. A model can only guess at what lies outside what it holds: walk away from the sofa and come back and it has changed. A finite window is a finite memory. What would let a world model stay consistent about places it has stopped looking at?
Interview prompts
- Why is the tokenizer a ceiling for a video world model, and how would you measure it? (§2 — a model cannot beat what its tokens keep; decode the tokens that differ with and without the object and compare.)
- Why does squared error on pixels draw ghosts, and why is a better PSNR not the answer? (§3 — it rewards the mean of the futures; a sample costs 2m/(m + 1) in squared error and is still a picture.)
- A model outputs a distribution per token. Why not take the most likely token, or sample each independently? (§3 — the mode deletes what the futures disagree about, independent draws mix futures; the chain rule conditions on earlier tokens.)
- How does attention cost grow with context, and what does that mean for memory? (§5 — (nT)² pairs: 4 times the window, 16 times the pairs, so windows stay at seconds.)
- How would you test that a video model obeys its action input? (§4 — the gap between outcomes for different actions against the gap between draws of one action.)
- A video model forgets the sofa after you walk away. Why is that not compounding error, and why is a longer window not the whole answer? (§6 — the true pictures are empty while the ball is hidden and the step stays; a longer window costs the square and only moves the step.)
Companion reads: Lesson 20 · The tokenizer is the ceiling (tokenizers as a recipe), Lesson 25 · Horizon, memory, and the quadratic bill (the window's cost), Generative Models · 11 VQ-VAE · VQGAN · FSQ (vector quantisation), and Computer Vision · 16 Video and temporal models.