World Models, from first principles
A reconstruction explains what already happened. An agent that has to act needs the other direction: if I do this, what happens next? This track builds the machine that answers it, one forced step at a time, in a world small enough to simulate in your browser and rich enough to hold every failure the field has met. It starts from the economics of trying things, derives what a model must carry in its head, makes its predictions honest, uses them to practise and to plan, scales the same ideas to pictures, places, things and bodies, and ends by grading the model the only way that matters: by the decisions it supports. The last fifteen lessons then ask how such a model is trained: which loss, which schedule, which conditioning, which mixture, which deadline, and what the whole bill comes to.
The exam
The whole track is graded by one question asked in different ways: if the agent does a, what happens? A model that answers it well enough lets the agent choose actions without trying them, and the measure of "well enough" is always a decision: how often the ball ends up in the goal when the agent acts on the model's answer. Lesson 1 shows why that is the right measure by pricing a trial against a thought; the later lessons are the moves forced by what goes wrong. Part VI turns the question round: what has to be true of the training for the model to pass.
- What to carry. The last observation is not enough, because velocity is not in a single reading and the curtain hides the ball. A belief fixes that for a world we can write down, a predictive state keeps what a picture hides in its noise, and a sequence model learns it from data (lessons 1–4).
- What to trust. One predicted future is the average of several, an average of a fork is a state the world never visits, and a model that eats its own predictions drifts. A distribution, a calibration check, a trustworthy horizon, and the difference between watching and doing are what make a prediction usable (lessons 5–7).
- How to use it. Practise inside the model without being fooled by its flattering errors, plan with it at the moment of decision, and think in bigger steps to see farther (lessons 8–10).
- How it scales. Pixels, places that must be remembered, things that can be counted, actions nobody labelled, and a body that touches the world (lessons 11–15).
- How to know. A ladder of exams, from how it looks to what it decides, with the failures that rank them differently (lesson 16).
- How to train one. What “trained” must mean for a simulator, the bit budget and the tokenizer ceiling, teacher forcing against rollout, how the action gets in and where its labels come from, heads that do not average, the quadratic bill for memory, staging and mixture, post-training, distillation into a real-time loop, training-side evaluation, and the whole bill (lessons 17–31).
Part I · Why predict, and what to carry (lessons 01–04)
The economics of consequence, the state an agent cannot see, what a learned state should keep, and the filter that produces it.
Part II · Futures you can trust (lessons 05–07)
Distributions instead of points, errors that compound, and the difference between watching and doing.
Part III · Using the model (lessons 08–10)
Practice inside the model, plan with it at decision time, and think coarser to see farther.
Part IV · Scaling the world (lessons 11–15)
Pixels, places, things, unlabeled actions and real time, and finally a body.
Part V · Knowing it works (lesson 16)
Grade the model by the decisions it supports, not by how it looks.
Part VI · Training a world model (lessons 17–31)
What "trained" has to mean for a simulator, the bit budget and the tokenizer ceiling, how the action gets in and where its labels come from, heads that do not average, the quadratic bill for memory, staging and mixture, post-training, distillation into a real-time loop, and the whole bill. These lessons keep the layout they were written in.
The whole derivation on one page
Read the second column of any row and the fourth column of the row above it: they are the same sentence. That is the linear thinking made visible. Each lesson's problem it inherits is exactly the previous lesson's problem it creates, and the middle column is the single move that resolves the first and manufactures the second. The first row inherits its problem from the last lesson of 3D Vision, from first principles, where a scene that changes can be reconstructed from what has happened but nothing predicts what will; the last row hands its problem to lesson 17, which opens Part VI below.
| # | The problem it inherits | The one move | The problem it creates |
|---|---|---|---|
| Part I · Why predict, and what to carry | |||
| 01 | A scene that changes can be reconstructed from what has already happened. An agent that has to act needs the other direction: what will happen next, and what would happen if it did something else. Nothing in a reconstruction answers that, because it has no actions and no future. What must a model carry in its head to answer "if I do this, what happens next?", and why would an agent want one at all? | Why predict? The world-model contract: Weigh thinking against trying: a model turns computation into trials, and the contract it must meet is "if I do this, what happens next?". | A model that answers "if I push this way, where does the ball end up?" makes thinking cheaper than trying: one push in the real world instead of dozens, provided the model is handed the true state of the ball. An agent is never handed the state. It sees a noisy position, and nothing at all while the ball is behind the curtain. What should the agent carry in its head instead of the state it cannot see? |
| 02 | A model that answers "if I push this way, where does the ball end up?" makes thinking cheaper than trying: one push in the real world instead of dozens, provided the model is handed the true state of the ball. An agent is never handed the state. It sees a noisy position, and nothing at all while the ball is behind the curtain. What should the agent carry in its head instead of the state it cannot see? | The state you cannot see: belief: Carry a distribution over the hidden state, updated by predict-then-correct, instead of the last observation. | The belief, a distribution over the hidden state updated by predict-then-correct, solves the curtain problem. But we built it from a model of the ball and a model of the sensor that we wrote down by hand, using state variables we chose. A real agent receives pixels, and nobody tells it which of a million numbers matter. What should an agent keep from what it sees, when the only teacher is the stream itself? |
| 03 | The belief, a distribution over the hidden state updated by predict-then-correct, solves the curtain problem. But we built it from a model of the ball and a model of the sensor that we wrote down by hand, using state variables we chose. A real agent receives pixels, and nobody tells it which of a million numbers matter. What should an agent keep from what it sees, when the only teacher is the stream itself? | What should a state keep?: Keep what predicts what matters: a predictive objective instead of a reconstructive one, and a guard against collapse. | A good state keeps what predicts the future that matters and drops the rest; reconstructing the observation is the wrong way to ask for it, and predicting in a latent space needs protection against collapse. We still lack the machine itself: something that turns a stream of observations and actions into that state and moves it forward. How is the filter learned from data? |
| 04 | A good state keeps what predicts the future that matters and drops the rest; reconstructing the observation is the wrong way to ask for it, and predicting in a latent space needs protection against collapse. We still lack the machine itself: something that turns a stream of observations and actions into that state and moves it forward. How is the filter learned from data? | Learning the filter and the dynamics: Train a sequence predictor: its hidden state becomes the belief, and with actions as inputs it becomes the dynamics model. | A recurrent state trained to predict the next observation becomes the belief, and with actions as inputs it becomes a dynamics model; in a deterministic world it is almost exact. But at the post the ball goes up or down depending on a difference smaller than the sensor can see, and a model trained to minimize squared error predicts the average, a ball that passes straight through the post. What should a model output when the future is not one point? |
| Part II · Futures you can trust | |||
| 05 | A recurrent state trained to predict the next observation becomes the belief, and with actions as inputs it becomes a dynamics model; in a deterministic world it is almost exact. But at the post the ball goes up or down depending on a difference smaller than the sensor can see, and a model trained to minimize squared error predicts the average, a ball that passes straight through the post. What should a model output when the future is not one point? | One future is a lie: distributions: Output a distribution: squared error predicts the mean, so use mixtures, discrete latents or samplers, and check calibration. | A model can now output a distribution and draw samples: several futures, each plausible. Using it means feeding each prediction back in as the next input, again and again, though the model was only ever trained on real inputs. A small error becomes the input of the next step. How fast do errors grow when a model predicts from its own predictions, and how far ahead can it be trusted? |
| 06 | A model can now output a distribution and draw samples: several futures, each plausible. Using it means feeding each prediction back in as the next input, again and again, though the model was only ever trained on real inputs. A small error becomes the input of the next step. How fast do errors grow when a model predicts from its own predictions, and how far ahead can it be trusted? | Errors compound: rollouts and horizons: Measure how error grows when a model eats its own predictions, train on rollouts, and re-observe to reset the clock. | A rollout has a trustworthy horizon that we can now estimate, and re-observing resets it. But every model so far was fitted to logs that someone generated by acting. If whoever acted reacted to something the model cannot see, such as a gust of wind, the logs confound what an action did with why it was taken. Can a model fitted to logs tell what an action causes? |
| 07 | A rollout has a trustworthy horizon that we can now estimate, and re-observing resets it. But every model so far was fitted to logs that someone generated by acting. If whoever acted reacted to something the model cannot see, such as a gust of wind, the logs confound what an action did with why it was taken. Can a model fitted to logs tell what an action causes? | What does my action cause?: Separate seeing from doing: randomize, adjust for what you can see, put the confounder into the state, and cover the actions you will consider. | A model fitted to varied, randomized actions can say what an action causes, at least where the data went. The cheapest use of it is practice: let a policy improve on the model's imagined rollouts instead of in the world. But a policy trained against a model finds the model's flattering mistakes and climbs them. How do we learn from imagined experience without being fooled by it? |
| Part III · Using the model | |||
| 08 | A model fitted to varied, randomized actions can say what an action causes, at least where the data went. The cheapest use of it is practice: let a policy improve on the model's imagined rollouts instead of in the world. But a policy trained against a model finds the model's flattering mistakes and climbs them. How do we learn from imagined experience without being fooled by it? | Practice in imagination: Improve a policy on the model's rollouts, pessimistically, because the optimizer will find every flattering error. | Practice produces a policy that is only as good as the model was where it practiced. When the goal changes, or the situation is new, the policy has nothing to say. What if the agent used the model at the moment of decision, searching over what to do next, and what does it do about a model that is trustworthy for only a few steps? |
| 09 | Practice produces a policy that is only as good as the model was where it practiced. When the goal changes, or the situation is new, the policy has nothing to say. What if the agent used the model at the moment of decision, searching over what to do next, and what does it do about a model that is trustworthy for only a few steps? | Plan with it: search at decision time: Search over actions with the model at the moment of decision, execute the first step, and plan again. | A planner can look past the horizon its model is trustworthy for, because it re-plans and uses the far end of a plan as a direction. But the maze already needed about fifty steps of look-ahead, each paid for at every decision, and the tasks that matter, crossing a building or cooking a meal, take thousands. A better model lowers the error of each step, not the number of steps; the way to look farther at the same price is to take fewer, bigger steps. How can a model predict in jumps, and what does it lose? |
| 10 | A planner can look past the horizon its model is trustworthy for, because it re-plans and uses the far end of a plan as a direction. But the maze already needed about fifty steps of look-ahead, each paid for at every decision, and the tasks that matter, crossing a building or cooking a meal, take thousands. A better model lowers the error of each step, not the number of steps; the way to look farther at the same price is to take fewer, bigger steps. How can a model predict in jumps, and what does it lose? | Think coarser to see farther: temporal abstraction: Predict in jumps: fewer, bigger steps push the horizon out, at the price of detail. | The mechanism is complete in a toy where the state is four numbers and the observation is a noisy dot. Real observation is a video: a million numbers per frame, almost all of them irrelevant, and the same ball can be drawn a million ways. Each piece we built, state, dynamics, uncertainty and rollouts, has to survive that. What happens to a world model when the observation is pixels at internet scale? |
| Part IV · Scaling the world | |||
| 11 | The mechanism is complete in a toy where the state is four numbers and the observation is a noisy dot. Real observation is a video: a million numbers per frame, almost all of them irrelevant, and the same ball can be drawn a million ways. Each piece we built, state, dynamics, uncertainty and rollouts, has to survive that. What happens to a world model when the observation is pixels at internet scale? | Pixels: the world model becomes a video model: Encode frames into tokens or latents, and predict them with a generative sequence model conditioned on actions. | A video model predicts the next frames convincingly, but its state is the last few hundred frames it is allowed to look at. Walk away from the sofa and come back and the sofa has changed. A finite window is a finite memory. What would let a world model stay consistent about places it has stopped looking at? |
| 12 | A video model predicts the next frames convincingly, but its state is the last few hundred frames it is allowed to look at. Walk away from the sofa and come back and the sofa has changed. A finite window is a finite memory. What would let a world model stay consistent about places it has stopped looking at? | Places: a world you can leave and return to: Remember places in a world frame, separate your own motion from the world's change, and read memory by pose. | A map remembers where things are, not what they are: it cannot say that this cup is the cup that was on the table a minute ago, or what happens when one thing pushes another. A world made of things needs a state made of things. How should a model represent entities and their interactions? |
| 13 | A map remembers where things are, not what they are: it cannot say that this cup is the cup that was on the table a minute ago, or what happens when one thing pushes another. A world made of things needs a state made of things. How should a model represent entities and their interactions? | Things: a world made of objects: Represent entities and let them interact: slots, identity, and pairwise mechanisms that generalize across counts and orders. | Structure inside the model fixes what it can represent, not what it can learn from. The video on the internet shows what happened and not what anyone did to make it happen, and an interactive world must answer a button press now, not after a minute of sampling. Where do actions come from when the data has none, and how fast can a model run? |
| 14 | Structure inside the model fixes what it can represent, not what it can learn from. The video on the internet shows what happened and not what anyone did to make it happen, and an interactive world must answer a button press now, not after a minute of sampling. Where do actions come from when the data has none, and how fast can a model run? | Actions without labels, worlds in real time: Infer the actions from the video, then make the model causal, few-step and fast enough to answer a button press. | Actions can be inferred from video and the model can run in real time, for a world seen through a screen. A robot's world is touched, not only seen: contact is discontinuous, force is invisible to a camera, and the goal arrives in words. What changes when the agent has a body, and the model must predict what it will feel? |
| 15 | Actions can be inferred from video and the model can run in real time, for a world seen through a screen. A robot's world is touched, not only seen: contact is discontinuous, force is invisible to a camera, and the goal arrives in words. What changes when the agent has a body, and the model must predict what it will feel? | Bodies, contact, and words: Add the body: proprioception, touch and force, contact modes, and goals given in language. | We now have every piece: state, dynamics, uncertainty, rollouts, causality, practice, planning, abstraction, pixels, memory, objects, actions and a body, each justified by an exam the previous model failed. How do we know that a particular model passes the exam that matters, a better decision in the world, and not only the ones that look good? |
| Part V · Knowing it works | |||
| 16 | We now have every piece: state, dynamics, uncertainty, rollouts, causality, practice, planning, abstraction, pixels, memory, objects, actions and a body, each justified by an exam the previous model failed. How do we know that a particular model passes the exam that matters, a better decision in the world, and not only the ones that look good? | The exam we never stopped taking: evaluation: Grade the model by the decisions it supports: a ladder of exams, rank reversals, exploitation, and a failure atlas. | Everything in this series assumed that a trained world model exists. How is one trained: which loss, which schedule, which data, and what does it cost? |
| Part VI · Training a world model | |||
| Lessons 17–31 keep the layout they were written in: each opens with a box that says what forced it and closes with a hand-off, but those sentences are not diffed against this table and their numbers are not re-derived by script. | |||
Part VI in brief
The first sixteen lessons derived what a world model is and what you do with one, assuming a trained model existed. Part VI discharges that assumption. It is the engineering of the forward map p(o′|o,a), which loss, which schedule, which conditioning, which mixture, which deadline, derived one forced step at a time. These lessons keep the layout they were written in, which is why the table above stops at lesson 16.
The one number that frames everything
Meta's V-JEPA 2 was pretrained on over 1,000,000 hours of internet video. Its action-conditioned counterpart, V-JEPA 2-AC, was trained on roughly 62 hours of robot video carrying an end-effector state.
free video : action-grounded video ≈ 1,000,000 : 62 ≈ 16,000 : 1Read it as a price, not a lament. Four orders of magnitude separate the column you can have from the column you need, and the field's entire method catalogue (inverse dynamics, latent actions, simulation, cross-embodiment pooling, training inside a learned model) is arbitrage across that spread.
Why this order and no other
Each lesson exists because the previous one's solution created a new failure. That is the only organising principle here.
How to read Part VI
- Run the free checks first. The blur floor (lesson 20) and the per-tick evaluation budget (lessons 18, 24, 29) are arithmetic, cost nothing, and determine your architecture. Every project that skips them rediscovers them expensively.
- Never let fidelity be the gate. It is one of four acceptance coordinates and the least important. Keep it as a cheap regression alarm.
- Distinguish "behaves wrongly" from "cannot see." The first is a training problem. The second is upstream, and no amount of training touches it.
- Treat every mixture weight as an epoch count. Divide by availability before believing the weight.
- Assume your consumer is an optimiser. It will find whatever your model got wrong, and the better it is, the better it finds it.
python3 tools/chain/validate_chain.py --series wm --widgets --oracles from the repository root: it re-derives each one in lessons 1–16 with an independent script.
Where this track sits
Before it: 3D Vision, from first principles ends on the hand-off this track begins with. Beside it: Reinforcement Learning · 07 Planning and 19 Offline RL cover the control side that lessons 8–9 build on; Computer Vision · 16 Video and temporal models and Generative Models · 11 VQ tokenizers cover the machinery lesson 11 puts to work. After it: Part VI of this track asks how such a model is trained, with which loss, schedule and data, and what it costs; Robot Model Training then follows the same questions into policies and into the data they eat.