all_lessons/World Models/index31 lessons · ~16h read

World Models, from first principles

A reconstruction explains what already happened. An agent that has to act needs the other direction: if I do this, what happens next? This track builds the machine that answers it, one forced step at a time, in a world small enough to simulate in your browser and rich enough to hold every failure the field has met. It starts from the economics of trying things, derives what a model must carry in its head, makes its predictions honest, uses them to practise and to plan, scales the same ideas to pictures, places, things and bodies, and ends by grading the model the only way that matters: by the decisions it supports. The last fifteen lessons then ask how such a model is trained: which loss, which schedule, which conditioning, which mixture, which deadline, and what the whole bill comes to.

The seed question
A ball rolls across a courtyard and slips behind a curtain, and you may nudge it once, around the moment it is out of sight. What would you have to be able to predict, and how would you know the prediction can be trusted, to make the ball finish in the goal?
Who this is for
An engineer or researcher who has met the words (Dreamer, MuZero, video world models, latent actions, JEPA) and wants to see why each exists. You need linear algebra, probability and the idea of training a network by gradient descent; the Reinforcement Learning track and the 3D Vision track are good company and neither is required. By the end you can state what a world model is for and what it must output; why an observation is not a state and what a belief is; why a predictive state collapses unless it is guarded; why squared error cannot represent a fork; how errors compound and how far a rollout can be trusted; why a log of actions does not tell you what an action does; how an agent improves a policy inside the model, and how it plans with it; and how the same ideas scale to video, memory, objects, real time and contact. Lesson 16 is a ladder of exams, because a model that looks right and a model that decides right are different things. The last fifteen lessons add what training one involves: which space you predict in and what a fixed bit budget buys you there; why the tokenizer, not the transformer, sets the ceiling on expressible physics; how to stop teacher forcing from hiding the only error that matters; how the action gets in and where action labels come from when they do not exist; why a mean is not a future; what a context window costs; the staging that keeps four objectives from fighting; mixture design against real availability; post-training; distillation into a real-time loop; and the evaluation that predicts downstream usefulness rather than flattering it.
How this track is built
An original first-principles track, not a survey of papers. Four things set lessons 1–16 apart. (1) A relay, not a list. Every lesson opens with the failure the previous one left standing and closes by stating the next; that hand-off is one sentence, the baton, which appears word for word at the end of one lesson, at the start of the next and in the table below, and a script diffs them. (2) One world. Everything happens in the Courtyard: an 8 m by 5 m floor, a rolling ball with friction and bouncing walls, a noisy sensor, a curtain it vanishes behind, a goal, and a few obstacles that appear when a lesson needs a new difficulty. Because the world is small and exactly known, every claim is a computation you can watch, and a model's mistake can be compared with the truth. (3) The widgets compute. Each lesson has one widget that runs the mechanism it is about (a filter, a trained network, a planner, a rollout) rather than illustrating it. (4) Every number is recomputed. Each quoted number is tagged, and an independent script re-derives it from scratch; facts about real systems come only from a verified list, are described as of October 2026, and are cited by name and year. Part VI is older. Lessons 17–31, on training a world model, were written before this method, in a different layout: each opens with a box that says what forced it and closes with a hand-off, but those sentences are not diffed against each other and their numbers are not re-derived by script.

The exam

The whole track is graded by one question asked in different ways: if the agent does a, what happens? A model that answers it well enough lets the agent choose actions without trying them, and the measure of "well enough" is always a decision: how often the ball ends up in the goal when the agent acts on the model's answer. Lesson 1 shows why that is the right measure by pricing a trial against a thought; the later lessons are the moves forced by what goes wrong. Part VI turns the question round: what has to be true of the training for the model to pass.

Part I · Why predict, and what to carry (lessons 01–04)

The economics of consequence, the state an agent cannot see, what a learned state should keep, and the filter that produces it.

01
Why predict? The world-model contract
Trial and error versus a model, the real cost of a trial, the oracle planner on the Courtyard, and the four ingredients any model of consequences needs.
02
The state you cannot see: belief
Why the observation is not the state, the Markov property, the Bayes filter, the Kalman filter in closed form, covariance growth behind the curtain, and calibration.
03
What should a state keep?
Sufficiency and minimality, why reconstruction keeps the loudest nuisance, predictive and latent-space objectives, collapse and its cures, value-equivalent states.
04
Learning the filter and the dynamics
Sequential prediction as the objective, a linear predictor that rediscovers the Kalman filter, memory length, recurrent and stochastic latent states, the surprise term.

Part II · Futures you can trust (lessons 05–07)

Distributions instead of points, errors that compound, and the difference between watching and doing.

05
One future is a lie: distributions
Why least squares gives the conditional mean, the knife-edge fork, mixture density heads, aleatoric versus epistemic uncertainty, ensembles, and calibration.
06
Errors compound: rollouts and horizons
Teacher forcing versus free running, the recursion e(k+1) = J·e(k) + δ, the trustworthy horizon, chaos, multi-step training, and closed-loop correction.
07
What does my action cause?
Conditional versus interventional distributions, a confounded log, positivity and coverage, exploration by disagreement, and latent confounders as belief.

Part III · Using the model (lessons 08–10)

Practice inside the model, plan with it at decision time, and think coarser to see farther.

08
Practice in imagination
Dyna and Dreamer in miniature, imagined returns, the optimizer's curse and its sqrt(2 ln K) law, ensembles and pessimism, short rollouts, mixing in real data.
09
Plan with it: search at decision time
Random shooting, the cross-entropy method, receding-horizon control, discrete search, value-equivalent models, and the horizon that balances myopia against model error.
10
Think coarser to see farther: temporal abstraction
Why long tasks outrun the trustworthy horizon, jumpy models, the optimal stride, options and skills, subgoals, and hierarchical planning.

Part IV · Scaling the world (lessons 11–15)

Pixels, places, things, unlabeled actions and real time, and finally a body.

11
Pixels: the world model becomes a video model
Tokenizers and the rate-distortion ceiling, autoregressive and diffusion dynamics, action conditioning, token and compute budgets, and what the window forgets.
12
Places: a world you can leave and return to
Why a finite window forgets, ego-motion versus scene change, spatial maps and keyframes, 3D state, and the revisit test.
13
Things: a world made of objects
Permutation symmetry, binding and identity, pairwise interaction models, compositional generalization, and where objects are the wrong abstraction.
14
Actions without labels, worlds in real time
Inverse dynamics and latent action models, controllability, causal streaming, few-step distillation, latency budgets, and consistency over long interaction.
15
Bodies, contact, and words
Hybrid contact dynamics, why smooth models smear stick and slip, force as a hidden variable, multimodal conditioning, and world models versus policies.

Part V · Knowing it works (lesson 16)

Grade the model by the decisions it supports, not by how it looks.

16
The exam we never stopped taking: evaluation
Visual quality versus decision value, state probes, calibration, intervention tests, closed-loop value, model exploitation, and designing a system by product.

Part VI · Training a world model (lessons 17–31)

What "trained" has to mean for a simulator, the bit budget and the tokenizer ceiling, how the action gets in and where its labels come from, heads that do not average, the quadratic bill for memory, staging and mixture, post-training, distillation into a real-time loop, and the whole bill. These lessons keep the layout they were written in.

17
Two seats, one tuple
A world model and a policy are forward and inverse conditionals of one triple, so they compete for one resource: action-labeled experience. The 16,000:1 price, and the arc for the two series.
18
What "trained" means for a simulator
Controllability, consistency and calibration, inside a latency budget. Why fidelity rises with compute while all three stay flat, and why FVD cannot stand in for any of them.
19
The prediction space is a bit budget
Reverse water-filling over three scene components proves a pixel-weighted loss defunds contact first — because variance and decision-relevance are anti-correlated. Pixels, tokens and latents, priced.
20
The tokenizer is the ceiling
Convert patch size into a physical blur floor in millimetres and check it against the task's clearance. Codebook collapse, the temporal Nyquist limit, and the fixes in cost order.
21
From teacher forcing to rollout
Teacher forcing optimises δ and leaves J unconstrained. Scheduled sampling's mode-averaging bias, k-step rollout's k× bill, and why diffusion forcing gets rollout-grade contraction at teacher-forcing prices.
22
How the handle gets in
Price the action in bits: with ~80% of a frame predictable from history, ignoring the handle is a self-reinforcing optimum. Injection ranked by how hard "ignore" is to express; dropout, guidance, and the interior optimum.
23
The actions do not exist — infer them
An action is encoded in its own consequences, so the inverse problem is the easy one. Inverse dynamics (VPT: 2k hours unlocked 70,000), latent actions, and the stopping rule dG/dL > 1.
24
Spreads, not averages
The conditional mean of two real futures is a state of zero probability sitting in a density minimum — invalid, not imprecise. Head choice by tolerance-versus-branching, priced against N ≤ T/ℓ.
25
Horizon, memory, and the quadratic bill
P(forget) = exp(−C/ḡ) rises exponentially in context while attention costs (C·T)², so there is a value-per-FLOP optimum. Why retrieval is the only mechanism priced independently of elapsed time.
26
The curriculum
Joint training loses 1/(1+c(m−1)) to gradient cancellation; staging loses (1−λ)^(m−1) to forgetting. The crossing point, why replay dominates both, and the discipline of evaluating jointly at every boundary.
27
The mixture is the model
A weight is a request; availability decides. 8% of a 50,000-hour budget on a 350-hour source is 11.4 epochs and a 3.3× discount. Why selecting a mixture on validation loss picks the one with the least handle.
28
Post-training a world model
Curated data buys the time constant; the representation sets the ceiling. Consistency as a rollout reward (and therefore gameable), what is capped and what is not, and the ~90%-of-headroom stopping rule.
29
Distilling into a real-time loop
Under a deadline, quality is q(N)·1[N·ℓ≤T] — so a 0.96 teacher at 450 ms scores zero. Free fixes first, consistency distillation next, and adversarial distillation as a calibration hazard.
30
Evaluating the thing you trained
A planner selects for optimistic error, so the report shows g+δ√(2 ln K) and the robot gets g−δ√(2 ln K). The search optimum, training signals ranked by predictiveness, and the same arithmetic applied to your own checkpoint selection.
31
Capstone: the whole bill
One bimanual assembly spec through all thirteen constraints. Two free day-zero checks determine the architecture; the sampling plan outweighs model size; the grounded corpus caps controllability.

The whole derivation on one page

Read the second column of any row and the fourth column of the row above it: they are the same sentence. That is the linear thinking made visible. Each lesson's problem it inherits is exactly the previous lesson's problem it creates, and the middle column is the single move that resolves the first and manufactures the second. The first row inherits its problem from the last lesson of 3D Vision, from first principles, where a scene that changes can be reconstructed from what has happened but nothing predicts what will; the last row hands its problem to lesson 17, which opens Part VI below.

#The problem it inheritsThe one moveThe problem it creates
Part I · Why predict, and what to carry
01A scene that changes can be reconstructed from what has already happened. An agent that has to act needs the other direction: what will happen next, and what would happen if it did something else. Nothing in a reconstruction answers that, because it has no actions and no future. What must a model carry in its head to answer "if I do this, what happens next?", and why would an agent want one at all?Why predict? The world-model contract: Weigh thinking against trying: a model turns computation into trials, and the contract it must meet is "if I do this, what happens next?".A model that answers "if I push this way, where does the ball end up?" makes thinking cheaper than trying: one push in the real world instead of dozens, provided the model is handed the true state of the ball. An agent is never handed the state. It sees a noisy position, and nothing at all while the ball is behind the curtain. What should the agent carry in its head instead of the state it cannot see?
02A model that answers "if I push this way, where does the ball end up?" makes thinking cheaper than trying: one push in the real world instead of dozens, provided the model is handed the true state of the ball. An agent is never handed the state. It sees a noisy position, and nothing at all while the ball is behind the curtain. What should the agent carry in its head instead of the state it cannot see?The state you cannot see: belief: Carry a distribution over the hidden state, updated by predict-then-correct, instead of the last observation.The belief, a distribution over the hidden state updated by predict-then-correct, solves the curtain problem. But we built it from a model of the ball and a model of the sensor that we wrote down by hand, using state variables we chose. A real agent receives pixels, and nobody tells it which of a million numbers matter. What should an agent keep from what it sees, when the only teacher is the stream itself?
03The belief, a distribution over the hidden state updated by predict-then-correct, solves the curtain problem. But we built it from a model of the ball and a model of the sensor that we wrote down by hand, using state variables we chose. A real agent receives pixels, and nobody tells it which of a million numbers matter. What should an agent keep from what it sees, when the only teacher is the stream itself?What should a state keep?: Keep what predicts what matters: a predictive objective instead of a reconstructive one, and a guard against collapse.A good state keeps what predicts the future that matters and drops the rest; reconstructing the observation is the wrong way to ask for it, and predicting in a latent space needs protection against collapse. We still lack the machine itself: something that turns a stream of observations and actions into that state and moves it forward. How is the filter learned from data?
04A good state keeps what predicts the future that matters and drops the rest; reconstructing the observation is the wrong way to ask for it, and predicting in a latent space needs protection against collapse. We still lack the machine itself: something that turns a stream of observations and actions into that state and moves it forward. How is the filter learned from data?Learning the filter and the dynamics: Train a sequence predictor: its hidden state becomes the belief, and with actions as inputs it becomes the dynamics model.A recurrent state trained to predict the next observation becomes the belief, and with actions as inputs it becomes a dynamics model; in a deterministic world it is almost exact. But at the post the ball goes up or down depending on a difference smaller than the sensor can see, and a model trained to minimize squared error predicts the average, a ball that passes straight through the post. What should a model output when the future is not one point?
Part II · Futures you can trust
05A recurrent state trained to predict the next observation becomes the belief, and with actions as inputs it becomes a dynamics model; in a deterministic world it is almost exact. But at the post the ball goes up or down depending on a difference smaller than the sensor can see, and a model trained to minimize squared error predicts the average, a ball that passes straight through the post. What should a model output when the future is not one point?One future is a lie: distributions: Output a distribution: squared error predicts the mean, so use mixtures, discrete latents or samplers, and check calibration.A model can now output a distribution and draw samples: several futures, each plausible. Using it means feeding each prediction back in as the next input, again and again, though the model was only ever trained on real inputs. A small error becomes the input of the next step. How fast do errors grow when a model predicts from its own predictions, and how far ahead can it be trusted?
06A model can now output a distribution and draw samples: several futures, each plausible. Using it means feeding each prediction back in as the next input, again and again, though the model was only ever trained on real inputs. A small error becomes the input of the next step. How fast do errors grow when a model predicts from its own predictions, and how far ahead can it be trusted?Errors compound: rollouts and horizons: Measure how error grows when a model eats its own predictions, train on rollouts, and re-observe to reset the clock.A rollout has a trustworthy horizon that we can now estimate, and re-observing resets it. But every model so far was fitted to logs that someone generated by acting. If whoever acted reacted to something the model cannot see, such as a gust of wind, the logs confound what an action did with why it was taken. Can a model fitted to logs tell what an action causes?
07A rollout has a trustworthy horizon that we can now estimate, and re-observing resets it. But every model so far was fitted to logs that someone generated by acting. If whoever acted reacted to something the model cannot see, such as a gust of wind, the logs confound what an action did with why it was taken. Can a model fitted to logs tell what an action causes?What does my action cause?: Separate seeing from doing: randomize, adjust for what you can see, put the confounder into the state, and cover the actions you will consider.A model fitted to varied, randomized actions can say what an action causes, at least where the data went. The cheapest use of it is practice: let a policy improve on the model's imagined rollouts instead of in the world. But a policy trained against a model finds the model's flattering mistakes and climbs them. How do we learn from imagined experience without being fooled by it?
Part III · Using the model
08A model fitted to varied, randomized actions can say what an action causes, at least where the data went. The cheapest use of it is practice: let a policy improve on the model's imagined rollouts instead of in the world. But a policy trained against a model finds the model's flattering mistakes and climbs them. How do we learn from imagined experience without being fooled by it?Practice in imagination: Improve a policy on the model's rollouts, pessimistically, because the optimizer will find every flattering error.Practice produces a policy that is only as good as the model was where it practiced. When the goal changes, or the situation is new, the policy has nothing to say. What if the agent used the model at the moment of decision, searching over what to do next, and what does it do about a model that is trustworthy for only a few steps?
09Practice produces a policy that is only as good as the model was where it practiced. When the goal changes, or the situation is new, the policy has nothing to say. What if the agent used the model at the moment of decision, searching over what to do next, and what does it do about a model that is trustworthy for only a few steps?Plan with it: search at decision time: Search over actions with the model at the moment of decision, execute the first step, and plan again.A planner can look past the horizon its model is trustworthy for, because it re-plans and uses the far end of a plan as a direction. But the maze already needed about fifty steps of look-ahead, each paid for at every decision, and the tasks that matter, crossing a building or cooking a meal, take thousands. A better model lowers the error of each step, not the number of steps; the way to look farther at the same price is to take fewer, bigger steps. How can a model predict in jumps, and what does it lose?
10A planner can look past the horizon its model is trustworthy for, because it re-plans and uses the far end of a plan as a direction. But the maze already needed about fifty steps of look-ahead, each paid for at every decision, and the tasks that matter, crossing a building or cooking a meal, take thousands. A better model lowers the error of each step, not the number of steps; the way to look farther at the same price is to take fewer, bigger steps. How can a model predict in jumps, and what does it lose?Think coarser to see farther: temporal abstraction: Predict in jumps: fewer, bigger steps push the horizon out, at the price of detail.The mechanism is complete in a toy where the state is four numbers and the observation is a noisy dot. Real observation is a video: a million numbers per frame, almost all of them irrelevant, and the same ball can be drawn a million ways. Each piece we built, state, dynamics, uncertainty and rollouts, has to survive that. What happens to a world model when the observation is pixels at internet scale?
Part IV · Scaling the world
11The mechanism is complete in a toy where the state is four numbers and the observation is a noisy dot. Real observation is a video: a million numbers per frame, almost all of them irrelevant, and the same ball can be drawn a million ways. Each piece we built, state, dynamics, uncertainty and rollouts, has to survive that. What happens to a world model when the observation is pixels at internet scale?Pixels: the world model becomes a video model: Encode frames into tokens or latents, and predict them with a generative sequence model conditioned on actions.A video model predicts the next frames convincingly, but its state is the last few hundred frames it is allowed to look at. Walk away from the sofa and come back and the sofa has changed. A finite window is a finite memory. What would let a world model stay consistent about places it has stopped looking at?
12A video model predicts the next frames convincingly, but its state is the last few hundred frames it is allowed to look at. Walk away from the sofa and come back and the sofa has changed. A finite window is a finite memory. What would let a world model stay consistent about places it has stopped looking at?Places: a world you can leave and return to: Remember places in a world frame, separate your own motion from the world's change, and read memory by pose.A map remembers where things are, not what they are: it cannot say that this cup is the cup that was on the table a minute ago, or what happens when one thing pushes another. A world made of things needs a state made of things. How should a model represent entities and their interactions?
13A map remembers where things are, not what they are: it cannot say that this cup is the cup that was on the table a minute ago, or what happens when one thing pushes another. A world made of things needs a state made of things. How should a model represent entities and their interactions?Things: a world made of objects: Represent entities and let them interact: slots, identity, and pairwise mechanisms that generalize across counts and orders.Structure inside the model fixes what it can represent, not what it can learn from. The video on the internet shows what happened and not what anyone did to make it happen, and an interactive world must answer a button press now, not after a minute of sampling. Where do actions come from when the data has none, and how fast can a model run?
14Structure inside the model fixes what it can represent, not what it can learn from. The video on the internet shows what happened and not what anyone did to make it happen, and an interactive world must answer a button press now, not after a minute of sampling. Where do actions come from when the data has none, and how fast can a model run?Actions without labels, worlds in real time: Infer the actions from the video, then make the model causal, few-step and fast enough to answer a button press.Actions can be inferred from video and the model can run in real time, for a world seen through a screen. A robot's world is touched, not only seen: contact is discontinuous, force is invisible to a camera, and the goal arrives in words. What changes when the agent has a body, and the model must predict what it will feel?
15Actions can be inferred from video and the model can run in real time, for a world seen through a screen. A robot's world is touched, not only seen: contact is discontinuous, force is invisible to a camera, and the goal arrives in words. What changes when the agent has a body, and the model must predict what it will feel?Bodies, contact, and words: Add the body: proprioception, touch and force, contact modes, and goals given in language.We now have every piece: state, dynamics, uncertainty, rollouts, causality, practice, planning, abstraction, pixels, memory, objects, actions and a body, each justified by an exam the previous model failed. How do we know that a particular model passes the exam that matters, a better decision in the world, and not only the ones that look good?
Part V · Knowing it works
16We now have every piece: state, dynamics, uncertainty, rollouts, causality, practice, planning, abstraction, pixels, memory, objects, actions and a body, each justified by an exam the previous model failed. How do we know that a particular model passes the exam that matters, a better decision in the world, and not only the ones that look good?The exam we never stopped taking: evaluation: Grade the model by the decisions it supports: a ladder of exams, rank reversals, exploitation, and a failure atlas.Everything in this series assumed that a trained world model exists. How is one trained: which loss, which schedule, which data, and what does it cost?
Part VI · Training a world model
Lessons 17–31 keep the layout they were written in: each opens with a box that says what forced it and closes with a hand-off, but those sentences are not diffed against this table and their numbers are not re-derived by script.

Part VI in brief

The first sixteen lessons derived what a world model is and what you do with one, assuming a trained model existed. Part VI discharges that assumption. It is the engineering of the forward map p(o′|o,a), which loss, which schedule, which conditioning, which mixture, which deadline, derived one forced step at a time. These lessons keep the layout they were written in, which is why the table above stops at lesson 16.

The seed this part grows from
A world model and a policy are the forward and inverse conditionals of one triple (o, a, o′). Condition on (o,a) and you get a simulator; condition on (o,g) and you get a policy. So both draw on the same supply, and in that supply the observation column is nearly free while the action column costs roughly four orders of magnitude more. Every technique in this part is an answer to that price.

The one number that frames everything

Meta's V-JEPA 2 was pretrained on over 1,000,000 hours of internet video. Its action-conditioned counterpart, V-JEPA 2-AC, was trained on roughly 62 hours of robot video carrying an end-effector state.

free video : action-grounded video ≈ 1,000,000 : 62 ≈ 16,000 : 1

Read it as a price, not a lament. Four orders of magnitude separate the column you can have from the column you need, and the field's entire method catalogue (inverse dynamics, latent actions, simulation, cross-embodiment pooling, training inside a learned model) is arbitrage across that spread.

The two answers to "what does this training consume the most?"
By token volume: video, and not close. A 256² frame at patch 16 is 256 tokens; at 8 fps that is 7.37 M tokens per hour, about 600× denser per wall-clock minute than speech. By marginal value per dollar: action-grounded, contact-rich, on-policy data. The two rankings are almost exactly inverted, and that inversion is the field's central engineering problem. Lessons 16 to 24 of the robot series make the ledger explicit.

Why this order and no other

Each lesson exists because the previous one's solution created a new failure. That is the only organising principle here.

17 two seats, one tuple → we are training a forward map from expensive tuples 18 but "trained" is undefined → three acceptance tests no pixel metric implies 19 tests need a quantity → the bit budget: a pixel loss starves what decides the task 20 tokens delegate the budget → the tokenizer's blur floor is a permanent ceiling 21 now train the dynamics → teacher forcing hides the only error that matters 22 stable but uncontrolled → the action is worth little likelihood, so it gets dropped 23 conditioning needs labels → and 99.99% of video has none: infer them 24 labels, but futures branch → the conditional mean is a state physics forbids 25 sampling × horizon binds → context is quadratic; permanence needs another mechanism 26 five objectives, one budget→ stage when conflict beats forgetting 27 each stage wants data → up-weighting a scarce source only re-reads it 28 broad training won't obey → post-training moves behaviour, never information 29 obedient but 450 ms slow → under a deadline, the student beats the teacher 30 two models, which is better→ the planner hunts your errors; search has an optimum 31 everything, costed → and the binding constraints were free to check on day zero

How to read Part VI

  1. Run the free checks first. The blur floor (lesson 20) and the per-tick evaluation budget (lessons 18, 24, 29) are arithmetic, cost nothing, and determine your architecture. Every project that skips them rediscovers them expensively.
  2. Never let fidelity be the gate. It is one of four acceptance coordinates and the least important. Keep it as a cheap regression alarm.
  3. Distinguish "behaves wrongly" from "cannot see." The first is a training problem. The second is upstream, and no amount of training touches it.
  4. Treat every mixture weight as an epoch count. Divide by availability before believing the weight.
  5. Assume your consumer is an optimiser. It will find whatever your model got wrong, and the better it is, the better it finds it.
Companion tracks, and the red lines
Lessons 1–16 own the objects Part VI trains: the transition objective (lesson 4), calibrated futures (5), the compounding recurrence (6), causality and intervention (7), model exploitation (8), the hierarchy of rates (10), places and memory (12), objects (13), latent actions as an interface (14). Synthetic Vision owns the sim data-generating process and the closed-loop evaluation harness (11). Generative Continuous owns diffusion and flow mathematics; Distillation owns compression generally; Reinforcement Learning owns policy optimisation. Part VI cites all of them and re-derives none.
How to read this
Straight through, in order: the track is one argument, and each lesson's "Linear position" box names the problem it inherits, the single idea it adds and the problem it hands on. If you have time only for the spine, read 01 → 02 → 04 → 05 → 06 → 08 → 09 → 16, and for training one 17 → 18 → 19 → 20 → 22 → 23 → 31. Each lesson has one widget that runs its mechanism, and a "What to try" paragraph that tells you which control to move first. To re-check every quoted number yourself, run python3 tools/chain/validate_chain.py --series wm --widgets --oracles from the repository root: it re-derives each one in lessons 1–16 with an independent script.

Where this track sits

Before it: 3D Vision, from first principles ends on the hand-off this track begins with. Beside it: Reinforcement Learning · 07 Planning and 19 Offline RL cover the control side that lessons 8–9 build on; Computer Vision · 16 Video and temporal models and Generative Models · 11 VQ tokenizers cover the machinery lesson 11 puts to work. After it: Part VI of this track asks how such a model is trained, with which loss, schedule and data, and what it costs; Robot Model Training then follows the same questions into policies and into the data they eat.