all_lessons/ World Models/ 24 · spreads, not averageslesson 24 / 31

Spreads, not averages

Knowing the action exactly does not make the future unique. And when two futures are both real, the squared-error answer is not a compromise between them — it is a third state that physics forbids. That distinction, imprecise versus invalid, changes which head you can use.

Where we are
Lesson 23 gave us an action column even where none was recorded. lesson 5 already derived why futures branch and what a calibrated branch distribution is. This lesson takes that as settled and asks the training-side question: which output head, and can you afford it inside the loop from lesson 18 §6?
Forced by 23Actions are available, even for unlabeled video. This stepChoose the output head: the conditional mean is an invalid state, so sample — and price the sampling. Forces 25Sampling multiplies with horizon, so the horizon budget is now the binding constraint.

1 · The mean of two valid futures is an invalid future

The standard objection to MSE is that it produces blur. True, but too weak — blur sounds like a cosmetic defect that a sharper loss would fix. The real problem is categorical.

Squared error has one minimiser, and it is the conditional mean:

argminŷ E[ ‖ŷ − y‖² | x ] = E[ y | x ]

Now suppose the mug, nudged at its balance point, tips left half the time and right half the time. Two futures, each entirely real, each occurring 50% of the time. The conditional mean is the mug upright — a state that occurs with probability zero. It is not a smeared version of the answer. It is a fourth possibility the world excludes.

Push the point: the conditional mean lands in a region of the future distribution where the density is a minimum. The optimiser is aiming, with perfect fidelity to its objective, at the least likely place. And every downstream consumer inherits the error. A planner evaluating "will the mug tip?" is told "no" — the model predicted upright — so it selects a plan whose entire safety rests on an outcome that never happens.

Why this is worse in a rollout than in a single prediction
The invalid mean state becomes the next input. The model has never seen a mug balanced in that impossible configuration, so its next prediction is not merely wrong but unconstrained — and lesson 21's recurrence now compounds an error that started off the data manifold entirely. Mode-averaging and drift are not two problems; the first is a generator for the second.

2 · The heads, and what each one is really claiming

Every output head is a claim about the shape of p(o′|o,a). Choose the head whose claim is true of your data.

headimplicit claimcost per stepfails when
deterministic (MSE)unimodal, and you only need its centre1 forward passthe future branches at all
Gaussian / mixturea few modes, known count1 passmode count varies with the situation
discrete categoricalthe space was quantised (lesson 20)1 pass per tokenthe codebook's ceiling binds
diffusion / flowarbitrary shape, no mode count neededN passesyou cannot afford N

Notice the discrete categorical row, because it is the quiet winner in many systems. If lesson 20 already quantised the world into tokens, a softmax over the codebook is a fully general distribution over next tokens at the cost of one forward pass. Multimodality is free — the softmax simply puts mass on two codes. You paid for this at tokenizer time; use it. The reason people reach for diffusion anyway is that the joint distribution over many tokens is not captured by independent per-token softmaxes, and coherence across a frame is exactly what matters.

3 · Price the sampling against the loop

Lesson 18 §6 established the constraint. A generative head needing N sampling steps at ℓ ms, inside a control tick T:

N ≤ T / ℓ

With T = 50 ms and ℓ = 9 ms, that is N ≤ 5. Five sampling steps. A diffusion head that needs 50 steps to produce a coherent frame is, at this stage of training, not deployable at all — and no amount of quality justifies a model that cannot answer before the next control tick arrives.

This is where the honest options narrow to three, and it is worth being blunt about them:

  1. Use a one-pass distributional head. Categorical over tokens, or a mixture. Coherence is weaker; the loop is met. Often the right answer, and under-used because it is unfashionable.
  2. Use a generative head and distil it. Train at 50 steps, deploy at 1–4. This is lesson 29, and it is the mainstream answer for interactive systems.
  3. Move the branching out of the fast loop. Sample futures at 2 Hz for planning, run a cheap deterministic model at 50 Hz for tracking. The hierarchy that lesson 10 derives from the rates the physics and the decision demand.
P(valid future) against the control tick
Violet is the probability a sampled future is a real one, rising with sampling steps; dashed cyan is throughput; the red line is the hard step limit your tick allows. The amber line is the MSE head — flat, and zero whenever the modes are separated by more than the tolerance.
Interactive sampling budget against validity.
MSE head P(valid)
—
generative P(valid)
—
steps the tick allows
—
achieved rate
—
diagnosis
—

Raise the tolerance until it exceeds the mode separation and the MSE readout jumps from 0.00 to 1.00. That discontinuity is the real content of this lesson: whether you need a generative head is not a question about your model, it is a question about whether your task's tolerance is smaller than your world's branching. Insertion with 2 mm clearance and a mug that can tip either way: tolerance is tiny, branching is large, you need samples. Predicting whether a room is bright in 5 seconds: tolerance is huge, take the mean and move on.

4 · Calibration is a separate obligation

A sampling head gives you a distribution. Lesson 18's third test asks whether it is the right one. These come apart in a specific, common way.

Diffusion and flow heads are trained by matching a denoising or velocity target, not by maximising likelihood, so nothing in the objective enforces that the sample frequencies match reality. A model can produce two sharp, plausible, well-separated modes — and put 90% of its mass on the one that occurs 40% of the time. Every sample looks great. The planner is systematically misled, and worse, it is misled confidently.

So measure it directly, per rollout: bin predicted probabilities of a discrete downstream event (did the mug tip? did the insertion seat?), compare to observed frequencies, report the gap. lesson 5 derives what calibration means for branch distributions; the training-side instruction is simply that sharpness is not calibration and a sharp miscalibrated model is more dangerous than a blurry one, because it invites trust.

Takeaway
Squared error's minimiser is the conditional mean, which for a branching future is a state of zero probability sitting in a density minimum — invalid, not merely imprecise — and because it becomes the next input it also seeds the drift of lesson 21. Pick the head by whether the task's tolerance is smaller than the world's branching: below tolerance, MSE is fine; above it, you need a distribution. Prefer a one-pass categorical head when the space is already quantised, and remember the hard constraint N ≤ T/ℓ — a 50 ms tick at 9 ms per step allows five sampling steps, which is why distillation (lesson 29) exists. Finally, sharpness is not calibration; measure the latter separately.

Where this points next

Sampling steps multiply with horizon length, and horizon length interacts with something we have not yet priced: how much of the past the model attends to. Context is what makes objects persist when they leave the frame — and it is quadratic. Lesson 25 puts the permanence benefit and the attention cost on the same axes and finds the context length that actually maximises value per FLOP.

Interview prompts