Spreads, not averages
Knowing the action exactly does not make the future unique. And when two futures are both real, the squared-error answer is not a compromise between them — it is a third state that physics forbids. That distinction, imprecise versus invalid, changes which head you can use.
1 · The mean of two valid futures is an invalid future
The standard objection to MSE is that it produces blur. True, but too weak — blur sounds like a cosmetic defect that a sharper loss would fix. The real problem is categorical.
Squared error has one minimiser, and it is the conditional mean:
argminŷ E[ ‖ŷ − y‖² | x ] = E[ y | x ]Now suppose the mug, nudged at its balance point, tips left half the time and right half the time. Two futures, each entirely real, each occurring 50% of the time. The conditional mean is the mug upright — a state that occurs with probability zero. It is not a smeared version of the answer. It is a fourth possibility the world excludes.
Push the point: the conditional mean lands in a region of the future distribution where the density is a minimum. The optimiser is aiming, with perfect fidelity to its objective, at the least likely place. And every downstream consumer inherits the error. A planner evaluating "will the mug tip?" is told "no" — the model predicted upright — so it selects a plan whose entire safety rests on an outcome that never happens.
2 · The heads, and what each one is really claiming
Every output head is a claim about the shape of p(o′|o,a). Choose the head whose claim is true of your data.
| head | implicit claim | cost per step | fails when |
|---|---|---|---|
| deterministic (MSE) | unimodal, and you only need its centre | 1 forward pass | the future branches at all |
| Gaussian / mixture | a few modes, known count | 1 pass | mode count varies with the situation |
| discrete categorical | the space was quantised (lesson 20) | 1 pass per token | the codebook's ceiling binds |
| diffusion / flow | arbitrary shape, no mode count needed | N passes | you cannot afford N |
Notice the discrete categorical row, because it is the quiet winner in many systems. If lesson 20 already quantised the world into tokens, a softmax over the codebook is a fully general distribution over next tokens at the cost of one forward pass. Multimodality is free — the softmax simply puts mass on two codes. You paid for this at tokenizer time; use it. The reason people reach for diffusion anyway is that the joint distribution over many tokens is not captured by independent per-token softmaxes, and coherence across a frame is exactly what matters.
3 · Price the sampling against the loop
Lesson 18 §6 established the constraint. A generative head needing N sampling steps at ℓ ms, inside a control tick T:
N ≤ T / ℓWith T = 50 ms and ℓ = 9 ms, that is N ≤ 5. Five sampling steps. A diffusion head that needs 50 steps to produce a coherent frame is, at this stage of training, not deployable at all — and no amount of quality justifies a model that cannot answer before the next control tick arrives.
This is where the honest options narrow to three, and it is worth being blunt about them:
- Use a one-pass distributional head. Categorical over tokens, or a mixture. Coherence is weaker; the loop is met. Often the right answer, and under-used because it is unfashionable.
- Use a generative head and distil it. Train at 50 steps, deploy at 1–4. This is lesson 29, and it is the mainstream answer for interactive systems.
- Move the branching out of the fast loop. Sample futures at 2 Hz for planning, run a cheap deterministic model at 50 Hz for tracking. The hierarchy that lesson 10 derives from the rates the physics and the decision demand.
Raise the tolerance until it exceeds the mode separation and the MSE readout jumps from 0.00 to 1.00. That discontinuity is the real content of this lesson: whether you need a generative head is not a question about your model, it is a question about whether your task's tolerance is smaller than your world's branching. Insertion with 2 mm clearance and a mug that can tip either way: tolerance is tiny, branching is large, you need samples. Predicting whether a room is bright in 5 seconds: tolerance is huge, take the mean and move on.
4 · Calibration is a separate obligation
A sampling head gives you a distribution. Lesson 18's third test asks whether it is the right one. These come apart in a specific, common way.
Diffusion and flow heads are trained by matching a denoising or velocity target, not by maximising likelihood, so nothing in the objective enforces that the sample frequencies match reality. A model can produce two sharp, plausible, well-separated modes — and put 90% of its mass on the one that occurs 40% of the time. Every sample looks great. The planner is systematically misled, and worse, it is misled confidently.
So measure it directly, per rollout: bin predicted probabilities of a discrete downstream event (did the mug tip? did the insertion seat?), compare to observed frequencies, report the gap. lesson 5 derives what calibration means for branch distributions; the training-side instruction is simply that sharpness is not calibration and a sharp miscalibrated model is more dangerous than a blurry one, because it invites trust.
Where this points next
Sampling steps multiply with horizon length, and horizon length interacts with something we have not yet priced: how much of the past the model attends to. Context is what makes objects persist when they leave the frame — and it is quadratic. Lesson 25 puts the permanence benefit and the attention cost on the same axes and finds the context length that actually maximises value per FLOP.
Interview prompts
- Why is "MSE causes blur" an understatement? The minimiser is the conditional mean, which for a branching future is a zero-probability state in a density minimum — an outcome the world excludes, not a smeared version of a real one. (§1)
- How does mode-averaging cause drift? The invalid mean state becomes the next input, a configuration the model has never seen, so its next prediction is unconstrained and lesson 21's recurrence compounds from off-manifold. (§1)
- When is a one-pass categorical head fully sufficient for multimodality? When the space is already quantised, since a softmax over the codebook can place mass on two codes at no extra cost — though it does not capture joint coherence across many tokens. (§2)
- Compute the sampling steps available in a 50 ms tick at 9 ms per step, and state the consequence. N ≤ 5, so a 50-step diffusion head is undeployable and must be distilled, replaced with a one-pass head, or moved out of the fast loop. (§3)
- What single comparison decides whether you need a generative head? Task tolerance versus the world's branching scale: if tolerance exceeds the mode separation, the mean is acceptable; if not, sampling is mandatory. (§3)
- Why can a diffusion head be sharp and badly calibrated? It is trained on a denoising or velocity target rather than likelihood, so nothing enforces that sample frequencies match real frequencies. (§4)
- Why is a sharp miscalibrated model more dangerous than a blurry one? Its samples look correct so it invites trust, and the planner is then confidently misled about the probability of the outcomes it depends on. (§4)