Distilling into a real-time loop
Latency is not an engineering concern to be handled after the science. It is part of the specification, and it changes which model is better — because a model that answers after the deadline has a quality of zero, whatever its samples look like.
1 · Deadlines make quality discontinuous
Here is the situation in one image. The controller asks a question every fifty milliseconds. Your model answers beautifully, in four hundred and fifty. By the time the answer arrives the arm has already moved eight times without it, so the answer is not slightly late — it is about a world that no longer exists. Nobody reads it. A better answer nobody reads is not a better model, and that sentence is the whole lesson.
Ordinarily we treat quality as continuous — more compute, slightly better output. A deadline breaks that. Define the quality the system receives:
Qsys(N) = q(N) · 1[ N·ℓ ≤ T ]That indicator is the whole lesson. Quality rises smoothly with sampling steps right up to the deadline and then falls off a cliff to zero, because a prediction that arrives after the controller needed it is not a worse prediction, it is no prediction. The consequence reorders every comparison you have been making:
And there is a second, subtler gain. With a cheap student you can afford more of the other factors. Lesson 25's bill was (C·T)²·H·N·K. Cutting N from 50 to 4 frees 12× that you can spend on horizon or on candidate plans — so the distilled system is often better on the acceptance tests even ignoring the deadline, because it can search more.
2 · Why few-step sampling is possible at all
It should be surprising that 50 steps compress to 4 without collapse. The reason is that the many-step sampler is not spending its steps on information — it is spending them on path accuracy.
A diffusion or flow sampler numerically integrates a trajectory from noise to data. Fifty steps is a fine-grained integration of that path. But the caller does not want the path; it wants the endpoint. So the question becomes: can a network learn the endpoint map directly, skipping the integration? And it can, because the endpoint map is a well-defined function — the many-step sampler computes it, just slowly. Distillation is fitting a direct approximation to a map you already have an expensive oracle for.
Three families do this, and they differ in what they match:
| family | what the student is trained to match | practical note |
|---|---|---|
| trajectory / progressive | the teacher's endpoint from the same noise; halve the step count repeatedly | stable, needs teacher rollouts, several rounds |
| consistency | self-consistency along the path — any point maps to the same endpoint | enables 1–4 steps; the mainstream choice |
| adversarial | a discriminator's judgement of realism | sharpest samples; silently ruins calibration |
3 · Latency has three other terms, and two are cheaper to fix
Sampling steps get all the attention because they are the largest factor. They are not the only one, and the others need no retraining:
- Cache the context. A rollout re-attends the same history at every step. Persisting the key-value cache turns repeated context encoding into one encode plus incremental updates. Pure engineering, often 2–3×, no quality cost at all. Do this first.
- Batch the candidate plans. The planner's K rollouts are independent, so they are one batched forward pass rather than K sequential ones. On a GPU this is nearly free up to the memory limit — K largely stops being a latency term and becomes a memory term.
- Shrink the model, not just the schedule. Quantisation and a smaller student both reduce ℓ directly, and N ≤ T/ℓ means halving ℓ doubles the steps you can afford. Combining schedule distillation with weight compression is multiplicative.
- Decouple the frequencies. Nothing requires the whole model to run at 50 Hz. Run expensive branching prediction at 2 Hz for planning and a cheap deterministic tracker at 50 Hz — this is the hierarchy lesson 10 derives from the rates the physics and the decision demand, and it is why robot foundation models are built as two systems rather than one.
Note the ordering again: caching and batching before distillation, because they cost nothing in quality. A surprising number of "we need to distil" conclusions are really "we re-encode the context 50 times per step."
Widen the tick to 500 ms and the verdict flips to "the teacher already fits — do not distil." That is the honest boundary of this lesson. Distillation is not universally good; it is the correct response to a binding latency constraint. For an offline data-generation use — where synthetic_vision generates frames and episodes in batch with no deadline — the teacher is simply the better model and distilling it is a pointless quality sacrifice.
4 · What to re-measure after distilling
A distilled student is a different model. The acceptance tests do not transfer, and two of them fail in characteristic ways.
The calibration row is the one to internalise. Compressing a sampler tends to compress its diversity — the student learns the map to the most likely endpoint more reliably than it learns the spread of endpoints. So a distilled model is systematically over-confident, and over-confidence in a simulator is what lesson 30's planner exploits. Measure ECE on the student directly; never inherit the teacher's number.
Where this points next
We now have two models and a genuine question: which is better? That is an evaluation problem, and it contains a trap that gets stronger the better your planner is. A planner searching K candidate plans does not sample your model's errors at random — it seeks them out, because an error that overstates a plan's value looks exactly like a good plan. Lesson 30 shows that this makes measured improvement and delivered improvement diverge, and computes the search budget past which more search makes things worse.
Interview prompts
- Write the system-level quality under a deadline and explain the discontinuity. Q_sys(N) = q(N)·1[N·ℓ ≤ T]; a prediction arriving after the controller needed it is no prediction, so quality drops to zero rather than degrading. (§1)
- Give a second reason distillation improves the acceptance tests beyond meeting the deadline. The per-tick bill is (C·T)²·H·N·K, so cutting N frees budget to spend on horizon or candidate plans — the cheap model can search more. (§1)
- Why is few-step sampling possible without collapse? Many steps buy path accuracy in a numerical integration, not information; the endpoint map is a well-defined function the slow sampler computes, so a network can fit it directly. (§2)
- Which distillation family is dangerous for a planner-facing model, and why? Adversarial — it optimises apparent realism with no term for matching sample frequencies, so it concentrates mass on the likeliest mode and destroys calibration. (§2)
- Name two latency fixes that cost no quality, and say why they come first. Persisting the key-value cache across rollout steps and batching the planner's independent candidate rollouts — both are pure engineering with no quality cost, so they should precede any distillation. (§3)
- When is distilling the wrong decision? When the latency constraint is not binding — for offline batch trajectory generation the teacher simply is the better model, and distilling sacrifices quality for nothing. (§3)
- Why is a distilled model systematically over-confident? Compression preserves the map to the most likely endpoint better than the spread of endpoints, so diversity shrinks and calibration degrades — which is exactly what a planner exploits. (§4)