all_lessons/ World Models/ 29 · distilling into a real-time looplesson 29 / 31

Distilling into a real-time loop

Latency is not an engineering concern to be handled after the science. It is part of the specification, and it changes which model is better — because a model that answers after the deadline has a quality of zero, whatever its samples look like.

Where we are
Lesson 24 established N ≤ T/ℓ and lesson 28 handed us an obedient, calibrated teacher needing 50 sampling steps. At 9 ms per step that is 450 ms per world-model step against a 50 ms tick — a 9× shortfall. This lesson closes it, and the interesting part is that closing it can raise the quality that gets measured.
Forced by 28An obedient, calibrated teacher — at 450 ms per step. This stepDistil the sampling schedule; the student wins under the deadline that actually applies. Forces 30Which of these two models is better is now an evaluation question with a trap in it.

1 · Deadlines make quality discontinuous

Here is the situation in one image. The controller asks a question every fifty milliseconds. Your model answers beautifully, in four hundred and fifty. By the time the answer arrives the arm has already moved eight times without it, so the answer is not slightly late — it is about a world that no longer exists. Nobody reads it. A better answer nobody reads is not a better model, and that sentence is the whole lesson.

Ordinarily we treat quality as continuous — more compute, slightly better output. A deadline breaks that. Define the quality the system receives:

Qsys(N) = q(N) · 1[ N·ℓ ≤ T ]

That indicator is the whole lesson. Quality rises smoothly with sampling steps right up to the deadline and then falls off a cliff to zero, because a prediction that arrives after the controller needed it is not a worse prediction, it is no prediction. The consequence reorders every comparison you have been making:

The reordering
Teacher at 50 steps, 0.96 quality, 450 ms: Qsys = 0. Student at 4 steps, 0.88 quality, 36 ms: Qsys = 0.88. The student is not "an acceptable compromise." Under the specification the caller actually stated, it is infinitely better, and the teacher is a non-functional artifact. Distillation is not a degradation step — it is what makes the model exist as a product.

And there is a second, subtler gain. With a cheap student you can afford more of the other factors. Lesson 25's bill was (C·T)²·H·N·K. Cutting N from 50 to 4 frees 12× that you can spend on horizon or on candidate plans — so the distilled system is often better on the acceptance tests even ignoring the deadline, because it can search more.

2 · Why few-step sampling is possible at all

It should be surprising that 50 steps compress to 4 without collapse. The reason is that the many-step sampler is not spending its steps on information — it is spending them on path accuracy.

A diffusion or flow sampler numerically integrates a trajectory from noise to data. Fifty steps is a fine-grained integration of that path. But the caller does not want the path; it wants the endpoint. So the question becomes: can a network learn the endpoint map directly, skipping the integration? And it can, because the endpoint map is a well-defined function — the many-step sampler computes it, just slowly. Distillation is fitting a direct approximation to a map you already have an expensive oracle for.

Three families do this, and they differ in what they match:

familywhat the student is trained to matchpractical note
trajectory / progressivethe teacher's endpoint from the same noise; halve the step count repeatedlystable, needs teacher rollouts, several rounds
consistencyself-consistency along the path — any point maps to the same endpointenables 1–4 steps; the mainstream choice
adversariala discriminator's judgement of realismsharpest samples; silently ruins calibration
The trap in the third row
Adversarial distillation optimises "looks real," which is lesson 18's test zero. It has no term for whether the sample frequencies match reality, and in practice it concentrates mass on the most-plausible mode — destroying exactly the calibration lesson 24 insisted on. For a generative video demo this is fine. For a model a planner will consume, it converts a calibrated simulator into a confident liar. If you use it, re-measure ECE afterwards, and expect to be unhappy.

3 · Latency has three other terms, and two are cheaper to fix

Sampling steps get all the attention because they are the largest factor. They are not the only one, and the others need no retraining:

  1. Cache the context. A rollout re-attends the same history at every step. Persisting the key-value cache turns repeated context encoding into one encode plus incremental updates. Pure engineering, often 2–3×, no quality cost at all. Do this first.
  2. Batch the candidate plans. The planner's K rollouts are independent, so they are one batched forward pass rather than K sequential ones. On a GPU this is nearly free up to the memory limit — K largely stops being a latency term and becomes a memory term.
  3. Shrink the model, not just the schedule. Quantisation and a smaller student both reduce ℓ directly, and N ≤ T/ℓ means halving ℓ doubles the steps you can afford. Combining schedule distillation with weight compression is multiplicative.
  4. Decouple the frequencies. Nothing requires the whole model to run at 50 Hz. Run expensive branching prediction at 2 Hz for planning and a cheap deterministic tracker at 50 Hz — this is the hierarchy lesson 10 derives from the rates the physics and the decision demand, and it is why robot foundation models are built as two systems rather than one.

Note the ordering again: caching and batching before distillation, because they cost nothing in quality. A surprising number of "we need to distil" conclusions are really "we re-encode the context 50 times per step."

Teacher versus student, inside the tick
Both curves are quality against latency, not against steps — so the teacher's slow, high ceiling and the student's fast, slightly lower ceiling can be compared fairly. The red line is the tick. Read the two readouts at the line: that comparison, not peak quality, is the one that decides deployment.
Interactive teacher-versus-student latency comparison.
steps inside the tick
—
teacher quality
—
student quality
—
loop rate
—
decision
—

Widen the tick to 500 ms and the verdict flips to "the teacher already fits — do not distil." That is the honest boundary of this lesson. Distillation is not universally good; it is the correct response to a binding latency constraint. For an offline data-generation use — where synthetic_vision generates frames and episodes in batch with no deadline — the teacher is simply the better model and distilling it is a pointless quality sacrifice.

4 · What to re-measure after distilling

A distilled student is a different model. The acceptance tests do not transfer, and two of them fail in characteristic ways.

Pixel fidelity — usually near-preserved; the easy onelow risk
Controllability — re-run paired intervention and re-sweep guidancemoderate risk
Consistency over the full horizon — few-step samplers drift differentlyhigh risk
Calibration — the first casualty, especially adversariallyhighest risk
Latency at the deployed step count, measured on target hardwarethe whole point

The calibration row is the one to internalise. Compressing a sampler tends to compress its diversity — the student learns the map to the most likely endpoint more reliably than it learns the spread of endpoints. So a distilled model is systematically over-confident, and over-confidence in a simulator is what lesson 30's planner exploits. Measure ECE on the student directly; never inherit the teacher's number.

Takeaway
Under a deadline the system's quality is q(N)·1[N·ℓ ≤ T], so a 0.96-quality teacher at 450 ms scores zero against a 50 ms tick while a 0.88-quality student at 36 ms scores 0.88 — distillation is what makes the model exist, not a degradation. It works because a many-step sampler spends its steps on path accuracy, not information, and the endpoint map can be fitted directly. Do the free things first (KV caching, batching candidates), prefer consistency-style distillation, and treat adversarial distillation as a calibration hazard. Re-measure calibration and consistency on the student — a compressed sampler is systematically over-confident.

Where this points next

We now have two models and a genuine question: which is better? That is an evaluation problem, and it contains a trap that gets stronger the better your planner is. A planner searching K candidate plans does not sample your model's errors at random — it seeks them out, because an error that overstates a plan's value looks exactly like a good plan. Lesson 30 shows that this makes measured improvement and delivered improvement diverge, and computes the search budget past which more search makes things worse.

Interview prompts