Post-training a world model
The pretrained model has coverage and no obedience. Post-training is the pass that fixes that, and it is extraordinarily effective — until it hits a wall it cannot climb, at which point continuing to push is the most expensive way to make no progress in this entire part.
1 · Three post-training moves, three different targets
Two failures look identical from the outside and need opposite responses. In the first, the model can see the pin is misaligned and moves the wrong way — it behaves badly on information it has. In the second, the pin is misaligned by one millimetre and the model's input cannot represent one millimetre at all — it cannot see the thing it is being asked about. Post-training fixes the first completely and the second not at all, so the first job is telling them apart.
Move 2 deserves a note, because it is the one that is structurally different. Everything before this lesson optimised a per-frame or per-transition loss. Consistency is not a property of any frame; it is a property of a trajectory. So it cannot be a supervised target — there is no ground-truth "consistent" label — and it has to be a reward computed over a generated rollout. The gradient then comes back through the sampling process, which is expensive and is why this move is a post-training pass rather than part of the main objective.
2 · The ceiling, and how to recognise it from the curve
Post-training changes which function the network computes within its representational reach. It cannot change the reach. Lesson 20 established the reach quantitatively: if the tokenizer's blur floor is 25 mm and the task's clearance is 2 mm, the distinction between success and failure is not in the representation. Post-training on ten thousand careful insertion demonstrations cannot teach a distinction the input does not encode.
The signature is diagnostic and you can read it off the learning curve. Capability rises, decelerates, and flattens at a level well below 1 — and continues to flatten no matter how much curated data you add:
capability(h) ≈ ceiling · (1 − e−h/h₀)Two things are worth extracting. The ceiling is set by the representation, not the data. The time constant h₀ is what your data buys. Confusing them is the mistake: a project sees the curve flatten, concludes "we need more data," adds 10× more, moves from 92% to 96% of the ceiling, and never asks what the ceiling is.
The marginal return makes the stopping rule mechanical:
d capability / dh = (ceiling / h₀) · e−h/h₀Once you have closed roughly 90% of the headroom, the derivative has fallen by an order of magnitude from its initial value. That is the signal to stop buying curated data and go re-examine the representation.
Drag the ceiling down to 0.3 and out to 0.95 while leaving the curated hours fixed. The curve's shape is identical in both cases — same knee, same saturation — and only its height differs. That is the uncomfortable part: the learning curve alone cannot tell you whether you are data-limited or representation-limited, so the diagnosis has to come from §3's arithmetic rather than from the plot.
The two curves are the whole diagnostic. Semantics — "the model understands that this is a kitchen and that mugs go on shelves" — genuinely does improve with curated data almost without limit, because the representation already encodes it. Contact control does not, because it does not. So the useful question is never "is post-training working?" but "which of my capabilities is capped, and by what?"
3 · What is capped, what is not
That last line is the whole lesson in seven words. If the failure is that the model behaves wrongly given what it can see, post-training fixes it. If the failure is that the model cannot see the distinguishing feature, post-training is the wrong instrument and there are only three real remedies — all of which mean going back upstream:
- Change the optics or add a sensor. Lesson 20 §3's ranking, still the cheapest option by a wide margin, and still the one people try last.
- Retrain the representation and redo stages 2–4. Expensive and correct. Budget for it as a possibility from the start rather than discovering it as a crisis.
- Change the task decomposition. If the model cannot resolve 2 mm, do not ask it to. Let a visual-servoing or force-controlled primitive own the last few millimetres and let the world model own the approach. Frequently the best engineering answer, and it costs no retraining at all.
4 · Two failure modes specific to post-training
Diversity collapse. The curated set is small and deliberately narrow. Fine-tuning hard on it can destroy the coverage that a million hours of broad video bought — the model becomes excellent on the post-training distribution and worse everywhere else. The fixes are lesson 26's: mix a slice of pretraining data into every post-training batch, and keep the learning rate low enough that you are adjusting rather than overwriting. The measurement is a broad held-out set evaluated before and after; if nobody looks, nobody notices.
Controllability that only exists at the guidance scale you tuned. Lesson 22's guidance optimum is a property of the model, and post-training moves it. A model post-trained for obedience may need a much lower s, and running it at the old value now over-steers into artifacts. Re-sweep the guidance scale after every post-training pass; it takes minutes and it is a common source of "the fine-tune made it worse."
Where this points next
We now have a model that is accurate, stable, obedient, consistent and calibrated. It answers in 450 milliseconds. Lesson 18's contract requires 50, so as an interactive environment it does not exist. Lesson 29 buys the order of magnitude back, and shows why the distilled student can be strictly better than its teacher at the only quality that is measured under the deadline.
Interview prompts
- Why must consistency be trained as a reward rather than a supervised target? Consistency is a property of a whole trajectory, not of any frame, so there is no per-frame ground-truth label; it must be a functional scored on generated rollouts. (§1)
- What should a controllability post-training set be curated for? Action diversity — the same action in varied contexts and varied actions in the same context — so the conditional mutual information per frame is high, not visual quality. (§1)
- Separate the two parameters of the post-training curve and say what each depends on. In ceiling·(1−e^(−h/h₀)), the ceiling is set by the representation and the time constant h₀ is what curated data buys. (§2)
- Give the stopping rule for buying curated post-training data. Once about 90% of the headroom is closed the marginal return has fallen an order of magnitude, so stop and re-examine whether the ceiling is the binding constraint. (§2)
- Name a capability post-training genuinely cannot improve, and why. Spatial precision finer than the tokenizer's blur floor, or forces absent from the input — the distinguishing information is not in the representation, and post-training moves behaviour, not information. (§3)
- A model cannot resolve a 2 mm clearance. Give the remedy that requires no retraining. Change the task decomposition: let a force-controlled or visual-servoing primitive own the final millimetres while the world model owns the approach. (§3)
- Why can a consistency reward make rollouts worse in a way the reward cannot see? It is an optimisable proxy, so the model finds cheap ways to score — for example generating fewer, larger, unambiguous objects to satisfy a re-detection-based permanence score. (§1)
- Two things to check immediately after any post-training pass. A broad held-out set before and after, to catch diversity collapse; and a re-sweep of the guidance scale, since post-training moves its optimum. (§4)