all_lessons/ World Models/ 28 · post-traininglesson 28 / 31

Post-training a world model

The pretrained model has coverage and no obedience. Post-training is the pass that fixes that, and it is extraordinarily effective — until it hits a wall it cannot climb, at which point continuing to push is the most expensive way to make no progress in this entire part.

Where we are
Lesson 27's mixture bought coverage from a million hours of broad video, and the anti-correlation it derived means that model listens poorly. So we do a final pass on a small curated set. The question worth being precise about is which failures that pass can fix, because two of the three categories it is routinely aimed at are immovable.
Forced by 27Broad pretraining maximises coverage and minimises the handle. This stepPost-train for control and consistency — and locate the ceiling that no post-training crosses. Forces 29An obedient model at 50 sampling steps still cannot close a control loop.

1 · Three post-training moves, three different targets

Two failures look identical from the outside and need opposite responses. In the first, the model can see the pin is misaligned and moves the wrong way — it behaves badly on information it has. In the second, the pin is misaligned by one millimetre and the model's input cannot represent one millimetre at all — it cannot see the thing it is being asked about. Post-training fixes the first completely and the second not at all, so the first job is telling them apart.

1 · curated supervised
Controllability fine-tune
A small set chosen for action diversity, not visual quality — the same actions in varied contexts and varied actions in the same context, so I(o′;a|o) is high per frame. Lesson 22's data fix, applied concentratedly.
2 · reward on rollouts
Consistency fine-tune
Score the model's own rollouts on properties no single-frame loss expresses — object permanence across a revisit, physical plausibility, no teleporting — then optimise that score. The reward is a rollout functional.
3 · real-world grounding
Sim-trained model → reality
Correct the sim-share discount from lesson 27 using a small real set. Not more data — corrective data, aimed at the specific gaps sim got wrong.

Move 2 deserves a note, because it is the one that is structurally different. Everything before this lesson optimised a per-frame or per-transition loss. Consistency is not a property of any frame; it is a property of a trajectory. So it cannot be a supervised target — there is no ground-truth "consistent" label — and it has to be a reward computed over a generated rollout. The gradient then comes back through the sampling process, which is expensive and is why this move is a post-training pass rather than part of the main objective.

And it is a reward, so it will be gamed
A rollout reward is an objective the model can now optimise, which means lesson 8's warning applies directly to training: whatever proxy you use for "physically plausible," the model will find the cheapest way to score well on it. A permanence reward computed by object re-detection can be maximised by generating scenes with fewer, larger, unambiguous objects. Hold out a locked real evaluation set that the reward model never touched — the same discipline as lesson 23's circularity warning.

2 · The ceiling, and how to recognise it from the curve

Post-training changes which function the network computes within its representational reach. It cannot change the reach. Lesson 20 established the reach quantitatively: if the tokenizer's blur floor is 25 mm and the task's clearance is 2 mm, the distinction between success and failure is not in the representation. Post-training on ten thousand careful insertion demonstrations cannot teach a distinction the input does not encode.

The signature is diagnostic and you can read it off the learning curve. Capability rises, decelerates, and flattens at a level well below 1 — and continues to flatten no matter how much curated data you add:

capability(h) ≈ ceiling · (1 − e−h/h₀)

Two things are worth extracting. The ceiling is set by the representation, not the data. The time constant h₀ is what your data buys. Confusing them is the mistake: a project sees the curve flatten, concludes "we need more data," adds 10× more, moves from 92% to 96% of the ceiling, and never asks what the ceiling is.

The marginal return makes the stopping rule mechanical:

d capability / dh = (ceiling / h₀) · e−h/h₀

Once you have closed roughly 90% of the headroom, the derivative has fallen by an order of magnitude from its initial value. That is the signal to stop buying curated data and go re-examine the representation.

Post-training against a representational ceiling
Violet is contact control, capped by the tokenizer ceiling; dashed green is a semantic capability with no such cap. Lower the ceiling to simulate a coarser tokenizer and note that the shape is unchanged — which is exactly why the curve alone cannot tell you whether you are data-limited or representation-limited.
Interactive post-training saturation against a ceiling.
capability
—
ceiling
—
headroom closed
—
marginal return
—
decision
—

Drag the ceiling down to 0.3 and out to 0.95 while leaving the curated hours fixed. The curve's shape is identical in both cases — same knee, same saturation — and only its height differs. That is the uncomfortable part: the learning curve alone cannot tell you whether you are data-limited or representation-limited, so the diagnosis has to come from §3's arithmetic rather than from the plot.

The two curves are the whole diagnostic. Semantics — "the model understands that this is a kitchen and that mugs go on shelves" — genuinely does improve with curated data almost without limit, because the representation already encodes it. Contact control does not, because it does not. So the useful question is never "is post-training working?" but "which of my capabilities is capped, and by what?"

3 · What is capped, what is not

Semantic knowledge — object identity, affordance, plausible arrangementsnot capped
Stylistic and distributional match to your domainnot capped
Obedience to the action channel, given the channel existsnot capped
Trajectory-level consistency, given sufficient contextpartly capped by context (08)
Spatial precision finer than the blur floorcapped by 03
Events shorter than the temporal sampling limitcapped by 03
Forces and contact state, when absent from the inputcapped, absolutely
Rulepost-training moves behaviour, never information

That last line is the whole lesson in seven words. If the failure is that the model behaves wrongly given what it can see, post-training fixes it. If the failure is that the model cannot see the distinguishing feature, post-training is the wrong instrument and there are only three real remedies — all of which mean going back upstream:

  1. Change the optics or add a sensor. Lesson 20 §3's ranking, still the cheapest option by a wide margin, and still the one people try last.
  2. Retrain the representation and redo stages 2–4. Expensive and correct. Budget for it as a possibility from the start rather than discovering it as a crisis.
  3. Change the task decomposition. If the model cannot resolve 2 mm, do not ask it to. Let a visual-servoing or force-controlled primitive own the last few millimetres and let the world model own the approach. Frequently the best engineering answer, and it costs no retraining at all.

4 · Two failure modes specific to post-training

Diversity collapse. The curated set is small and deliberately narrow. Fine-tuning hard on it can destroy the coverage that a million hours of broad video bought — the model becomes excellent on the post-training distribution and worse everywhere else. The fixes are lesson 26's: mix a slice of pretraining data into every post-training batch, and keep the learning rate low enough that you are adjusting rather than overwriting. The measurement is a broad held-out set evaluated before and after; if nobody looks, nobody notices.

Controllability that only exists at the guidance scale you tuned. Lesson 22's guidance optimum is a property of the model, and post-training moves it. A model post-trained for obedience may need a much lower s, and running it at the old value now over-steers into artifacts. Re-sweep the guidance scale after every post-training pass; it takes minutes and it is a common source of "the fine-tune made it worse."

Takeaway
Post-training has three distinct jobs: curated action-diverse data for controllability, a rollout-level reward for consistency (which is a trajectory property and therefore cannot be a supervised target — and which will be gamed), and small corrective real data to pay down the sim discount. Its curve is ceiling·(1−e^(−h/h₀)), where data buys h₀ and the representation sets the ceiling — so when ~90% of headroom is closed, stop buying data and go re-examine lesson 20. Post-training moves behaviour, never information; watch for diversity collapse and re-sweep the guidance scale afterwards.

Where this points next

We now have a model that is accurate, stable, obedient, consistent and calibrated. It answers in 450 milliseconds. Lesson 18's contract requires 50, so as an interactive environment it does not exist. Lesson 29 buys the order of magnitude back, and shows why the distilled student can be strictly better than its teacher at the only quality that is measured under the deadline.

Interview prompts