all_lessons/ World Models/ 21 · teacher forcing to rolloutlesson 21 / 31

From teacher forcing to rollout

One line of the training loop decides whether your model survives deployment: at step k, does it consume the recorded token or its own last guess? The convenient choice is also the one that hides the only error a planner will ever meet — and there is now a way to avoid it without paying for it.

Where we are
Lesson 20 fixed the representation and checked its ceiling. Now we train the dynamics on top of it. lesson 6 already derived what happens when errors feed forward — the recurrence ek+1≈J·ek+δ. This lesson takes that recurrence as given and asks the training question it implies: which schedule buys a smaller J, and what does each one cost per step?
Forced by 20A representation with a known ceiling, ready for a dynamics model. This stepPick the input schedule: teacher forcing, scheduled sampling, k-step rollout, diffusion forcing. Forces 22Robust rollouts are worthless if the action never influences them.

1 · The exam with the answer key open

Teacher forcing is the default because it is fast. Every position in the sequence is supervised in parallel, each conditioned on the true prefix, so one forward pass trains the whole horizon at once. It is also, viewed as an exam, a student solving each problem with the previous answer already filled in correctly.

Deployment is the same exam with the key removed and each answer feeding the next. The distributions the model is asked to handle are simply different objects:

train:   ẑk+1 = Fθ(zktrue, ak)      deploy:   ẑk+1 = Fθ(ẑk, ak)

This is exposure bias, and here is the part that catches people: it is not merely that the model is untested on its own outputs. It is that teacher forcing gives the model no reason to be contractive. Recall the two pieces of the recurrence — old error amplified by the Jacobian J, plus fresh error δ. Teacher forcing only ever penalises δ, because ek is reset to zero at every position. The quantity that decides whether a rollout survives is invisible to the objective.

The asymmetry in one line
Teacher forcing optimises δ and leaves J to chance. Rollout training optimises both. Two models with identical one-step loss can differ by orders of magnitude at horizon 20 purely through J — which is why one-step validation curves are almost uninformative about the thing you are building.

2 · The four schedules

Each schedule is a different answer to "how much of its own output does the model eat during training," and each has a cost, because own-output means sequential.

schedulewhat the model consumeseffect on Jtraining cost
teacher forcingalways the true prefixnone — unconstrained1× (fully parallel)
scheduled samplingown sample with probability p, ramped up over trainingpartial, ∝ p≈1× (still one pass; mixed inputs)
k-step rolloutits own outputs for k steps, then re-anchorfull, up to kk× (sequential unroll)
diffusion forcinga prefix at independent per-token noise levelsnear-full≈1× (still parallel)

The first three are the classical ladder: pay compute, buy robustness. The fourth is the interesting one, so let us see why it escapes the trade.

3 · Why independent noise levels buy rollout robustness for free

The trick is to stop thinking of the prefix as "true or self-generated" and think of it as "noisy to some degree." Corrupt each token in the context with its own independently drawn noise level σk, and train the model to denoise the next token given that ragged context.

Now count what the model has been taught. A single training example presents some tokens nearly clean and others heavily corrupted — in every combination, because the levels are drawn independently. Two consequences follow directly:

And crucially the levels are drawn per-token, not per-sequence, so nothing is sequential and the whole context still trains in one parallel pass. One extra benefit falls out: because any noise-level pattern is valid, the same trained model rolls out to any horizon, including horizons longer than any training sequence, without a separate schedule.

Four schedules, same δ, very different rollouts
All four iterate the derived recurrence ek+1=Jeff·ek+δ with the same one-step error δ; only Jeff differs by what the schedule taught. Watch the cost readout: k-step rollout and diffusion forcing land near each other on error, and nowhere near each other on price.
Interactive comparison of four training schedules.
teacher forcingscheduled samplingk-step rolloutdiffusion forcing
e(H) teacher
—
e(H) scheduled
—
e(H) rollout
—
e(H) diffusion
—
training cost
—
diagnosis
—

Set δ to its minimum and note that the curves still separate. That is the lesson's whole point: the schedules do not differ in one-step accuracy, they differ in J. You cannot buy your way out of a bad schedule with a better one-step fit.

4 · The trap inside scheduled sampling

Scheduled sampling looks like a free lunch: same cost, partial robustness. It has a specific bias, and in a world model the bias matters more than it does in language.

Suppose the true future branches — the mug can tip left or right. The model samples a prefix in which it tipped left. Scheduled sampling then trains it toward the recorded continuation, in which the mug tipped right. The gradient says: "given the mug is falling left, predict that it lands right." That target is not merely unhelpful, it is physically incoherent, and repeated exposure teaches the model to hedge — to produce the blurry average of incompatible worlds, which is exactly what lesson 24 is about avoiding.

The rule that follows
Scheduled sampling is safe in proportion to how unimodal your dynamics are. Deterministic, smooth, contact-free motion: fine. Branching, contact-rich, multimodal: it actively teaches mode-averaging. Diffusion forcing does not have this problem, because its target is always the true token and the corruption is on the input, never a mismatched output.

5 · What to log so you can see any of this

One-step validation loss will not show you a schedule problem. Three cheap diagnostics will.

  1. The horizon curve. Free-running error against k, for k out to 2–4× your deployment horizon. Not a number — the curve. Its shape tells you whether J<1 (saturates to the bound δ/(1−J)) or J>1 (bends upward).
  2. The empirical Jacobian norm. Perturb a rollout state by a small realistic ε, re-run, measure divergence growth. This is J measured rather than inferred, and it is the number you are actually training when you change the schedule.
  3. Recovery from off-manifold states. Inject a plausible-but-wrong state and check whether the rollout returns toward the data manifold or leaves it. A contractive model recovers; a teacher-forced one usually cannot, having never been asked to.

All three come from lesson 6's analysis; the point here is that they are training-time instruments, cheap enough to run every checkpoint, and they are how you tell schedules apart before spending a deployment cycle finding out.

Takeaway
Teacher forcing supervises the whole horizon in one parallel pass by resetting the error to zero at every position — which optimises δ and leaves J, the amplification that actually decides rollout survival, entirely unconstrained. Scheduled sampling buys partial contraction at nearly no cost but teaches mode-averaging wherever dynamics branch, because it pairs a self-sampled prefix with a mismatched recorded target. k-step rollout buys full contraction at k× cost. Diffusion forcing — independent per-token noise levels — buys near-full contraction at ≈1× cost and rolls out to any horizon, because denoising a corrupted prefix is the contraction requirement, stated per token and trainable in parallel.

Where this points next

We can now roll out a long way without falling apart. But a stable rollout that ignores the controller is a movie, not a simulator — and lesson 18 already warned that under a pixel-weighted objective, ignoring the action is a favourable trade the optimiser will happily take. Lesson 22 quantifies how much likelihood the action is actually worth, shows why weak conditioning gets dropped, and finds the interior optimum in guidance strength.

Interview prompts