From teacher forcing to rollout
One line of the training loop decides whether your model survives deployment: at step k, does it consume the recorded token or its own last guess? The convenient choice is also the one that hides the only error a planner will ever meet — and there is now a way to avoid it without paying for it.
1 · The exam with the answer key open
Teacher forcing is the default because it is fast. Every position in the sequence is supervised in parallel, each conditioned on the true prefix, so one forward pass trains the whole horizon at once. It is also, viewed as an exam, a student solving each problem with the previous answer already filled in correctly.
Deployment is the same exam with the key removed and each answer feeding the next. The distributions the model is asked to handle are simply different objects:
train: ẑk+1 = Fθ(zktrue, ak) deploy: ẑk+1 = Fθ(ẑk, ak)This is exposure bias, and here is the part that catches people: it is not merely that the model is untested on its own outputs. It is that teacher forcing gives the model no reason to be contractive. Recall the two pieces of the recurrence — old error amplified by the Jacobian J, plus fresh error δ. Teacher forcing only ever penalises δ, because ek is reset to zero at every position. The quantity that decides whether a rollout survives is invisible to the objective.
2 · The four schedules
Each schedule is a different answer to "how much of its own output does the model eat during training," and each has a cost, because own-output means sequential.
| schedule | what the model consumes | effect on J | training cost |
|---|---|---|---|
| teacher forcing | always the true prefix | none — unconstrained | 1× (fully parallel) |
| scheduled sampling | own sample with probability p, ramped up over training | partial, ∝ p | ≈1× (still one pass; mixed inputs) |
| k-step rollout | its own outputs for k steps, then re-anchor | full, up to k | k× (sequential unroll) |
| diffusion forcing | a prefix at independent per-token noise levels | near-full | ≈1× (still parallel) |
The first three are the classical ladder: pay compute, buy robustness. The fourth is the interesting one, so let us see why it escapes the trade.
3 · Why independent noise levels buy rollout robustness for free
The trick is to stop thinking of the prefix as "true or self-generated" and think of it as "noisy to some degree." Corrupt each token in the context with its own independently drawn noise level σk, and train the model to denoise the next token given that ragged context.
Now count what the model has been taught. A single training example presents some tokens nearly clean and others heavily corrupted — in every combination, because the levels are drawn independently. Two consequences follow directly:
- It has seen degraded prefixes. A self-generated token is, to first order, a clean token plus error. Training across all noise levels covers that condition, so it is no longer out of distribution.
- It has been forced to be contractive. Mapping a noisy input to a clean target is the statement |∂F/∂z| < 1 in the corrupted directions. Denoising and contraction are the same requirement, so J is now under gradient pressure — which teacher forcing never achieved.
And crucially the levels are drawn per-token, not per-sequence, so nothing is sequential and the whole context still trains in one parallel pass. One extra benefit falls out: because any noise-level pattern is valid, the same trained model rolls out to any horizon, including horizons longer than any training sequence, without a separate schedule.
Set δ to its minimum and note that the curves still separate. That is the lesson's whole point: the schedules do not differ in one-step accuracy, they differ in J. You cannot buy your way out of a bad schedule with a better one-step fit.
4 · The trap inside scheduled sampling
Scheduled sampling looks like a free lunch: same cost, partial robustness. It has a specific bias, and in a world model the bias matters more than it does in language.
Suppose the true future branches — the mug can tip left or right. The model samples a prefix in which it tipped left. Scheduled sampling then trains it toward the recorded continuation, in which the mug tipped right. The gradient says: "given the mug is falling left, predict that it lands right." That target is not merely unhelpful, it is physically incoherent, and repeated exposure teaches the model to hedge — to produce the blurry average of incompatible worlds, which is exactly what lesson 24 is about avoiding.
5 · What to log so you can see any of this
One-step validation loss will not show you a schedule problem. Three cheap diagnostics will.
- The horizon curve. Free-running error against k, for k out to 2–4× your deployment horizon. Not a number — the curve. Its shape tells you whether J<1 (saturates to the bound δ/(1−J)) or J>1 (bends upward).
- The empirical Jacobian norm. Perturb a rollout state by a small realistic ε, re-run, measure divergence growth. This is J measured rather than inferred, and it is the number you are actually training when you change the schedule.
- Recovery from off-manifold states. Inject a plausible-but-wrong state and check whether the rollout returns toward the data manifold or leaves it. A contractive model recovers; a teacher-forced one usually cannot, having never been asked to.
All three come from lesson 6's analysis; the point here is that they are training-time instruments, cheap enough to run every checkpoint, and they are how you tell schedules apart before spending a deployment cycle finding out.
Where this points next
We can now roll out a long way without falling apart. But a stable rollout that ignores the controller is a movie, not a simulator — and lesson 18 already warned that under a pixel-weighted objective, ignoring the action is a favourable trade the optimiser will happily take. Lesson 22 quantifies how much likelihood the action is actually worth, shows why weak conditioning gets dropped, and finds the interior optimum in guidance strength.
Interview prompts
- Why does teacher forcing leave the Jacobian unconstrained? It resets the input error to zero at every position, so the objective only ever penalises the fresh one-step error δ, never the amplification of accumulated error. (§1)
- Two models have identical one-step loss but very different 20-step rollouts. What differs? Their Jacobian norms: the recurrence e(k+1)=J·e(k)+δ compounds through J, which one-step loss does not measure. (§1)
- Explain the specific bias of scheduled sampling in a branching world. A self-sampled prefix describing one branch is paired with the recorded continuation of another, so the gradient asks for a physically incoherent transition and teaches mode-averaging. (§4)
- Why does diffusion forcing stay parallel while k-step rollout does not? Noise levels are drawn independently per token, so no token's input depends on the model's own generated output and the whole context trains in one pass. (§3)
- Argue that denoising and contraction are the same requirement. Mapping a corrupted input to the clean target means perturbations shrink under the map, which is exactly |∂F/∂z| < 1 in the corrupted directions. (§3)
- Why can a diffusion-forced model roll out beyond its training sequence length? Any pattern of per-token noise levels is a valid training configuration, so no single horizon is baked into the schedule. (§3)
- Name three training-time diagnostics that distinguish schedules. The free-running error-versus-horizon curve's shape, the empirically measured Jacobian norm from perturbed rollouts, and recovery from injected off-manifold states. (§5)