all_lessons/ World Models/ 30 · evaluating what you trainedlesson 30 / 31

Evaluating the thing you trained

Your model's consumer is not a spectator, it is an optimiser — and an optimiser does not sample your errors at random. It hunts for them, because an error that overstates a plan's value is indistinguishable from a good plan. This makes measured improvement and delivered improvement come apart, in a way you can compute.

Where we are
We have a trained, post-trained, distilled model and a decision to make about it. Lesson 18 gave the acceptance tests; this lesson is about running them without being deceived. synthetic_vision/11 owns the closed-loop evaluation harness in full and lesson 16 owns the failure taxonomy — we cite both rather than duplicating them, and focus on the training-side decisions: which signals predict downstream value, and how search budget corrupts them.
Forced by 29Two candidate models, and a decision to make between them. This stepMeasure what the planner will actually get — optimiser pressure makes reports and reality diverge. Forces 31With acceptance and evaluation settled, assemble the whole bill.

1 · The optimiser is an adversary you built yourself

Nobody intends this, which is why it is so effective. The planner enumerates K candidate plans, evaluates each in your model, and selects the maximum. Now: your model has error. Some of that error is optimistic — it overstates a plan's outcome. And selecting the maximum is selecting for optimistic error.

Quantify it with the expected maximum of K draws. For zero-mean error of scale δ, the largest of K samples sits roughly

E[max of K] ≈ δ · √(2 ln K)

above the mean — the standard extreme-value result, derived for this setting in lesson 8. So the plan the planner picks is inflated by that much, and two different quantities follow:

what your report says
true gain + δ√(2 ln K)
The model's own estimate of the selected plan's value, which is the number that lands in the slide deck.
what the robot gets
true gain − δ√(2 ln K)
The real outcome of a plan selected because the model was wrong about it — the optimism was the reason it won.

The gap between them is 2δ√(2 ln K), and it widens with search. A better planner does not extract more value from a flawed model; it extracts more of the flaw. This is why "we improved the planner and performance dropped" is a coherent and common observation rather than a paradox.

2 · Therefore search has an optimum

Two opposing effects in K. Genuine benefit — with more candidates you find genuinely better plans — saturates, since the good plans are found early. Exploitation grows as √(ln K), slowly but without bound. A saturating gain minus an unbounded loss has a maximum:

realised(K) = gtrue·(1 − e−K/K₀) − δ·√(2 ln K)

Differentiate and there is an interior K*. Below it, search more. Above it, searching harder actively costs you performance. And note what K* depends on: your model's error scale δ. A more accurate model earns the right to search harder. Search budget is not a planner hyperparameter to be maximised; it is a quantity you derive from your model's measured error.

Search budget: report versus reality
Amber is the model's own estimate of the plan it selected; green is what that plan actually returns. They diverge as 2δ√(2 ln K). The marker is K* — raise δ and watch it collapse toward almost no search at all.
Interactive search-budget optimum.
exploitation gap
—
optimal search
—
realised return
—
diagnosis
—

Push δ above about 0.8 and the realised curve never leaves negative territory: every plan the optimiser likes is a model bug. That regime is real, it is where many first-generation systems live, and the correct response is not a better planner or a bigger search — it is to stop planning against the model until δ comes down.

3 · Which training-time signal predicts downstream value

Ranked by how well they correlate with what a planner ultimately gets. This ranking is the practical payload of the lesson.

1. Return under planner-selected actions, scored in reality or a trusted simthe ground truth
2. Error on planner-selected action sequences, not dataset onesvery strong
3. Calibration (ECE) on discrete downstream eventsstrong — bounds the gap
4. Free-running horizon curve at 2–4× deployment horizonstrong
5. Paired-intervention controllabilitynecessary, not sufficient
6. One-step error on dataset actionsweak
7. Pixel or distributional video fidelitynear-zero

Row 2 is the cheap trick worth adopting today. You do not need a full closed-loop harness to get most of the benefit — you only need to evaluate your model on the actions a planner would choose rather than the actions in your validation set. Those are different distributions, and the gap between error on them is a direct estimate of your exploitation exposure. It costs one planning pass per eval batch.

Row 3 deserves its placement because calibration is what bounds the exploitation gap rather than merely correlating with it. A calibrated model reports the uncertainty that makes δ√(2 ln K) predictable, so a risk-aware planner can discount for it — which is why lesson 29's warning about distillation destroying calibration is not an aesthetic complaint.

4 · Three ways an honest team still fools itself

Evaluating on the actions you trained on. Your validation set's actions came from the same behaviour policy as training. The planner's will not. A model can be excellent on the demonstrator's action distribution and arbitrarily bad off it — and that off-distribution region is precisely where the planner operates, because a planner that only proposed demonstrated actions would be pointless.

Selecting a checkpoint on the metric you report. Choose the best of 40 checkpoints by a metric and that metric is now inflated by δ√(2 ln 40) ≈ 2.7δ — the same extreme-value arithmetic, applied to your own model selection. Select on one held-out split, report on a second that was never used for selection.

Letting the reward model that post-trained the model also evaluate it. Lesson 28's consistency reward and lesson 23's IDM are both models you trained. If they also score the final system, you have measured agreement between your own artifacts. Keep a locked set whose labels were measured, and let it be the only gate.

The one-line protocol
Pre-register the gate before the run: which metric, which split, which threshold, which search budget K. Every degree of freedom you leave open after seeing results is worth roughly δ√(2 ln(#choices)) of illusory improvement. This is the same mathematics as §1, applied to the experimenter instead of the planner — which is a genuinely useful way to remember it.
Takeaway
A planner selecting the best of K plans selects for optimistic model error, so the report shows g+δ√(2 ln K) while the robot gets g−δ√(2 ln K) — a gap that widens with better search. Hence an interior optimum K* that depends on your model's error scale: accuracy earns the right to search. Rank training signals by correlation with delivered value — planner-selected-action error and calibration near the top, pixel fidelity near zero — and remember the same extreme-value arithmetic applies to checkpoint selection and to every un-pre-registered analysis choice you make.

Where this points next

Every constraint is now on the table: acceptance tests, representational ceiling, schedule, conditioning, label economics, head choice, memory budget, staging, mixture, post-training, latency, and an evaluation protocol that resists optimiser pressure. Lesson 31 puts a single concrete specification through all of it and produces the numbers — tokens, FLOPs, GPU-hours, dollars, wall-clock — then hands off to the actor's seat.

Interview prompts