Evaluating the thing you trained
Your model's consumer is not a spectator, it is an optimiser — and an optimiser does not sample your errors at random. It hunts for them, because an error that overstates a plan's value is indistinguishable from a good plan. This makes measured improvement and delivered improvement come apart, in a way you can compute.
1 · The optimiser is an adversary you built yourself
Nobody intends this, which is why it is so effective. The planner enumerates K candidate plans, evaluates each in your model, and selects the maximum. Now: your model has error. Some of that error is optimistic — it overstates a plan's outcome. And selecting the maximum is selecting for optimistic error.
Quantify it with the expected maximum of K draws. For zero-mean error of scale δ, the largest of K samples sits roughly
E[max of K] ≈ δ · √(2 ln K)above the mean — the standard extreme-value result, derived for this setting in lesson 8. So the plan the planner picks is inflated by that much, and two different quantities follow:
The gap between them is 2δ√(2 ln K), and it widens with search. A better planner does not extract more value from a flawed model; it extracts more of the flaw. This is why "we improved the planner and performance dropped" is a coherent and common observation rather than a paradox.
2 · Therefore search has an optimum
Two opposing effects in K. Genuine benefit — with more candidates you find genuinely better plans — saturates, since the good plans are found early. Exploitation grows as √(ln K), slowly but without bound. A saturating gain minus an unbounded loss has a maximum:
realised(K) = gtrue·(1 − e−K/K₀) − δ·√(2 ln K)Differentiate and there is an interior K*. Below it, search more. Above it, searching harder actively costs you performance. And note what K* depends on: your model's error scale δ. A more accurate model earns the right to search harder. Search budget is not a planner hyperparameter to be maximised; it is a quantity you derive from your model's measured error.
Push δ above about 0.8 and the realised curve never leaves negative territory: every plan the optimiser likes is a model bug. That regime is real, it is where many first-generation systems live, and the correct response is not a better planner or a bigger search — it is to stop planning against the model until δ comes down.
3 · Which training-time signal predicts downstream value
Ranked by how well they correlate with what a planner ultimately gets. This ranking is the practical payload of the lesson.
Row 2 is the cheap trick worth adopting today. You do not need a full closed-loop harness to get most of the benefit — you only need to evaluate your model on the actions a planner would choose rather than the actions in your validation set. Those are different distributions, and the gap between error on them is a direct estimate of your exploitation exposure. It costs one planning pass per eval batch.
Row 3 deserves its placement because calibration is what bounds the exploitation gap rather than merely correlating with it. A calibrated model reports the uncertainty that makes δ√(2 ln K) predictable, so a risk-aware planner can discount for it — which is why lesson 29's warning about distillation destroying calibration is not an aesthetic complaint.
4 · Three ways an honest team still fools itself
Evaluating on the actions you trained on. Your validation set's actions came from the same behaviour policy as training. The planner's will not. A model can be excellent on the demonstrator's action distribution and arbitrarily bad off it — and that off-distribution region is precisely where the planner operates, because a planner that only proposed demonstrated actions would be pointless.
Selecting a checkpoint on the metric you report. Choose the best of 40 checkpoints by a metric and that metric is now inflated by δ√(2 ln 40) ≈ 2.7δ — the same extreme-value arithmetic, applied to your own model selection. Select on one held-out split, report on a second that was never used for selection.
Letting the reward model that post-trained the model also evaluate it. Lesson 28's consistency reward and lesson 23's IDM are both models you trained. If they also score the final system, you have measured agreement between your own artifacts. Keep a locked set whose labels were measured, and let it be the only gate.
Where this points next
Every constraint is now on the table: acceptance tests, representational ceiling, schedule, conditioning, label economics, head choice, memory budget, staging, mixture, post-training, latency, and an evaluation protocol that resists optimiser pressure. Lesson 31 puts a single concrete specification through all of it and produces the numbers — tokens, FLOPs, GPU-hours, dollars, wall-clock — then hands off to the actor's seat.
Interview prompts
- Why does a planner amplify model error rather than average it out? It selects the maximum over K candidates, and selecting the maximum selects for optimistic error, since an error that overstates a plan's value is indistinguishable from a good plan. (§1)
- Write the two divergent quantities and their gap. Reported g+δ√(2 ln K) versus realised g−δ√(2 ln K), a gap of 2δ√(2 ln K) that widens with search. (§1)
- Explain how improving the planner can lower delivered performance. More search extracts more of the model's optimistic error, so past K* the exploitation term outgrows the saturating genuine benefit. (§2)
- What does the optimal search budget depend on? The model's error scale δ — a more accurate model earns the right to search harder, so K* is derived from measurement rather than tuned. (§2)
- Which single cheap change most improves a training-time metric's predictiveness? Evaluate on planner-selected action sequences rather than dataset actions; the gap between the two directly estimates exploitation exposure and costs one planning pass per batch. (§3)
- Why does calibration bound the exploitation gap rather than merely correlate with it? A calibrated model reports uncertainty that makes δ√(2 ln K) predictable, so a risk-aware planner can discount for it. (§3)
- How much illusory gain does picking the best of 40 checkpoints buy? About δ√(2 ln 40) ≈ 2.7δ — the same extreme-value arithmetic applied to model selection, which is why reporting needs a split never used for selection. (§4)
- Why must the post-training reward model not evaluate the final system? Both it and the IDM are models you trained, so scoring with them measures agreement among your own artifacts rather than agreement with reality. (§4)