all_lessons/World Models/16 · Evaluationlesson 16 / 31

The exam we never stopped taking: evaluation

Lesson 15 ended with a decoy: a better one-step score, a worse closed loop. This lesson makes the decoy a census. Twelve one-jump models of the Courtyard, differing in capacity, data, the nudges in their logs and a hidden confounder, sit five cheap exams and the decision they exist for; Kendall's τ says which exam ranks them as the decision does. None does for every use: the held-out error is best at the series' goal and worse than chance at a moved one, where the exam that tests the whole action range is best. The decision can be measured, at a price we derive: how many launches, what zero failures prove, how far a winner is flattered. Ten failures from lessons 3 to 15 each have an exam that passes them and one that catches them. The lesson says how good a model is, not how to make it better.

The thesis, here
A world model has no score of its own, only a score for a use. The exam that matters is the decision the model supports; every cheaper exam is a proxy whose agreement with it must be measured, and on twelve models the proxies disagree: the model with the lowest held-out error is not the one whose plans succeed. Which proxy is least wrong depends on the use, and the decision costs a computable number of real trials. Evaluation is a ladder of exams ordered by cost, each with a blind spot another rung covers; spend the real trials where the cheap rungs are blind.
Linear position
Forced by: We now have every piece: state, dynamics, uncertainty, rollouts, causality, practice, planning, abstraction, pixels, memory, objects, actions and a body, each justified by an exam the previous model failed. How do we know that a particular model passes the exam that matters, a better decision in the world, and not only the ones that look good?
New idea: grade a world model by the decision it supports, and treat every cheaper exam as a proxy whose agreement with it and whose blind spot have been measured. The ladder shows where each proxy is blind; binomial arithmetic prices the real trials that look there.
Forces next: Everything in this series assumed that a trained world model exists. How is one trained: which loss, which schedule, which data, and what does it cost?
The plan
Six moves. (1) Build a zoo of twelve models and take the exam that matters. (2) Climb the ladder of cheaper exams and score each by Kendall's τ. (3) Find why the rungs disagree: the log, the use, the planner's search. (4) Price the decision. (5) Read the failure atlas. (6) Design by product, and see what a better model costs.

1 · Twelve models of one world, and the exam that matters

The task is lesson 1's nudge: one impulse of at most 3 m/s at step 12, then 80 steps of coasting; the ball must end inside the goal disc at (6.6, 2.5), radius 0.5 m. A planner with budget K draws K candidate nudges from the disc |a| ≤ 3 m/s, asks a model for the ending of each, and executes the one with the smallest imagined miss in the real Courtyard (lesson 9). A model's decision score is the share of 100 fixed launches on which that push ends in the goal: the exam that matters, and the only one that costs real launches. With the simulator as the model it is the ceiling every model shares: 6 % at K = 1, 43 % at 16, 85 % at 64, 100 % at 128.

A model here is lesson 8's kind: state at the nudge and nudge in, ending out; a ridge regression on M random tanh features, fitted in closed form to a log of n real trials, with E = 4 bootstrap members (lesson 5) whose mean is the answer. It trains in milliseconds, so twenty zoos of twelve are cheap.

modelM · n · nudges in the logwhat differs
1 tiny8 · 500 · ≤ 3 m/scapacity (lesson 4)
2 small24 · 500 · ≤ 3capacity
3 mid64 · 500 · ≤ 3the reference
4 twin64 · 500 · ≤ 3the reference again, another log: the control
5 big128 · 500 · ≤ 3capacity
6 scarce64 · 100 · ≤ 3data
7 plenty64 · 2000 · ≤ 3data
8 cautious64 · 500 · ≤ 1.2lesson 8's operator
9 narrow64 · 500 · ≤ 0.5a more cautious operator
10 weighted64 · 500 · ≤ 3loss weighted towards endings near the goal
11 confounded64 · 500 · ≤ 0.3windy days, operator pushes against each gust (lesson 7)
12 big-cautious128 · 500 · ≤ 1.2capacity, cautious operator

Rank the twelve by the exam everyone runs first, the error on held-out trials from the model's own log. narrow is first with 0.28 m; big is fifth with 0.70 m, 2.5 times worse. Rank them by the decision at K = 64: big scores 45 %, narrow 24 %. The model with the smaller error decides about half as well. Over twenty zoos, each retrained from its own logs, the lowest held-out error is never the best decision (0 of 20 agree); narrow has the lowest error in 19 of them, big the best decision in 13. Lesson 15's decoy was one such pair: its smooth networks beat a model that had never met a wall on one-step error, and lost to it on decisions.

2 · A ladder of exams, each scored against the decision

rungthe exam (lesson); what it needsin the zoo
1one-step or pixel error on held-out data (4, 11); a logthe held-out error
2rollout error, the model on its own output (6); long runs= rung 1 for a one-jump model; §5, row 4
3state probe: read the state out of the code (3); labelled states§5, rows 1–2
4calibration: does the 90 % interval hold 90 %? (2, 5); a logthe coverage of the members' interval
5intervention error: error on actions set by fiat, not drawn from the log (7); a simulator or randomised trialsthe intervention error on 300 nudges over the whole disc
6closed loop inside the model (8, 9); no real trialthe imagined success; the success judged by the other half of the members
7the decision (1); real launchesthe real success

A zoo model answers in one jump, so its one-step and rollout errors coincide and it has no code to probe; rungs 2 and 3 appear in the atlas of §5. Five exams apply to the zoo, in the order of the exam slider: held-out error, coverage, intervention error, imagined success, judged success.

To score an exam against the decision, count the pairs of models that the two rankings order alike (c) and oppositely (d). Kendall's τ is (c − d) / (c + d), corrected for ties: 1 for the same ranking, 0 for none, −1 for the reverse; at τ = 0.4 a random untied pair is ordered as the decision orders it with probability (1 + 0.4)/2 = 70 %. Over twenty zoos (K = 64, the series' goal) the five exams have median τ 0.38, 0.29, 0.25, −0.08 and 0.05. No cheap exam is far above 0.4, and the one that sounds most like the decision, the success the model imagines for its own plan, is the worst. Real launches do better. Against the success on a thousand other launches (twelve zoos), ten score τ = 0.58, twenty 0.71 and two hundred 0.89, where the held-out error, the best cheap exam, scores 0.45 (0.38 over the twenty).

The field met the same thing. A human study "confirms that FVD correlates well with qualitative human judgment" (FVD; Unterthiner et al., 2018), yet FVD "increases only slightly with large temporal corruption" (Ge et al., CVPR 2024). Video generators' "physical understanding is severely limited, and unrelated to visual realism" (Physics-IQ; Motamed et al., 2025). In a closed loop "visual quality alone does not guarantee task success, controllability matters more" (World-in-World; Zhang et al., ICLR 2026), and WorldArena (Shang et al., 2026) reports "a significant perception-functionality gap". Not every cheap benchmark fails: DreamGen Bench (Jang et al., 2025) "shows a strong correlation between benchmark performance and downstream policy success". Whether a cheap exam tracks the decision depends on the exam and the use (§3).

3 · Why the rungs disagree: the log, the use and the planner's search

The exam samples from the log; the planner does not. A held-out trial is drawn from the model's own recipe, so its nudge is one the log could hold. The planner draws over the whole disc and keeps the nudge the model likes best, so it asks where the log did not go: lesson 7's seeing against doing. The intervention error measures that, on 300 nudges set by fiat over the whole disc.

model, K = 64nudges in logheld-out errorerror under nudges set by fiatpicks beyond the logreal success
narrow≤ 0.50.28 m9.37 m88 %24 %
confounded≤ 0.30.36 m10.99 m37 %13 %
big≤ 30.70 m0.68 m0 %45 %

For big, whose log covers the disc, the two errors agree. For narrow the second is 33 times the first: 88 % of its picks lie beyond the largest nudge in its log, and there it hits 16 times in 88, against 8 in 12 inside. confounded is the sharper case: third lowest held-out error, worst error under nudges set by fiat. If the operator pushes against each gust (lesson 7), a = u − κw (u a random nudge of spread σu = 0.15 m/s, w a gust of spread σw = 0.05 m/s², κ = 5.76 s the nudge that cancels a unit gust), the ending is f(s) + Bu (f the ending without a nudge): the gust cancels out of it, and a regression of ending on nudge recovers only the fraction σu² / (σu² + κ²σw²) = 0.213 of the true effect B on an open floor (40,000 simulated trials: 0.213) and −0.04 in the walled Courtyard. The log says that nudges do nothing. The model fitted to it has a held-out error 3.1 times below the reference's, 30 times larger under nudges set by fiat, and a real success of 13 %. Over twenty zoos both narrow and confounded beat the reference on held-out error in 20 of 20, and in 20 of 20 the intervention error exposes them (more than five times their held-out error).

The exam is for a use. Move the goal from (6.6, 2.5) to (4.0, 1.2). Of the nudges that hit the series' goal, 28 % are at most 1.2 m/s, the reach of a cautious log; of those that hit the moved goal, 1.8 %. So cautious falls from 36 % to 10 % (median of twenty zoos, K = 64) while big goes from 43 % to 41 %, and the exams reshuffle. The median τ of the held-out error goes from 0.38, the best of the five, to −0.15; that of the intervention error goes from 0.25 to 0.76. An exam is valid for a use to the extent that it asks what the use will ask. The intervention error asks what a planner asks, of the whole disc, and stays positive for both goals where the held-out error does not.

The planner finds the holes. The success a model imagines for its own plan is the exam that sounds right, and at the series' goal the weakest. Among 128 candidates the planner keeps the one with the smallest imagined miss, the one whose error flatters most: lesson 8's optimiser's curse. At K = 128 a zoo model imagines a hit on 91 % of launches on average and really hits on 26 %; raising K pushes the imagined number towards 100 % for every model, so it ranks none. An exam can hold out the planner: plan with two members of the ensemble, let the other two grade the pick. The judged success averages 66 %, closer and still flattering, because the members share their features and log, and so their mistakes.

The zoo: which exam ranks the models like the decision?
Left: one dot per model (numbers as in §1); across, the exam you choose, better to the right; up, the real success of its planner; τ is the rank agreement; ★ marks the best model, purple is model A, amber model B. Right: for one launch, the nudge each picks (arrow), the ending it imagines (ring) and the one it gets (dot); below, success against K, solid real, dashed imagined, grey exact.
τ of this exam with the decision
—
τ: held-out · coverage · intervention · imagined · judged
—
best model on this exam
—
model A
—
model B
—
lowest held-out error · best decision
—
launches per policy to tell A from B
—
this launch
—
Show the core JS
best = Infinity; bi = 0;
for (k = 0; k < K; k++) {
  q = (i * K + k) * 2; m = Math.hypot(P[q] - g.x, P[q + 1] - g.y);
  if (m < best) { best = m; bi = k; }
  if (Math.hypot(S.TE[(i * K + bi) * 2] - g.x, S.TE[(i * K + bi) * 2 + 1] - g.y) < g.r) real[k] += 1 / N;
  if (best < g.r) imag[k] += 1 / N;                                           // what the model said it would do
}
...
a = Math.sign(x[i] - x[j]); b = Math.sign(y[i] - y[j]);
if (a === 0 && b === 0) continue; if (a === 0) tx++; else if (b === 0) ty++; else if (a === b) c++; else d++;
return (c - d) / Math.sqrt((c + d + tx) * (c + d + ty));

What to try. It opens on the held-out error, K = 64, the series' goal, narrow (A) against big (B). (1) τ = 0.46 for this zoo; the star marks narrow, the best model on this exam, and the readout puts the pair side by side: 0.28 m and 24 % against 0.70 m and 45 %. (2) Step the exam slider through the five exams; the second readout lists their τ in order, and imagined success is the lowest, −0.11. (3) Switch to the moved goal: the five become −0.23, 0.20, 0.72, 0.69, 0.56, and narrow falls to 1 %. (4) Back on the series' goal, raise K to 128: the imagined curves of A and B reach 69 % and 97 %, the real ones 26 % and 50 %, the exact model 100 %. (5) On launch #40, A imagines landing 0.04 m from the goal's centre and lands 2.20 m away; B imagines 0.16 m and lands 0.41 m, a hit. (6) twin against mid, one recipe with other luck: 26 % against 19 %; telling them apart would take 558 launches per policy.

4 · What the decision costs

A success rate measured on n launches has standard error √(pq/n), q = 1 − p: 5.0 points at p = 0.45 and n = 100. To tell two policies apart, the gap must stand out from the error of the difference, whose variance is 2p̄q̄/n if the policies are equal and (p1q1 + p2q2)/n if not. A 5 % test rejects beyond zα/2 = 1.96 standard errors of the first; it has power 80 % when the true gap clears that threshold by zβ = 0.84 standard errors of the second. Solving for n per policy:

n = ( zα/2·√(2p̄q̄) + zβ·√(p1q1 + p2q2) )² / (p1 − p2)²

For 55 % against 40 % that is 172.8, so 173 launches per policy. Simulated at that size, the gap is found in 79 % of experiments (target 80) and two equal policies are called different in 4.7 % (target 5). The widget's pair, 24 % against 45 %, needs 80.

A model is a sample. Retrain the reference 60 times, each with its own log, features and bootstrap, and measure each on 900 other launches: real success runs from 13 % to 49 %, mean 27 %, standard deviation 7.9 points, 1.8 times the 4.4-point binomial error of a 100-launch benchmark. Two single training runs differ by that spread: the twin differs from the mid by 9 points (median over twenty zoos), by up to 33. The cheapest protection is the control the zoo carries: train the reference twice and believe only gaps larger than the A/A gap, the gap between a recipe and itself.

The winner of a benchmark is flattered. Pick the best of the twelve by its score on n launches and measure it on other launches (every disjoint block of a thousand fresh launches, twelve zoos, K = 64):

benchmark launches n102050100200
winner's score above its fresh score, points12.48.84.02.50.9
lesson 8's law, points17.110.04.62.41.2

The second row is lesson 8's law with models for candidates: for true successes with spread σ (13 points here) measured with noise δ = √(p̄q̄/n), the winner of 12 is flattered by δ²/√(σ² + δ²) · E[max12], E[max12] = 1.63 being the expected maximum of twelve standard normals.

Zero failures proves less than it sounds. If the true failure rate were p, n launches would all be clean with probability (1 − p)n. The largest p that leaves this at least 5 % solves (1 − p)n = 0.05, that is p = 1 − 0.051/n. Since ln(1 − p) ≈ −p for small p, p ≈ ln 20 / n = 3.00 / n: the rule of three, that zero failures in n launches bound the failure rate at about 3/n with 95 % confidence. After 100 clean launches the exact bound is 2.95 %, after 300 it is 0.99 %. By simulation, with the rate at the 100-launch bound the launches are clean in 4.9 % of 40,000 runs; at 2 % in 13.5 % (exactly 13.3 %). Showing a rate below 1 % takes 299 clean launches, below 0.1 % 2,995.

A model can judge other policies, not its own. When launches are scarce a world model ranks policies. 1X did so for its Redwood checkpoints (16 June 2025): "Given a true real-world success rate gap of 15% between two policies, a World Model with 70% accuracy can accurately predict the better policy with 90% success." One reading, ours and not theirs: the model's verdict on an episode is right with probability a = 0.7 whoever the policy is. The gap then shrinks by 2a − 1 = 0.4, 55 % against 40 % is seen as 52 % against 46 %, and a 90 % chance of ranking the pair right takes about 227 judged episodes per policy (simulation at 227: 89 %). In the zoo each model grades the plans of the other eleven with median τ = 0.49 and a mean bias of −2.7 points, but grades its own plan 58 points too high. Gemini Robotics in a Veo World Simulator (11 Dec 2025) reports accurate prediction of relative policy performance in nominal and out-of-distribution conditions, validated by over 1,600 real evaluations of 8 checkpoints on 5 bimanual tasks: 40 per checkpoint and task if spread evenly, where a 15-point gap between two cells is detected 27 % of the time. The validation is about ranking across many cells, not a certificate for one.

5 · The failure atlas: ten failures, two exams each

Ten failures of lessons 3 to 15 each have a row: a cheap exam that the broken model passes and another that catches it. Rows 5 and 6 are read off the zoo above. The others are small separate runs of that lesson's experiment, not part of the widget: lesson 3's instruments (rows 1–2), the fork of lesson 5 (3), a one-step model fitted to 150 launches (4, 10), the exact model with a short-sighted planner (7), windows of readings behind a 1.6 m curtain (8), two identical balls (9).

failure (lesson)an exam it passesthe exam that catches it
1 · nuisance kept (3)rebuild the picture: the code explains 0.92 of the picture's variance (a predictive code: 0.00)state probe: the ball is read with r² = 0.04 (predictive code: 0.96)
2 · collapse (3)next-code loss: 0.000 (a code that keeps its variance: 0.070)rank of the code: 1.01 against 12.0; probe r² 0.01 against 0.97
3 · mean of modes (5, 11)squared error: 0.49 m², at the floor of 0.52 for any single answerwhere the answer lies: only 11 % of endings are within 0.3 m of the model's 2.45 m
4 · compounding (6)one-step error: 4.5 mmending after 80 steps: 1.04 m, 232 times more
5 · confounding (7)held-out error: 0.36 m (the reference: 1.13 m)error under nudges set by fiat: 11.0 m
6 · exploitation (8)success the model imagines: 91 % at K = 128success judged by held-out members: 66 %; in the world: 26 %
7 · myopia (9, 10)model error: none, the model is the simulatorclosed loop (K = 128), planner scoring the ball 40, 60, 81 steps after the nudge: 0 %, 51 %, 100 %
8 · forgetting (4, 11, 12)next reading while visible: 0.021 m with a window of 8, 0.019 m with 16reappearance after 10 hidden steps: right in 22 % of launches with 8, 100 % with 16
9 · identity swaps (13)error of the set of two positions: zero, whichever way the balls wentwhich is which: matching to the last sighting is wrong in 9.2 % of launches (the balls crossed in 9.1 %)
10 · contact blur (14, 15)one-step error over all steps: 4.5 mmerror at contact steps: 49 mm, 56 times free flight, on 0.8 % of the steps

The cheap exam is never wrong; it answers a question the failure does not touch. The atlas is not the whole series: a force a camera cannot see and a goal given in words (lesson 15) have no analogue in a ball-in-a-yard zoo, and the contrast between a world model and a policy is not tested here; none has a row.

6 · Design by product, and what to measure first

The one real push succeeds only if some candidate hits and the model's favourite is one of the hits. The first factor, reach, belongs to the task and the budget: the exact planner has 85 % at K = 64. The second, selection, belongs to the model: real success divided by reach. big picks a hit 53 % of the time when one is on offer, narrow 28 %, so 45 % = 85 % × 53 % and 24 % = 85 % × 28 %. Real systems are chosen on more factors: the horizon over which a model can be trusted (lesson 6), how well it obeys the action (controllability, lessons 11 and 14), the fidelity where the decision is made, the cost. When factors multiply, a small one caps the product however large the others, and an exam of one factor cannot rank systems that differ in another.

What does better selection cost? Capacity and data, at K = 128 (mean of three zoos; the exact model scores 100 %). A trial is 81 steps, 8.1 s of world time, plus a reset.

trials in the log (world time)32 units128 units256 units
250 (0.56 h)17 %35 %41 %
1000 (2.25 h)18 %53 %70 %
4000 (9 h)23 %56 %81 %

Neither alone is enough: sixteen times the data lifts 32 units from 17 % to 23 %, eight times the units lift 250 trials to 41 %, both together to 81 %.

What to measure first. (1) Compare the inputs the use will present with the range of the log. (2) Run the exam that samples those inputs, the intervention error if the planner will leave the log. (3) Grade plans with models that did not choose them. (4) Spend real launches on the finalists, 173 per policy for a 15-point gap, once the A/A control has shown the noise floor. Never rank on imagined success; use the held-out error only to choose between retrains of one recipe.

Road not taken · trust human preference alone
Raters see ghosts and broken physics at a glance, and VBench (Huang et al., 2023) has human-preference annotations for its 16 dimensions. But a rater judges what a model draws for the nudges a log holds, as the held-out error does; the decision asks about nudges nobody drew. On logged nudges narrow draws endings closer to the truth (0.28 m) than big (0.70 m), as such a rater would see, and succeeds on 24 % against 45 %.
Road not taken · evaluate only on held-out data
It is supervised learning's standard and lesson 7's seeing: the held-out set comes from the process that made the log. It is the right exam for comparing retrains of one recipe (τ = 0.55 among 60 copies) and the wrong one across logs: confounded has 3.1 times lower error than the reference and a real success of 13 %.
Road not taken · more benchmarks
If one benchmark can be gamed, use sixteen, as VBench does. Each is still a finite sample selected on: the best of twelve on 100 launches is flattered by 2.5 points, on 20 by 8.8; the more models are ranked, the larger E[max] grows, and a model tuned to a benchmark is no longer measured by it. The repair is a held-back set, rotated and refreshed, not more dimensions.
What this lesson did not do
It graded one-jump models of one task in a world whose truth is a simulator; real truth costs a robot, and τ over twelve models is itself noisy. The zoo varied capacity, data, the operator's range and one confounder, not rollout length or noise; rollouts and probes appear only in the atlas, as separate small runs. It did not grade video models, latent-action models or language-conditioned planners, and it did not say how to make a model better: that is training.

Common mistakes / failure modes

"the model with the lowest held-out error is the best"
In 0 of 20 zoos it was the best by decision (§1, §3).
"the success a model imagines tells us how good its plan is"
At K = 128 it imagines 91 % and gets 26 % (§3).
"no failures in 100 launches means a failure rate of zero"
It bounds the rate at 3.0 % with 95 % confidence, about 3/n (§4).
"the benchmark winner will score that much in deployment"
On 20 launches it is flattered by 8.8 points, on 100 by 2.5 (§4).

Checkpoint exercise

Try it
A planner runs 300 launches and never fails. (a) What failure rate can you rule out with 95 % confidence? (b) How many clean launches would show a rate below 1 %? (c) Two policies are at 60 % and 50 %: how many launches per policy give an 80 % chance of seeing the difference at the 5 % level? Answer: (a) 1 − 0.051/300 = 0.99 %, about 3/300: rates above that are ruled out. (b) ln 0.05 / ln 0.99, rounded up, is 299. (c) the formula of §4 gives 387.3, so 388 per policy: a 10-point gap costs more than twice what a 15-point gap does.

Where this points next

The ladder says how good a model is, and the price list says what knowing costs. The best model in the zoo succeeds on 45 % of launches at K = 64 where the exact model manages 85 %. No exam closes that gap: a better model is found by training, and the grid shows that how large, on how much, and from where move the decision from 17 % to 81 %, with 256 units and 9 hours of world time. Everything in this series assumed that a trained world model exists. How is one trained: which loss, which schedule, which data, and what does it cost?

Takeaway
A world model has a score for a use, not a score of its own. The exam that matters is the decision: the share of real launches on which the plan made with the model succeeds. Every cheaper exam is a proxy whose Kendall τ with the decision is a measurement, at best about 0.38 in the median zoo, and which proxy is best changes with the use. The planner exploits the model, so an exam must use it as the use will, over the whole action range, graded by models that did not choose the plan. The decision costs launches: 173 per policy for a 15-point gap, a bound of 3/n after n clean ones, a retraining spread of 7.9 points, a flattered winner. Spend real launches where the cheap rungs are blind.

Interview prompts

Companion reads: Lesson 30 · Training-side evaluation (the same exams during training), Robot Model Training · 15 Evaluation you can afford (launch budgets for robots) and Synthetic Vision · 11 What braking changes (the closed-loop harness).