Evaluation you can afford
Lesson 14 left a recipe and a checkpoint, and the question of whether the checkpoint is better than the last one. A success rate is a count of trials, so it carries a spread that can be computed, and the spread decides what an evaluation can say. This lesson runs the evaluation of two Bench checkpoints that differ by eight points, 76 and 68 %, over and over; puts an interval on a count; derives the trials per checkpoint that a gap needs; and prices them in robot-hours. Twenty trials per checkpoint reach a defensible verdict in 8.4 % of evaluations, a ten-point gap takes hundreds, and pairing, a graded score and stopping early each buy some of that back at a price of their own. It cannot say what to spend the robot-hours on instead.
New idea: an evaluation is an experiment on a count: its resolution is set by the number of trials, so the trials per checkpoint are chosen from the gap to be detected, and priced in robot time. The interval, pairing, the choice of metric and the stopping rule are ways to read that resolution or to buy it.
Forces next: Twenty trials cannot tell seventy-six percent from sixty-eight, telling a ten-point improvement apart takes hundreds of trials per policy, and a programme that evaluates honestly spends more robot time on evaluation than on training. The programme also needs data of several kinds, each from a source with a different price, a different usefulness and a different shelf life. How much of each kind do we need, in what unit can they even be compared, and what does each cost?
1 · What twenty trials say
A checkpoint comes out of training, and whether to ship it depends on whether it beats the last one. The instrument is the robot: a trial is one attempt from a start the experimenter sets up, scored finished or not. Published evaluations use few of them. The real-robot rows of π0, Octo, Diffusion Policy and ACT hold 10 to 40 trials per task and method, OpenVLA's headline rows pool 170 and 60 rollouts across tasks, only OpenVLA prints an uncertainty, and none of the five tests a policy comparison for significance. Take twenty trials per checkpoint as the case.
The Bench lets us do what a laboratory cannot, and know the truth. Two checkpoints of the five-post course (lesson 1) are each run on the same 2,000 trials at gust 0.05 rad/s, and the share each finishes is, by definition, its true success rate: 68.0 % for the previous checkpoint and 76.0 % for the new one. Both are tables of 20 demonstrations that all start at the marked point, recorded under gusts of 0.0403 and 0.0537 rad/s (lesson 2's noisier demonstrations); the digits were tuned once so that the pool gives exactly 68.0 and 76.0 %. The new checkpoint is better by 8.0 points. An evaluation of n trials draws n of the 2,000 for each checkpoint. What does one show?
| what the evaluation shows | 20 trials per checkpoint | 495 trials per checkpoint |
|---|---|---|
| the new checkpoint has more successes | 65.2 % | 99.7 % |
| the two have the same number | 12.0 % | 0.1 % |
| the previous checkpoint has more | 22.8 % | 0.2 % |
| a gap significant at 5 %, new ahead | 8.4 % | 80.2 % |
| a gap significant at 5 %, previous ahead | 0.6 % | 0.0 % |
A checkpoint that is truly eight points better is not ahead in 34.8 % of the 20-trial evaluations (a tie or worse), a gap that would hold up at the 5 % level appears in 8.4 %, and the worse checkpoint appears significantly ahead in 0.6 %. The evaluation was not careless: it is as large as the published ones. The count is too small for the question, and the second column is what the rest of this lesson derives.
2 · A count has an interval
A trial finishes with some probability p that depends on the checkpoint and the conditions and not on the trials before it. Run n of them under the same conditions and the successes K are binomial, with mean np and variance np(1 − p), so the estimate p̂ = K/n scatters around p with standard deviation √(p(1 − p)/n): 10.0 points at p = 0.72 and n = 20, halved by four times the trials. An interval turns the scatter into a statement, the true rates the count does not rule out. Three ways to build one, for 15, 19 and 20 successes in 20 trials (95 %, in percent):
| successes in 20 | normal interval | Wilson | Clopper–Pearson |
|---|---|---|---|
| 15 | 56.0–94.0 | 53.1–88.8 | 50.9–91.3 |
| 19 | 85.4–104.6 | 76.4–99.1 | 75.1–99.9 |
| 20 | 100.0–100.0 | 83.9–100.0 | 83.2–100.0 |
The normal interval, p̂ ± 1.96√(p̂(1 − p̂)/n), is centred on the count and cannot know that p ≤ 1: at 19 of 20 it ends above 100 %, and at 20 of 20 it has zero width, an evaluation that certifies perfection after twenty trials. Clopper and Pearson (1934) invert the exact binomial; their interval never covers less than its label, and is the widest. Wilson (1927) inverts the test instead: keep every p for which the count is not surprising, |p̂ − p| ≤ z√(p(1 − p)/n). Squared, that is a quadratic in p, (1 + z²/n)p² − (2p̂ + z²/n)p + p̂² = 0, and its roots are
p = [ p̂ + z²/2n ± z √( p̂(1 − p̂)/n + z²/4n² ) ] / (1 + z²/n), z = 1.96
They stay inside [0, 1] and do not collapse at 0 or n. Averaged over true rates from 0.5 to 0.95, a nominal 95 % interval at n = 20 covers the truth in 89.6 % of evaluations for the normal interval, 95.3 % for Wilson and 97.5 % for Clopper–Pearson. Wilson is the interval this track has used since lesson 1.
Fifteen of 20 is 75 % and compatible with anything from 53.1 to 88.8 %. Ten trials are no better: 5 of 10 and 8 of 10 have Wilson intervals of 23.7 to 76.3 % and 49.0 to 94.3 %, which overlap. At the new checkpoint's 76 % the interval is 35.3 points wide at 20 trials, 16.5 at 100 and 7.5 at 500.
3 · What a gap needs
Two intervals are a picture; a verdict needs a rule fixed before the trials, with a known error rate. Write d̂ = p̂new − p̂prev for the difference of the two success rates from n trials each, and p̄ for the mean of the true rates. If the checkpoints are equal, d̂ has mean 0 and standard deviation σ = √(2p̄(1 − p̄)/n). The rule: declare a difference when |d̂| > 1.96 σ. That declares one in 5 % of evaluations of equal checkpoints (the pooled two-proportion z-test). If the true gap is Δ, d̂ is centred on Δ instead, and the rule fires in a share 1 − β of evaluations, the power, when Δ = (1.96 + zβ) σ; 80 % power is zβ = 0.8416. Solving for Δ and for n:
Δmin(n) = (1.96 + 0.8416) √(2p̄(1 − p̄)/n), n = 2(1.96 + 0.8416)² p̄(1 − p̄) / Δ² = 15.7 p̄(1 − p̄) / Δ²
The smallest gap that n trials resolve falls as 1/√n, and the trials a gap needs grow as 1/Δ²: half the gap, four times the trials. At p̄ = 0.72, the mean of our pair, 20 trials per checkpoint resolve 39.8 points, 100 resolve 17.8, 500 resolve 8.0 and 1,000 resolve 5.6. Our pair needs 495 trials per checkpoint, a ten-point gap 317; lesson 3's repair, from 63.5 to 95.6 %, needs only 25: gaps of tens of points are cheap to see, and the ones that remain are a few points. The formula leans on a normal approximation; summing the binomial probabilities of every pair of counts instead, 495 trials per checkpoint reach the right verdict in 80.2 % of evaluations (§1's table).
The rule matters as much as the number. Three ways to read the same two counts, on our pair:
| rule | names the new checkpoint when the two are equal at 72 % (20 trials) | trials per checkpoint for 80 % power |
|---|---|---|
| the higher count wins | 43.0 % | none: it names a winner of equal checkpoints at any n |
| the two 95 % intervals do not overlap | 0.2 % | 820 |
| the z-test above | 2.6 % | 495 |
A higher count is a coin toss that settles every evaluation whose counts differ: it names the new checkpoint in 43.0 % of the evaluations of two equal ones, and the previous in as many. Non-overlap is valid and strict: two 95 % intervals can overlap while the gap is significant, because the standard deviation of a difference is √2 times one rate's, not twice, so it needs 1.66 times the trials of the test (the TRI LBM Team, 2025, make the same point). The z-test has the stated error rate and the smaller bill of the two rules that control the error.
4 · The price
A trial is a run of the robot, a reset of the scene and a score. A Bench run lasts at most 12 s; a real episode takes minutes. AutoEval (Zhou et al., 2025) logged about 850 evaluation episodes in a 24-hour run with three human interventions in all, 1.7 minutes an episode when a machine resets the scene. Take 3 minutes a trial as the working figure; the widget lets you change it. A comparison costs 2n trials, 2nt/60 robot-hours at t minutes a trial, and a robot-week is 40 hours. At p̄ = 0.72:
| gap to detect | trials per checkpoint | trials in all | robot-hours | robot-weeks |
|---|---|---|---|---|
| 30 points | 36 | 72 | 3.6 | 0.09 |
| 20 points | 80 | 160 | 8.0 | 0.20 |
| 10 points | 317 | 634 | 31.7 | 0.79 |
| 8 points (our pair) | 495 | 990 | 49.5 | 1.24 |
| 5 points | 1,266 | 2,532 | 126.6 | 3.17 |
| 3 points | 3,517 | 7,034 | 351.7 | 8.79 |
For our pair the honest evaluation is 990 trials, 49.5 robot-hours, 1.24 robot-weeks: one checkpoint against one other, on one task. Set that against what it judges. Each checkpoint is 20 demonstrations, 20 runs of the robot, so the two took 40 runs, 2.0 robot-hours at the same 3 minutes, and the evaluation costs 24.8 times the robot time that made what it measures. A sweep makes it worse. The mixture ratios of lesson 14, or any sweep of settings over one dataset, train on data already collected, which costs GPU time and no robot time, and six honest comparisons cost 297 robot-hours, 7.4 robot-weeks, all of it evaluation. AutoEval takes the person out (more than 99 % of the evaluator's time, by its own count) and not the robot: at 1.7 minutes the 990 trials still take 28 hours.
The widget
What to try. Leave the defaults: 20 trials per checkpoint, the previous checkpoint a table, independent trials, success, 3 minutes a trial. The true rates read 68.0 and 76.0 %. One 95 % interval is 35.3 points wide and the smallest gap 20 trials resolve is 39.8 points. The new checkpoint has the higher count in 65.2 % of evaluations and an evaluation names it the better one in 8.4 %; count the green squares among the 30 rows, press the button for 30 more and count again: the rows move, the probability does not. Slide N to 100: the interval is 16.5 points wide, the resolvable gap 17.8 and the right verdict 24.2 %; to 200, 43.1 %. At 500 the right verdict is 80.6 %, the cost 50.0 robot-hours or 1.25 robot-weeks, 25.0 trials for every training run, and the formula for 80 % power reads 495. At 1,000 it is 97.9 % and 100.0 hours. Set the minutes per trial to 1.7, AutoEval's pace, and the 500-trial evaluation takes 28.3 hours.
5 · Buying trials back
Pairing. Run both checkpoints from the same starting conditions and the outcomes come in pairs: both finish, both fail, only the new one finishes (b pairs), only the previous one (c pairs). Only the discordant pairs say which is better (McNemar, 1947). If the checkpoints are equal, b is binomial with b + c trials and probability ½, and the exact test rejects when the smaller count falls in the 5 % tail. Twenty pairs with b = 8 and c = 2 give p = 2 × (1 + 10 + 45)/1024 = 0.109: not significant, though the new checkpoint won 8 of the 10 discordant pairs. With ψ = (b + c)/n the share of discordant pairs and Δ = (b − c)/n, the pairs needed are
n = ( 1.96 √ψ + 0.8416 √(ψ − Δ²) )² / Δ²
For independent outcomes ψ = pnew(1 − pprev) + pprev(1 − pnew) = 40.6 %, and the formula gives 497, within half a percent of §3's 495. Pairing helps as far as ψ is smaller, which it is when the two outcomes share their cause. A robot can match the starting condition and cannot replay the disturbance; a simulator can replay both. Three designs on two pairs, the previous table against the new one and a smooth fit (lesson 2's road not taken, also 68.0 %, §6) against the new one, need this many pairs for 80 % power:
| design | previous table to new: discordant share | pairs | smooth fit to new: discordant share | pairs |
|---|---|---|---|---|
| independent trials | 40.6 % | 495 | 40.6 % | 495 |
| matched starts | 35.9 % | 438 | 41.6 % | 508 |
| same starts and same gusts | 8.0 % | 96 | 34.6 % | 422 |
Matched starts buy little. The start offset decides some of the tables' trials, because their demonstrations all began at the marked point: the new table finishes 52 % of the fifth of trials that start lowest and 87 % of the middle fifth, the smooth fit 72 % and 70 %. Matching starts trims the tables' bill by 11.5 % and the cross-recipe pair's by nothing. Replaying the gusts too turns two similar checkpoints into near copies: the new one never loses a trial the previous one wins (0.0 % of pairs), 8.0 % of pairs are discordant, and 96 pairs suffice, 5.2 times fewer. Different recipes share little even then. Pairing buys what the pair shares, and it needs enough discordant pairs: at ψ = 8 %, twenty pairs hold 1.6 on average and the exact test needs at least 6, so 20 same-gust pairs reach the right verdict in 0.4 % of evaluations. In the widget, 100 same-gust pairs reach it in 82.0 %, and 500 pairs with matched starts in 83.4 % against 80.6 % for independent trials.
6 · The metric is part of the question
Success uses one bit of each trial. A score that gives credit for getting part of the way, here the share of the course length reached before the run ended, uses more. π0 reports a normalised task-progress score instead of success, and RoboArena (Atreya et al., 2025) asks evaluators for a progress score and a preference on every A/B episode and calls the two rankings complementary. But a graded score can order two checkpoints the other way. The third checkpoint is the smooth fit, which finishes 68.0 % of the pool, like the previous table:
| checkpoint | finishes | progress, mean (sd) | trials ending at post 1 | post 2 | post 3 | post 4 | post 5 |
|---|---|---|---|---|---|---|---|
| previous table | 68.0 % | 81.8 (30.0) | 13.6 % | 4.8 | 4.8 | 7.0 | 1.9 |
| new table | 76.0 % | 85.9 (27.8) | 11.4 % | 2.9 | 3.9 | 4.8 | 1.0 |
| smooth fit | 68.0 % | 90.8 (14.8) | 0.6 % | 0.1 | 2.9 | 22.4 | 6.0 |
The tables fail mostly at the first post, when the start is low: of the new table's trials in the lowest fifth of start offsets (0.8 to 3.7 cm below the path) 40 % end there, against 3 % in the middle fifth; the smooth fit fails at the fourth. Success counts a failure at post 1 and one at post 4 the same; progress does not, so the new table against the smooth fit is a reversal: success ranks the new table first by 8.0 points, progress ranks the smooth fit first by 4.9. Progress is the area under the curve of how many runs are still going at each point of the course, success is the curve's last value, and curves that cross are ranked differently by the two. Neither is wrong; they answer different questions, and what the checkpoint is for fixes the right one, before the trials. The metric also sets the price: for the previous table against the new one, progress needs 797 trials per checkpoint (a gap of 4.1 points against a standard deviation near 30) where success needs 495; for the smooth fit against the new table it needs 329. In the widget, set the metric to progress with each previous checkpoint in turn.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
A success rate is a count, and on the Bench twenty trials per checkpoint name the better of two checkpoints eight points apart in 8.4 % of evaluations; a ten-point gap needs 317 trials per checkpoint and our pair 495. At three minutes a trial that is 49.5 robot-hours for one comparison, 24.8 times the robot time behind the two checkpoints, and 7.4 robot-weeks for a sweep of six variants that train on data already collected. Pairing cut the bill by a factor of one to five on the Bench, a graded score cut it or raised it, a sequential test trims a fraction, and none removes it. A programme that evaluates honestly therefore spends more robot time on evaluation than on training, so robot time is the scarce thing and the data the model learns from compete with the evaluation for it. Those data come in several kinds, our own robot's demonstrations, other bodies, video, simulation, each with a price, a usefulness and a shelf life, and an evaluation result has a shelf life too: a robot that wears and a scene that moves make last month's score of the incumbent useless for this month's comparison. How much of each kind do we need, in what unit can they even be compared, and what does each cost?
Interview prompts
- A policy succeeded in 15 of 20 trials. What can you say about its true success rate? (§2 — a Wilson interval, 53.1 to 88.8 %; the normal interval is centred on the count and leaves [0, 1] near the ends.)
- Why can't 20 trials per policy separate 76 % from 68 %? (§1, §3 — the standard deviation of the difference of two counts of 20 is about 14.2 points, so only a gap near 39.8 points clears the test reliably; the right verdict comes in 8.4 % of evaluations.)
- How many trials per policy detect a ten-point improvement at 80 % power? (§3 — n = 2(1.96 + 0.8416)² p̄(1 − p̄)/Δ², which is 317 around 72 %.)
- Why do overlapping 95 % intervals not mean the policies are the same? (§3 — the standard deviation of a difference is √2 times one rate's, not twice; non-overlap costs 1.66 times the trials.)
- What does a paired comparison buy, and when nothing? (§5 — only discordant pairs count, so it buys what the pair shares: 5.2 times from replayed gusts on similar checkpoints, nothing between different recipes.)
- Why is stopping at the first p < 0.05 after every trial invalid? (§5 — every look is another chance: 33.3 % false positives for identical checkpoints.)
- Two checkpoints rank oppositely on success and on progress. What do you do? (§6 — fix the metric before the trials from what the checkpoint is for, report both with intervals, size the experiment for the chosen one.)
Companion reads: World Models · 16 The exam we never stopped taking (which cheap exam ranks policies as the decision does), World Models · 30 Evaluating the thing you trained (why measured and delivered improvement come apart) and Search, Ads and Recsys · 33 A/B testing deep dive (randomisation, power and the launch decision).