all_lessons/Robot Model Training/15 · Evaluationlesson 15 / 24

Evaluation you can afford

Lesson 14 left a recipe and a checkpoint, and the question of whether the checkpoint is better than the last one. A success rate is a count of trials, so it carries a spread that can be computed, and the spread decides what an evaluation can say. This lesson runs the evaluation of two Bench checkpoints that differ by eight points, 76 and 68 %, over and over; puts an interval on a count; derives the trials per checkpoint that a gap needs; and prices them in robot-hours. Twenty trials per checkpoint reach a defensible verdict in 8.4 % of evaluations, a ten-point gap takes hundreds, and pairing, a graded score and stopping early each buy some of that back at a price of their own. It cannot say what to spend the robot-hours on instead.

The thesis, here
A success rate is not a property of a policy. It is the count of successes in some number of trials, and a count has a spread that shrinks only as the square root of that number. Two checkpoints can be told apart only if the gap between them is larger than the spread, so the trials an evaluation needs are set by the smallest gap worth finding, and every trial costs a robot and a person for minutes. Pairing the trials, scoring progress instead of success and stopping early buy resolution per trial; none of them makes the bill small.
Linear position
Forced by: Staging the training and mixing the old data into the new keeps the language the model started with, at a measurable price in speed and in the tasks it learns. We now have a recipe and a checkpoint, and need to say whether the checkpoint is better than the last one. An evaluation is a count of trials, and each trial costs a robot and a person for minutes. How many trials does it take to tell one policy from another, and what does that do to the cost of the whole programme?
New idea: an evaluation is an experiment on a count: its resolution is set by the number of trials, so the trials per checkpoint are chosen from the gap to be detected, and priced in robot time. The interval, pairing, the choice of metric and the stopping rule are ways to read that resolution or to buy it.
Forces next: Twenty trials cannot tell seventy-six percent from sixty-eight, telling a ten-point improvement apart takes hundreds of trials per policy, and a programme that evaluates honestly spends more robot time on evaluation than on training. The programme also needs data of several kinds, each from a source with a different price, a different usefulness and a different shelf life. How much of each kind do we need, in what unit can they even be compared, and what does each cost?
The plan
Six moves. (1) Run one evaluation many times and see what twenty trials say. (2) Put an interval on a count. (3) Ask what a gap needs: the smallest detectable difference and the trials per checkpoint. (4) Price the trials in robot-hours and set the bill against the training it judges. (5) Buy trials back, by pairing, and see why stopping early is not free. (6) Fix the metric before the trials, because the ranking can reverse.

1 · What twenty trials say

A checkpoint comes out of training, and whether to ship it depends on whether it beats the last one. The instrument is the robot: a trial is one attempt from a start the experimenter sets up, scored finished or not. Published evaluations use few of them. The real-robot rows of π0, Octo, Diffusion Policy and ACT hold 10 to 40 trials per task and method, OpenVLA's headline rows pool 170 and 60 rollouts across tasks, only OpenVLA prints an uncertainty, and none of the five tests a policy comparison for significance. Take twenty trials per checkpoint as the case.

The Bench lets us do what a laboratory cannot, and know the truth. Two checkpoints of the five-post course (lesson 1) are each run on the same 2,000 trials at gust 0.05 rad/s, and the share each finishes is, by definition, its true success rate: 68.0 % for the previous checkpoint and 76.0 % for the new one. Both are tables of 20 demonstrations that all start at the marked point, recorded under gusts of 0.0403 and 0.0537 rad/s (lesson 2's noisier demonstrations); the digits were tuned once so that the pool gives exactly 68.0 and 76.0 %. The new checkpoint is better by 8.0 points. An evaluation of n trials draws n of the 2,000 for each checkpoint. What does one show?

what the evaluation shows20 trials per checkpoint495 trials per checkpoint
the new checkpoint has more successes65.2 %99.7 %
the two have the same number12.0 %0.1 %
the previous checkpoint has more22.8 %0.2 %
a gap significant at 5 %, new ahead8.4 %80.2 %
a gap significant at 5 %, previous ahead0.6 %0.0 %

A checkpoint that is truly eight points better is not ahead in 34.8 % of the 20-trial evaluations (a tie or worse), a gap that would hold up at the 5 % level appears in 8.4 %, and the worse checkpoint appears significantly ahead in 0.6 %. The evaluation was not careless: it is as large as the published ones. The count is too small for the question, and the second column is what the rest of this lesson derives.

2 · A count has an interval

A trial finishes with some probability p that depends on the checkpoint and the conditions and not on the trials before it. Run n of them under the same conditions and the successes K are binomial, with mean np and variance np(1 − p), so the estimate p̂ = K/n scatters around p with standard deviation √(p(1 − p)/n): 10.0 points at p = 0.72 and n = 20, halved by four times the trials. An interval turns the scatter into a statement, the true rates the count does not rule out. Three ways to build one, for 15, 19 and 20 successes in 20 trials (95 %, in percent):

successes in 20normal intervalWilsonClopper–Pearson
1556.0–94.053.1–88.850.9–91.3
1985.4–104.676.4–99.175.1–99.9
20100.0–100.083.9–100.083.2–100.0

The normal interval, p̂ ± 1.96√(p̂(1 − p̂)/n), is centred on the count and cannot know that p ≤ 1: at 19 of 20 it ends above 100 %, and at 20 of 20 it has zero width, an evaluation that certifies perfection after twenty trials. Clopper and Pearson (1934) invert the exact binomial; their interval never covers less than its label, and is the widest. Wilson (1927) inverts the test instead: keep every p for which the count is not surprising, |p̂ − p| ≤ z√(p(1 − p)/n). Squared, that is a quadratic in p, (1 + z²/n)p² − (2p̂ + z²/n)p + p̂² = 0, and its roots are

p = [ p̂ + z²/2n ± z √( p̂(1 − p̂)/n + z²/4n² ) ] / (1 + z²/n), z = 1.96

They stay inside [0, 1] and do not collapse at 0 or n. Averaged over true rates from 0.5 to 0.95, a nominal 95 % interval at n = 20 covers the truth in 89.6 % of evaluations for the normal interval, 95.3 % for Wilson and 97.5 % for Clopper–Pearson. Wilson is the interval this track has used since lesson 1.

Fifteen of 20 is 75 % and compatible with anything from 53.1 to 88.8 %. Ten trials are no better: 5 of 10 and 8 of 10 have Wilson intervals of 23.7 to 76.3 % and 49.0 to 94.3 %, which overlap. At the new checkpoint's 76 % the interval is 35.3 points wide at 20 trials, 16.5 at 100 and 7.5 at 500.

3 · What a gap needs

Two intervals are a picture; a verdict needs a rule fixed before the trials, with a known error rate. Write d̂ = p̂new − p̂prev for the difference of the two success rates from n trials each, and p̄ for the mean of the true rates. If the checkpoints are equal, d̂ has mean 0 and standard deviation σ = √(2p̄(1 − p̄)/n). The rule: declare a difference when |d̂| > 1.96 σ. That declares one in 5 % of evaluations of equal checkpoints (the pooled two-proportion z-test). If the true gap is Δ, d̂ is centred on Δ instead, and the rule fires in a share 1 − β of evaluations, the power, when Δ = (1.96 + zβ) σ; 80 % power is zβ = 0.8416. Solving for Δ and for n:

Δmin(n) = (1.96 + 0.8416) √(2p̄(1 − p̄)/n), n = 2(1.96 + 0.8416)² p̄(1 − p̄) / Δ² = 15.7 p̄(1 − p̄) / Δ²

The smallest gap that n trials resolve falls as 1/√n, and the trials a gap needs grow as 1/Δ²: half the gap, four times the trials. At p̄ = 0.72, the mean of our pair, 20 trials per checkpoint resolve 39.8 points, 100 resolve 17.8, 500 resolve 8.0 and 1,000 resolve 5.6. Our pair needs 495 trials per checkpoint, a ten-point gap 317; lesson 3's repair, from 63.5 to 95.6 %, needs only 25: gaps of tens of points are cheap to see, and the ones that remain are a few points. The formula leans on a normal approximation; summing the binomial probabilities of every pair of counts instead, 495 trials per checkpoint reach the right verdict in 80.2 % of evaluations (§1's table).

The rule matters as much as the number. Three ways to read the same two counts, on our pair:

rulenames the new checkpoint when the two are equal at 72 % (20 trials)trials per checkpoint for 80 % power
the higher count wins43.0 %none: it names a winner of equal checkpoints at any n
the two 95 % intervals do not overlap0.2 %820
the z-test above2.6 %495

A higher count is a coin toss that settles every evaluation whose counts differ: it names the new checkpoint in 43.0 % of the evaluations of two equal ones, and the previous in as many. Non-overlap is valid and strict: two 95 % intervals can overlap while the gap is significant, because the standard deviation of a difference is √2 times one rate's, not twice, so it needs 1.66 times the trials of the test (the TRI LBM Team, 2025, make the same point). The z-test has the stated error rate and the smaller bill of the two rules that control the error.

4 · The price

A trial is a run of the robot, a reset of the scene and a score. A Bench run lasts at most 12 s; a real episode takes minutes. AutoEval (Zhou et al., 2025) logged about 850 evaluation episodes in a 24-hour run with three human interventions in all, 1.7 minutes an episode when a machine resets the scene. Take 3 minutes a trial as the working figure; the widget lets you change it. A comparison costs 2n trials, 2nt/60 robot-hours at t minutes a trial, and a robot-week is 40 hours. At p̄ = 0.72:

gap to detecttrials per checkpointtrials in allrobot-hoursrobot-weeks
30 points36723.60.09
20 points801608.00.20
10 points31763431.70.79
8 points (our pair)49599049.51.24
5 points1,2662,532126.63.17
3 points3,5177,034351.78.79

For our pair the honest evaluation is 990 trials, 49.5 robot-hours, 1.24 robot-weeks: one checkpoint against one other, on one task. Set that against what it judges. Each checkpoint is 20 demonstrations, 20 runs of the robot, so the two took 40 runs, 2.0 robot-hours at the same 3 minutes, and the evaluation costs 24.8 times the robot time that made what it measures. A sweep makes it worse. The mixture ratios of lesson 14, or any sweep of settings over one dataset, train on data already collected, which costs GPU time and no robot time, and six honest comparisons cost 297 robot-hours, 7.4 robot-weeks, all of it evaluation. AutoEval takes the person out (more than 99 % of the evaluator's time, by its own count) and not the robot: at 1.7 minutes the 990 trials still take 28 hours.

The widget

Evaluate the pair, thirty times over
Each of the 30 rows is one evaluation of N trials per checkpoint, drawn from the 2,000-trial pool of each: the previous checkpoint's 95 % interval in cyan, the new one's in amber, dashed lines at the pool's true values. The square at the right of a row is the verdict: green, the better checkpoint declared better; grey, no verdict; red, the worse one declared better. Bottom left: the chance that an evaluation concludes correctly (teal) and that its higher count belongs to the better checkpoint (dashed), against N, exact for success and resampled for progress. Bottom right: the robot time of one comparison against the training behind the two checkpoints. The first control is N; sections 5 and 6 explain the selectors.
previous checkpoint, true value
—
new checkpoint, true value
—
one 95 % interval of the new one
—
smallest gap N resolves (80 % power)
—
evaluation names the better one
—
higher count is the better one's
—
N for 80 % power, by the formula
—
robot-hours, both arms
—
robot-weeks (40 h)
—
trials per training run (40 runs)
—
of the 30 rows: right / none / wrong
—
Show the core JS
EL.zTest = function (k1, k2, n) {
  var s = k1 + k2; if (s === 0 || s === 2 * n) return 0;
  var pp = s / (2 * n), z = (k2 - k1) / n / Math.sqrt(pp * (1 - pp) * 2 / n);
  return z > EL.Z ? 1 : z < -EL.Z ? -1 : 0;
};
EL.powerIndep = function (p1, p2, n) {
  var w1 = EL.win(n, p1), w2 = EL.win(n, p2), a = EL.pmf(n, p1), b = EL.pmf(n, p2), o = { right: 0, wrong: 0, more: 0, tie: 0, less: 0 }, i, j;
  for (i = w1[0]; i <= w1[1]; i++) for (j = w2[0]; j <= w2[1]; j++) {
    var w = a[i] * b[j], v = EL.zTest(i, j, n);
    if (v > 0) o.right += w; else if (v < 0) o.wrong += w;
    if (j > i) o.more += w; else if (j === i) o.tie += w; else o.less += w;
  }
  return o;
};
EL.nIndep = function (p1, p2) { var pb = (p1 + p2) / 2; return Math.ceil(2 * Math.pow(EL.Z + EL.ZB, 2) * pb * (1 - pb) / Math.pow(p2 - p1, 2)); };
EL.mdd = function (n, pbar) { return (EL.Z + EL.ZB) * Math.sqrt(2 * pbar * (1 - pbar) / n); };

What to try. Leave the defaults: 20 trials per checkpoint, the previous checkpoint a table, independent trials, success, 3 minutes a trial. The true rates read 68.0 and 76.0 %. One 95 % interval is 35.3 points wide and the smallest gap 20 trials resolve is 39.8 points. The new checkpoint has the higher count in 65.2 % of evaluations and an evaluation names it the better one in 8.4 %; count the green squares among the 30 rows, press the button for 30 more and count again: the rows move, the probability does not. Slide N to 100: the interval is 16.5 points wide, the resolvable gap 17.8 and the right verdict 24.2 %; to 200, 43.1 %. At 500 the right verdict is 80.6 %, the cost 50.0 robot-hours or 1.25 robot-weeks, 25.0 trials for every training run, and the formula for 80 % power reads 495. At 1,000 it is 97.9 % and 100.0 hours. Set the minutes per trial to 1.7, AutoEval's pace, and the 500-trial evaluation takes 28.3 hours.

5 · Buying trials back

Pairing. Run both checkpoints from the same starting conditions and the outcomes come in pairs: both finish, both fail, only the new one finishes (b pairs), only the previous one (c pairs). Only the discordant pairs say which is better (McNemar, 1947). If the checkpoints are equal, b is binomial with b + c trials and probability ½, and the exact test rejects when the smaller count falls in the 5 % tail. Twenty pairs with b = 8 and c = 2 give p = 2 × (1 + 10 + 45)/1024 = 0.109: not significant, though the new checkpoint won 8 of the 10 discordant pairs. With ψ = (b + c)/n the share of discordant pairs and Δ = (b − c)/n, the pairs needed are

n = ( 1.96 √ψ + 0.8416 √(ψ − Δ²) )² / Δ²

For independent outcomes ψ = pnew(1 − pprev) + pprev(1 − pnew) = 40.6 %, and the formula gives 497, within half a percent of §3's 495. Pairing helps as far as ψ is smaller, which it is when the two outcomes share their cause. A robot can match the starting condition and cannot replay the disturbance; a simulator can replay both. Three designs on two pairs, the previous table against the new one and a smooth fit (lesson 2's road not taken, also 68.0 %, §6) against the new one, need this many pairs for 80 % power:

designprevious table to new: discordant sharepairssmooth fit to new: discordant sharepairs
independent trials40.6 %49540.6 %495
matched starts35.9 %43841.6 %508
same starts and same gusts8.0 %9634.6 %422

Matched starts buy little. The start offset decides some of the tables' trials, because their demonstrations all began at the marked point: the new table finishes 52 % of the fifth of trials that start lowest and 87 % of the middle fifth, the smooth fit 72 % and 70 %. Matching starts trims the tables' bill by 11.5 % and the cross-recipe pair's by nothing. Replaying the gusts too turns two similar checkpoints into near copies: the new one never loses a trial the previous one wins (0.0 % of pairs), 8.0 % of pairs are discordant, and 96 pairs suffice, 5.2 times fewer. Different recipes share little even then. Pairing buys what the pair shares, and it needs enough discordant pairs: at ψ = 8 %, twenty pairs hold 1.6 on average and the exact test needs at least 6, so 20 same-gust pairs reach the right verdict in 0.4 % of evaluations. In the widget, 100 same-gust pairs reach it in 82.0 %, and 500 pairs with matched starts in 83.4 % against 80.6 % for independent trials.

Road not taken · stop as soon as it looks significant
Trials are expensive, so look after every pair of trials and stop when the test first rejects. The 5 % error rate belongs to one look at a number of trials fixed in advance; every extra look is another chance. Two identical checkpoints at 72 %, tested after every pair of trials from the 10th to the 200th, are declared different in 33.3 % of evaluations; looking every 10 pairs gives 24.4 %, and a single look at 200 gives 5.0 % (exact, by dynamic programming over the counts). A sequential test is built to keep its error rate through the looks: Wald's (1945) probability ratio test, or for two policies STEP (Snyder et al., 2025), which reports up to 32 % fewer trials than the best earlier sequential tests at the same error rates. That trims the bill by a fraction and does not change how it grows with the gap, and it needs the stopping rule written down before the first trial.

6 · The metric is part of the question

Success uses one bit of each trial. A score that gives credit for getting part of the way, here the share of the course length reached before the run ended, uses more. π0 reports a normalised task-progress score instead of success, and RoboArena (Atreya et al., 2025) asks evaluators for a progress score and a preference on every A/B episode and calls the two rankings complementary. But a graded score can order two checkpoints the other way. The third checkpoint is the smooth fit, which finishes 68.0 % of the pool, like the previous table:

checkpointfinishesprogress, mean (sd)trials ending at post 1post 2post 3post 4post 5
previous table68.0 %81.8 (30.0)13.6 %4.84.87.01.9
new table76.0 %85.9 (27.8)11.4 %2.93.94.81.0
smooth fit68.0 %90.8 (14.8)0.6 %0.12.922.46.0

The tables fail mostly at the first post, when the start is low: of the new table's trials in the lowest fifth of start offsets (0.8 to 3.7 cm below the path) 40 % end there, against 3 % in the middle fifth; the smooth fit fails at the fourth. Success counts a failure at post 1 and one at post 4 the same; progress does not, so the new table against the smooth fit is a reversal: success ranks the new table first by 8.0 points, progress ranks the smooth fit first by 4.9. Progress is the area under the curve of how many runs are still going at each point of the course, success is the curve's last value, and curves that cross are ranked differently by the two. Neither is wrong; they answer different questions, and what the checkpoint is for fixes the right one, before the trials. The metric also sets the price: for the previous table against the new one, progress needs 797 trials per checkpoint (a gap of 4.1 points against a standard deviation near 30) where success needs 495; for the smooth fit against the new table it needs 329. In the widget, set the metric to progress with each previous checkpoint in turn.

Road not taken · evaluate somewhere cheaper
A simulated trial costs seconds, so the eight-point question costs 990 cheap trials. SIMPLER (Li et al., 2024) builds simulated environments for common real setups; its success rates track real ones with Pearson r = 0.890 and a mean maximum rank violation of 0.014 on BridgeData V2 (per task, averaged), and rank six Google-robot checkpoints far better than their validation action error does, the score lesson 1 distrusted. Laboratories can also pool trials: RoboArena collected 612 pairwise real-robot comparisons across seven universities for seven policies, found its ranking settled within about 100 comparisons for that pool, and agreed better with the all-policies ranking than the conventional protocol of 17 fixed tasks and 44 episodes per policy. What breaks is that a proxy is itself a measurement: its agreement with the real ranking has to be established on real trials, and a simulator that replays gusts (5.2 times fewer trials for similar checkpoints, §5) does not promise that the replayed world ranks policies as the real one does (lesson 11).
What this lesson did not do
It compared two checkpoints on one task, one metric at a time, and treated a 2,000-trial pool as the truth; a real success rate drifts as the robot wears and the scene moves, so an evaluation result has a shelf life, which lesson 19 takes up for data. It gave one comparison, not a suite: k checkpoints make k(k − 1)/2 pairwise tests, each needing a stricter level, which the TRI LBM Team (2025) apply as a Bonferroni correction, and which raises the trials; across tasks plotted together they say the error is not controlled at all. It measured the randomness of one fixed checkpoint, not of training: a second training seed is a second checkpoint, and the TRI LBM Team say their intervals cover only the first. It did not implement a sequential test, preference-based evaluation or a learned judge, and it did not say where the robot-hours should go instead; that is the next lesson.

Common mistakes / failure modes

"15 of 20 is 75 %, so the policy is 75 % reliable"
The Wilson interval is 53.1 to 88.8 % (§2).
"the new checkpoint scored higher, so it is better"
For a checkpoint truly eight points better, 20 trials leave it not ahead in 34.8 % of evaluations (§1).
"the intervals overlap, so there is no difference"
Non-overlap is the stricter rule: 820 trials per checkpoint where the z-test needs 495 (§3).
"stop as soon as it is significant"
Looking after every pair from the 10th to the 200th calls two identical checkpoints different 33.3 % of the time (§5).
"pair the trials and the bill shrinks"
By what the pair shares: 5.2 times for similar checkpoints with replayed gusts, 11.5 % from matched starts, nothing between recipes (§5).
"a graded score is a cheaper measure of the same thing"
It ranks the smooth fit above the new table by 4.9 points where success ranks the table above by 8.0; one pair needs 797 trials against 495 (§6).
"evaluation is overhead on the training"
One comparison is 24.8 times the robot time of the training it judges; a six-variant sweep is 7.4 robot-weeks (§4).

Checkpoint exercise

Try it
A lab expects a new recipe to lift a task from 55 % to 65 % and wants an evaluation that finds a lift of that size 80 times in 100, at the 5 % level. (a) How many trials per policy? (b) At 3 minutes a trial, how many robot-hours and robot-weeks does the comparison take? (c) The lab can afford only 30 trials per policy: what is the smallest lift it can resolve? Answer: (a) p̄ = 0.6, so n = 15.7 × 0.6 × 0.4 / 0.10² = 377 per policy. (b) 2 × 377 = 754 trials, 37.7 robot-hours, 0.94 robot-weeks. (c) Δmin = 2.8016 × √(2 × 0.6 × 0.4 / 30) = 35.4 points.

Where this points next

A success rate is a count, and on the Bench twenty trials per checkpoint name the better of two checkpoints eight points apart in 8.4 % of evaluations; a ten-point gap needs 317 trials per checkpoint and our pair 495. At three minutes a trial that is 49.5 robot-hours for one comparison, 24.8 times the robot time behind the two checkpoints, and 7.4 robot-weeks for a sweep of six variants that train on data already collected. Pairing cut the bill by a factor of one to five on the Bench, a graded score cut it or raised it, a sequential test trims a fraction, and none removes it. A programme that evaluates honestly therefore spends more robot time on evaluation than on training, so robot time is the scarce thing and the data the model learns from compete with the evaluation for it. Those data come in several kinds, our own robot's demonstrations, other bodies, video, simulation, each with a price, a usefulness and a shelf life, and an evaluation result has a shelf life too: a robot that wears and a scene that moves make last month's score of the incumbent useless for this month's comparison. How much of each kind do we need, in what unit can they even be compared, and what does each cost?

Takeaway
A success rate is a count of successes in n trials, so it has a spread that falls only as 1/√n, and an interval puts the count in its place: 15 of 20 is anything from 53.1 to 88.8 %. To tell two checkpoints apart the trials per checkpoint must be 15.7 p̄(1 − p̄)/Δ², which is 495 for eight points at 72 % and 317 for ten; twenty trials reach the right verdict on our pair in 8.4 % of evaluations. At 3 minutes a trial that is 49.5 robot-hours for one comparison, 24.8 times the robot time behind both checkpoints, and 7.4 robot-weeks for a six-variant sweep. Pairing buys what the pair shares (5.2 times from replayed gusts on similar checkpoints, 11.5 % from matched starts, nothing between recipes), looking early lifts false positives to 33.3 %, and a progress score can rank two checkpoints the other way from success, at a different price. Choose the metric and the number of trials before the first trial, and write the bill down.

Interview prompts

Companion reads: World Models · 16 The exam we never stopped taking (which cheap exam ranks policies as the decision does), World Models · 30 Evaluating the thing you trained (why measured and delivered improvement come apart) and Search, Ads and Recsys · 33 A/B testing deep dive (randomisation, power and the launch decision).