Proof
Lesson 11 ended on numbers that describe the program, and moving one assumption of it moved a band of collision rates by 19 points. What would count as proof about the street? Only real step-outs, and how many is set by how rare the failure is: one collision in a thousand takes 2,995 failure-free ones, 29,950 hours at an assumed one per ten, and no program shortens that. A repaired program can rank candidate systems as the street does, though showing so takes about 450 real step-outs per candidate, and it can aim real tests, though that makes a proof dearer. It cannot say how to spend the bill.
New idea: the only evidence about the street's failures is the street's failures, so a proof costs real trials in proportion to how rare the failure is, and a program is worth what its rankings agree with the street's. A claim of one collision in a thousand takes 2,995 failure-free step-outs; a program that ranks twelve candidates as the street does (rank correlation 1.00) needs 5,400 real step-outs to show it.
Forces next: A factory that makes labelled frames, a ledger that prices each stage's gap, a pipeline that can be trusted, a handful of real frames to calibrate with, and an exam only the real street can grade: what remains is a budget. A claim that a step-out ends in a collision less than once in a thousand takes 2,995 failure-free real step-outs, 29,950 hours at an assumed one step-out in ten, and no program shortens that; a program with the street's light and scene ranks twelve candidate systems as the street does (rank correlation 0.99 and 1.00), and showing so takes about 450 real step-outs for each of them. Of all the real observations the series' ledger asks of the street, 85% are such step-outs, and 99% if the failure is ten times rarer, while the frames and labels stay where they were. Every real hour, rendered frame and human label has a price and a value, and the proof of safety sets how many real hours cannot be avoided. How should the budget be spent, and how is a model of the street trained when frames are the cheap part and consequences the expensive one?
1 · What would count as proof
Lesson 11's loop gives each candidate a number, the fraction psim of the program's step-outs that end in a collision. The claim a shuttle needs is another kind of object: p < ε, where p is the probability that a step-out on the street ends in a collision and ε a level fixed beforehand. The only evidence about p that is not a statement about the program is what real step-outs do. Run n of them, independently, and see f collisions. If p were ε or more, seeing none would have probability at most (1 − ε)n, and the claim stands at 95% confidence once that falls to 0.05:
(1 − ε)n = 0.05 ⟹ n = ln 0.05 / ln(1 − ε) ≈ 3/ε (−ln 0.05 = 2.996, and −ln(1 − ε) ≈ ε)
That is the rule of three (Hanley and Lippman-Hand, 1983; Jovanovic and Levy, 1997). With f collisions the same argument asks the chance of f or fewer to be at most 0.05 at p = ε, the exact binomial (Clopper-Pearson) limit; for small ε it takes 4.74/ε trials for one collision and 7.75/ε for three, the Poisson limits. Counts, and the driving they are if a step-out comes once per ten hours:
| Claim p < ε | no collision, exact (3/ε) | one collision | three collisions | hours, no collision, at 1 per 10 h |
|---|---|---|---|---|
| 1 in 10 | 29 (30) | 46 | 76 | 290 |
| 1 in 100 | 299 (300) | 473 | 773 | 2,990 |
| 1 in 1,000 | 2,995 (3,000) | 4,742 | 7,752 | 29,950 |
| 1 in 104 | 29,956 (30,000) | 47,437 | 77,535 | 299,560 |
| 1 in 105 | 299,572 (300,000) | 474,385 | 775,364 | 2,995,720 |
| 1 in 106 | 2,995,731 (3,000,000) | 4,743,863 | 7,753,655 | 29,957,310 |
The hours rest on an assumption, not a measurement: that a step-out of the kind the lab draws (a pedestrian from behind a parked van) comes once per T = 10 hours of driving. One in a thousand is then 29,950 hours, 3.4 years for one shuttle around the clock and 12 days for a hundred; at T = 1 and 100 it is 2,995 and 299,500 hours. Ten times rarer is ten times more (the table's rows). Road vehicles face the same arithmetic. Kalra and Paddock (2016) take the 2013 US rate of 1.09 fatalities per 100 million miles: 275 million failure-free miles show, at 95% confidence, a rate no worse than that, about 12.5 years for 100 vehicles driving around the clock at 25 mph. They report about 5 billion miles to show a rate 20% lower at 95% confidence (more than 11 billion with the usual 80% power); estimating the rate to within 20% takes 1.96²/0.2² = 96 fatalities, 8.8 billion miles.
The price belongs to the claim, not to the system: a perfect system and a mediocre one that shows no failure are charged alike. And a collision in the sample is dear: one raises the count 1.58 times, 4.74 against 3.00.
2 · Why the program cannot stand in for the trials
The program's estimate psim is a property of the program. Renders are free, so it can be read to any precision; what it says about the street is psim + (p − psim), and the second term belongs to the pair, program and street. To use psim one must assume that term bounded, say p ≤ k·psim. A program that reads psim = ε/10, with k = 3, gives p < ε with no real trial. But the assumption is a claim about the street, and testing it at the same confidence means certifying p ≤ 0.3ε, which takes 3/(0.3ε) = 10/ε failure-free trials (9,985 at ε = 10−3), 3.3 times the cost of proving p < ε directly. An assumption strong enough to certify ε costs 3/ε or more to test: the program moves the bill from the claim to the assumption and never lowers it.
The same arithmetic for a rate. To know p within ±δ at 95% takes n = 1.96² p(1 − p)/δ² real episodes, and the interval around p̂ − psim is as wide as the one around p̂. Take the lab's natural candidate, the detector of the street's own scene, light and camera: its collision probability on the street is 14.0% (2,000 stored episodes) and 15.4% in the best program (fresh draws of the situations: Lesson 11's 14.6% and 13.6% lie within a standard error of these). The half-width of the exact 95% interval after n real episodes, in points, averaged over 400 simulated campaigns (each draws n of the street's stored episodes), and by the formula:
| Real episodes n | 25 | 100 | 400 | 1,600 | 6,400 |
|---|---|---|---|---|---|
| simulated, exact interval | 14.6 | 7.2 | 3.5 | 1.7 | 0.86 |
| 1.96 √(p(1 − p)/n) | 13.6 | 6.8 | 3.4 | 1.7 | 0.85 |
Each factor of four in n halves the width. Telling the program's 1.4-point error from zero takes about 2,360 real episodes; confirming the program within 5, 2 and 1 points takes 185, 1,156 and 4,625. The error is made of large parts that cancel (Lesson 11). In the closest band the program says 98.5% and the street 79.8%: 18 episodes in that band see it, about 303 step-outs in all. In the farthest band the program draws 0 collisions in 275 episodes and the street's rate is 1.9%: more program episodes only repeat a zero that its single walking speed produces.
3 · What the program is for: ranking candidates
A project's questions are mostly comparisons, so an evaluator that orders candidates as the street would is worth having, and that can be measured. Take the twelve distinct detectors of Lesson 2's sixteen programs as 12 candidate systems (the four programs with the naive scene repeat the detector of their label-rule twin). A candidate is a detector with its exam threshold, the one that lets a tenth of the street's pedestrian-free frames alarm, which a project reads off free unlabelled logs (Lesson 10). Run each through Lesson 11's loop (two alarms in three frames brake; 7 m/s, 3 m/s²) on natural step-outs in five worlds: the naive program, the programs of Lessons 3, 4 and 5 (camera, then light, then scene repaired; each with a walking speed of 1.4 m/s) and the street, a lab privilege no project has. Each program meets 1,000 situations per candidate and the street 2,000; within a world every candidate meets the same ones. Rankings are compared by Spearman's ρ (1904), the correlation of the ranks with tied values sharing their mean, and by Kendall's τ (1938), the surplus of concordant over discordant pairs, corrected for ties; their scales differ for the same two rankings (τ is the smaller here), so both are given.
| Evaluator (what it costs a project) | ρ | τ | Street's collision rate of its best candidate | Mean absolute gap in collision rate, points |
|---|---|---|---|---|
| Frame exam (the lab's 3,000 real labelled test frames) | 0.98 | 0.94 | 14.0% | — |
| Naive program (free logs set the thresholds) | 0.85 | 0.71 | 21.0% | 21.7 |
| + camera (a bench kit) | 0.81 | 0.63 | 21.0% | 11.4 |
| + light (1,000 unlabelled logs) | 0.99 | 0.97 | 14.0% | 13.3 |
| + scene (outlines on labelled frames) | 1.00 | 1.00 | 14.0% | 1.1 |
On the street the best candidate collides in 14.0% of step-outs, the worst in 34.8%. Three readings. First, the street's thresholds carry the street into every world, so even the naive program orders the candidates at ρ = 0.85 (spread 0.03 when both sides' episodes are resampled); but its best pick collides in 21.0% on the street, a regret of 7.0 points. The camera alone does not lift it (0.81, spread 0.04); the light does (0.99), and the scene completes it (1.00, spread 0.02). Lesson 2's price list is for the detector; for the evaluator the light is what counts. Second, rank agreement is not rate agreement: the light-repaired world ranks at 0.99 and overstates every candidate's collision rate, by 13.3 points on average, because its pedestrians stand nearer (Lesson 5); the scene-repaired world is within 1.1 points on average. Third, the frame exam ranks as well as the loop (0.98): twelve candidates spread from 46.7 to 80.5% of frames missed do not need a loop to be ordered. Codevilla and colleagues (2018) report that offline error is not necessarily correlated with driving quality, and Dauner and colleagues (2023) call short-term planning and long-horizon ego-forecasting "fundamentally misaligned"; Lesson 11 showed where the exam fails, close detectors and the decision rule, which a ranking of well-separated candidates cannot see. Leaving out the one candidate a world's own program trained moves ρ by at most 0.05.
4 · The second bill: is the evaluator right?
An evaluator is trusted when its ranking agrees with the street's, and the street's ranking has to be estimated from real episodes, one campaign per candidate (a real situation cannot be replayed with another candidate). A campaign's estimate of a collision probability has standard error √(p(1 − p)/n), 2.7 points at n = 160, so the estimated ranking is a noisy copy of the street's, and an evaluator that agrees perfectly with the street still shows less than perfect agreement with the copy. Simulate it: draw n of the street's 2,000 stored episodes for each candidate, independently, 400 times, and ask how often the evaluator's agreement with the estimated ranking exceeds 0.9.
| Real episodes per candidate | 20 | 80 | 160 | 320 | 640 | 1,280 |
|---|---|---|---|---|---|---|
| scene-repaired program: mean agreement | 0.55 | 0.80 | 0.88 | 0.93 | 0.95 | 0.97 |
| chance the agreement exceeds 0.9 | 0.02 | 0.14 | 0.42 | 0.81 | 0.96 | 1.00 |
| the same for the naive program | 0.00 | 0.02 | 0.06 | 0.07 | 0.11 | 0.15 |
| scene-repaired, ranking by stopping margin | 0.19 | 0.77 | 0.92 | 0.97 | 1.00 | 1.00 |
| chance the campaign's best is within a point of the street's best | 0.46 | 0.84 | 0.94 | 0.99 | 1.00 | 1.00 |
The program that agrees with the street at 1.00 shows a mean agreement of 0.55 after 20 real episodes per candidate and 0.95 after 640, and clears "above 0.9, with probability 0.9" at about 450 (a scan of 4,000 campaigns per setting): twelve candidates, 5,400 real step-outs. The naive program never does (0.15 at best): a real campaign rejects it from about 640 episodes, where its mean is 0.84 against 0.95. The stopping margin, a length that Lesson 11 showed to be quieter than a coin, needs only about 140 per candidate (1,680 in all) but must be measured in every real episode, decision time and distance, not only its outcome. Picking one best candidate is cheaper than ordering twelve: 0.94 at 160. What the bill buys is an evaluator that orders new candidates without a real episode, so long as they resemble the twelve; real episodes alone order the twelve at a mean agreement of 0.88 with 160 per candidate and 0.97 with 1,280. UN Regulation No. 157 likewise requires a simulation toolchain to be validated by correlating its outcomes with physical tests, and does not let simulation replace them.
What to try. (1) Leave none and T = 10 and slide to 2,560: the bound is 0.12%, the rule of three says 0.12%, and one in a thousand is not yet claimed (it takes 2,995); at 10,000 it is 0.030%, and the driving is 11.4 years of one shuttle. (2) At 320 trials the bound is 0.93% with no collision, 1.47% with one and 2.41% with three. (3) With the scene-repaired evaluator ranked by collisions, n = 160 gives a mean agreement of 0.88 and a chance of 0.42; at 640 they are 0.95 and 0.96. (4) Switch to the naive program at 640: 0.84 and 0.11, and the dots wander off the bars. (5) Rank by stopping margin at 160: the chance is 0.92, where collisions give 0.42. (6) Camera repaired: the chance never exceeds 0.03, at any n.
5 · Aiming: the program's third use
A program that predicts where collisions fall can aim real trials. Stratify the step-outs by distance, as Lessons 5, 6 and 11 did, and give each stratum its share of the street's step-outs (a million scene draws), the program's predicted collision rate for the natural candidate and the street's:
| Step-out distance | Share of step-outs (natural trials) | Program says | Street says | Share of the street's collisions | Aimed share of trials |
|---|---|---|---|---|---|
| closer than 12 m | 5.9% | 98.5% | 79.8% | 33% | 4.2% |
| 12 to 16 m | 20.6% | 32.9% | 30.9% | 45% | 44.0% |
| 16 to 22 m | 45.6% | 4.1% | 5.7% | 18% | 46.3% |
| 22 m and beyond | 27.9% | 0.0% | 1.9% | 4% | 5.6% |
The two nearest bands, 27% of step-outs, hold 78% of the collisions. Estimating the street's rate (14.2%) from n trials at the natural shares has variance p(1 − p)/n = 0.1217/n. Allocating in proportion to the strata's shares removes the variance between strata and leaves ΣPhph(1 − ph) = 0.0831, 1.46 times smaller. Neyman's allocation, nh ∝ Ph√(ph(1 − ph)), uses the predicted risks and leaves (ΣPh√(ph(1 − ph)))² = 0.0689 when the risks are the street's, 1.77 times smaller. With the program's risks it asks for no trial beyond 22 m, where the program says zero and the street's rate is 1.9%: the estimate is then blind to 4% of the collisions and its variance is unbounded (Lessons 6 and 11: no weight repairs a case with probability zero). Keeping a fraction λ = 0.2 of the natural shares, the aimed column, leaves a variance 1.44 times smaller and finds 19.6 collisions per 100 trials against 14.2. Simulated at n = 400 from the street's stored episodes, the spread of the estimate is 1.76, 1.44 and 1.48 points for natural, proportional and aimed trials; the formulas say 1.74, 1.44 and 1.45.
The proof is another matter. With no collision seen, the largest rate the trials cannot rule out is the largest ΣPhph whose chance of no failure is still above 5%. At the natural shares that is 1 − 0.051/n, the unstratified bound, and no allocation does better, since a stratum left thin can hide the whole failure budget. With the aimed shares the stratum the program calls safe gets a fifth of its natural share and the bound is 5.0 times looser: the 2,995 trials that prove one in a thousand at the natural shares prove only 0.496%, and the same proof takes 14,850. Aiming lowers the real trials needed to find failures and to estimate a rate; it cannot raise the proof, and a proof wants no thin stratum exactly where the program is confident.
6 · The ledger
The series in one table: what each repair bought, in points of the exam miss rate (Lesson 2's prices, recomputed here from the sixteen programs), and what it took from the street.
| Line | What it bought | What it took from the street |
|---|---|---|
| Camera (Lesson 3) | 16.3 points | a bench kit: 11 grey-card exposures, 1 edge shot and 30 exposure-log frames, 42 frames, one afternoon |
| Light (Lesson 4) | 8.5 points | 1,000 unlabelled logs |
| Scene (Lesson 5) | 8.4 points | outlines on the first 50 of the labelled frames below |
| Label rule, rare case, renderer, checks, loop (Lessons 6 to 9, 11) | −0.1 points for the label rule; weights, contract, trust and consequences | renders and computing only |
| Calibrate, fine-tune, grade (Lesson 10) | the detector from 80.3% to 48.8% missed, certified at most 57.3% | 400 labelled frames: 100 to train, 300 to grade |
| Validate the evaluator (this lesson) | an agreement above 0.9, shown | 12 candidates × 450 = 5,400 real step-outs |
| Prove one in a thousand (this lesson) | the bound of §1 | 2,995 real step-outs |
In the order they take from the street: 42 bench frames, 1,000 logs, 400 labelled frames, then 2,995 and 5,400 step-outs. Counting a step-out as one observation, no more than a frame, which understates it (it costs T hours of driving, a frame seconds), step-outs are 8,395 of 9,837 real observations: 85%. The validation dominates, 1.8 times the proof; the proof overtakes it for claims below one in 1,800. At ten hours a step-out the two lines are 83,950 hours, 9.6 years of a shuttle around the clock.
Now make the failure ten times rarer. The proof takes 29,956 step-outs. The validation takes more than ten times its old price: with every candidate's collision probability divided by ten, a campaign clears the same test at about 6,000 episodes per candidate, 13 times as many, 72,000 in all: the gaps between candidates shrink tenfold, so their squares a hundredfold, while the variance of each rate falls by less than ten. The frames, labels and renders do not move: 1,442. The step-outs are then 101,956 of 103,398 observations, 99%, and at ten hours each 116 years of one shuttle. The ledger's total is not a number of renders, which are cheap and unlimited here: it is real step-outs, labelled frames and human attention, and it grows with the rarity of the failure while the frames stay put.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
This lesson priced the proof and the program's place beside it. One collision in a thousand takes 2,995 failure-free real step-outs, 29,950 hours at an assumed one per ten, and no program shortens that; the program ranks twelve candidates as the street does (0.99 and 1.00 once the light and the scene are right) and showing so costs 450 real step-outs per candidate. Of everything the series' ledger asks of the street, 85% are such step-outs, 99% if the failure is ten times rarer, while the frames and labels stay at 1,442. What it cannot do is say how to spend such a budget, nor how to train a model of the street when what is cheap is not what is needed. How should the budget be spent, and how is a model of the street trained when frames are the cheap part and consequences the expensive one?
Interview prompts
- Your simulator ran ten million episodes with no collision. What can you say about the street? (§2 — Only about the simulator: using the number for the street assumes it agrees, and testing that assumption takes 3/ε or more real trials.)
- How many failure-free trials show a failure probability below 10−4, and how long is that? (§1 — 29,956, about 3/ε; at one step-out per ten hours 299,560 hours.)
- A simulator ranks your candidates perfectly but gets the rates wrong. Is it useful? (§3 — For choosing among candidates yes: the light-repaired world ranked at 0.99 with rates 13.3 points high; not for a claim about a rate.)
- How would you decide a simulator is good enough to rank systems, and what does it cost? (§4 — Correlate its ranking with real campaigns: above 0.9 with probability 0.9 takes about 450 real episodes per candidate, about 140 by stopping margin.)
- Why not put every real trial where the simulator says failures are? (§5 — Where it says zero it gets none, the estimate goes blind there, and a proof gets 5.0 times dearer; keep a floor.)
- What dominates the budget of a simulator-based program? (§6 — Real consequences: 85% of the real observations are step-outs, and the share grows as the failure gets rarer.)
Companion reads: Training a Robot Model (the data economics of robot hours, the next series), World Models (a model of consequences, and what it costs to train) and Reinforcement Learning · 19 Offline RL (evidence that lacks the action you need).