all_lessons/Synthetic Vision Data/12 · Prooflesson 12 / 12

Proof

Lesson 11 ended on numbers that describe the program, and moving one assumption of it moved a band of collision rates by 19 points. What would count as proof about the street? Only real step-outs, and how many is set by how rare the failure is: one collision in a thousand takes 2,995 failure-free ones, 29,950 hours at an assumed one per ten, and no program shortens that. A repaired program can rank candidate systems as the street does, though showing so takes about 450 real step-outs per candidate, and it can aim real tests, though that makes a proof dearer. It cannot say how to spend the bill.

The thesis, here
Proof about the street can only come from the street, and its price is set by how rare the failure is, not by anything a program can do. A program's work is to predict, to rank and to aim; it is as good an evaluator as its rankings agree with the street's, and that agreement has to be paid for in real events.
Linear position
Forced by: Because a program can branch, the same step-out can be replayed with and without braking, and with the random numbers shared between the branches the difference between two detectors needs about 7 times fewer runs than two independent campaigns; the stopping margin then separates two detectors that 1,000 episodes of collisions cannot (2.9 standard errors against 1.6), while the exam, which grades frames, cannot see the decision rule that moves one detector's collisions from 9.3% to 20.0%. Every number so far is still a statement about the program, and the program is the thing under test: change only the pedestrian's walking speed, keeping its mean, and the closest band's collision rate moves by 19 points. What would count as proof that the system is safe on the real street, and how much real driving does that proof cost?
New idea: the only evidence about the street's failures is the street's failures, so a proof costs real trials in proportion to how rare the failure is, and a program is worth what its rankings agree with the street's. A claim of one collision in a thousand takes 2,995 failure-free step-outs; a program that ranks twelve candidates as the street does (rank correlation 1.00) needs 5,400 real step-outs to show it.
Forces next: A factory that makes labelled frames, a ledger that prices each stage's gap, a pipeline that can be trusted, a handful of real frames to calibrate with, and an exam only the real street can grade: what remains is a budget. A claim that a step-out ends in a collision less than once in a thousand takes 2,995 failure-free real step-outs, 29,950 hours at an assumed one step-out in ten, and no program shortens that; a program with the street's light and scene ranks twelve candidate systems as the street does (rank correlation 0.99 and 1.00), and showing so takes about 450 real step-outs for each of them. Of all the real observations the series' ledger asks of the street, 85% are such step-outs, and 99% if the failure is ten times rarer, while the frames and labels stay where they were. Every real hour, rendered frame and human label has a price and a value, and the proof of safety sets how many real hours cannot be avoided. How should the budget be spent, and how is a model of the street trained when frames are the cheap part and consequences the expensive one?
The plan
Six moves. (1) State what would count as proof, and count the trials it takes. (2) Show that no estimate from the program shortens the count. (3) Measure the program as an evaluator: twelve candidates, five worlds. (4) Price the evidence that an evaluator is right. (5) Aim real trials with the program, and see what aiming cannot buy. (6) Add up the series' ledger, and make the failure ten times rarer.

1 · What would count as proof

Lesson 11's loop gives each candidate a number, the fraction psim of the program's step-outs that end in a collision. The claim a shuttle needs is another kind of object: p < ε, where p is the probability that a step-out on the street ends in a collision and ε a level fixed beforehand. The only evidence about p that is not a statement about the program is what real step-outs do. Run n of them, independently, and see f collisions. If p were ε or more, seeing none would have probability at most (1 − ε)n, and the claim stands at 95% confidence once that falls to 0.05:

(1 − ε)n = 0.05  ⟹  n = ln 0.05 / ln(1 − ε) ≈ 3/ε   (−ln 0.05 = 2.996, and −ln(1 − ε) ≈ ε)

That is the rule of three (Hanley and Lippman-Hand, 1983; Jovanovic and Levy, 1997). With f collisions the same argument asks the chance of f or fewer to be at most 0.05 at p = ε, the exact binomial (Clopper-Pearson) limit; for small ε it takes 4.74/ε trials for one collision and 7.75/ε for three, the Poisson limits. Counts, and the driving they are if a step-out comes once per ten hours:

Claim p < εno collision, exact (3/ε)one collisionthree collisionshours, no collision, at 1 per 10 h
1 in 1029 (30)4676290
1 in 100299 (300)4737732,990
1 in 1,0002,995 (3,000)4,7427,75229,950
1 in 10429,956 (30,000)47,43777,535299,560
1 in 105299,572 (300,000)474,385775,3642,995,720
1 in 1062,995,731 (3,000,000)4,743,8637,753,65529,957,310

The hours rest on an assumption, not a measurement: that a step-out of the kind the lab draws (a pedestrian from behind a parked van) comes once per T = 10 hours of driving. One in a thousand is then 29,950 hours, 3.4 years for one shuttle around the clock and 12 days for a hundred; at T = 1 and 100 it is 2,995 and 299,500 hours. Ten times rarer is ten times more (the table's rows). Road vehicles face the same arithmetic. Kalra and Paddock (2016) take the 2013 US rate of 1.09 fatalities per 100 million miles: 275 million failure-free miles show, at 95% confidence, a rate no worse than that, about 12.5 years for 100 vehicles driving around the clock at 25 mph. They report about 5 billion miles to show a rate 20% lower at 95% confidence (more than 11 billion with the usual 80% power); estimating the rate to within 20% takes 1.96²/0.2² = 96 fatalities, 8.8 billion miles.

The price belongs to the claim, not to the system: a perfect system and a mediocre one that shows no failure are charged alike. And a collision in the sample is dear: one raises the count 1.58 times, 4.74 against 3.00.

2 · Why the program cannot stand in for the trials

The program's estimate psim is a property of the program. Renders are free, so it can be read to any precision; what it says about the street is psim + (p − psim), and the second term belongs to the pair, program and street. To use psim one must assume that term bounded, say p ≤ k·psim. A program that reads psim = ε/10, with k = 3, gives p < ε with no real trial. But the assumption is a claim about the street, and testing it at the same confidence means certifying p ≤ 0.3ε, which takes 3/(0.3ε) = 10/ε failure-free trials (9,985 at ε = 10−3), 3.3 times the cost of proving p < ε directly. An assumption strong enough to certify ε costs 3/ε or more to test: the program moves the bill from the claim to the assumption and never lowers it.

The same arithmetic for a rate. To know p within ±δ at 95% takes n = 1.96² p(1 − p)/δ² real episodes, and the interval around p̂ − psim is as wide as the one around p̂. Take the lab's natural candidate, the detector of the street's own scene, light and camera: its collision probability on the street is 14.0% (2,000 stored episodes) and 15.4% in the best program (fresh draws of the situations: Lesson 11's 14.6% and 13.6% lie within a standard error of these). The half-width of the exact 95% interval after n real episodes, in points, averaged over 400 simulated campaigns (each draws n of the street's stored episodes), and by the formula:

Real episodes n251004001,6006,400
simulated, exact interval14.67.23.51.70.86
1.96 √(p(1 − p)/n)13.66.83.41.70.85

Each factor of four in n halves the width. Telling the program's 1.4-point error from zero takes about 2,360 real episodes; confirming the program within 5, 2 and 1 points takes 185, 1,156 and 4,625. The error is made of large parts that cancel (Lesson 11). In the closest band the program says 98.5% and the street 79.8%: 18 episodes in that band see it, about 303 step-outs in all. In the farthest band the program draws 0 collisions in 275 episodes and the street's rate is 1.9%: more program episodes only repeat a zero that its single walking speed produces.

Road not taken · multiply the program's rate by a safety factor
Claim p ≤ 10·psim and be done. A factor is the assumption k of this section, priced above. It also fails where the program says zero: ten times zero is zero, and the street's farthest band is 1.9%. And no constant fits: the best program's rates over the street's, candidate by candidate, run from 0.98 to 1.10, the naive program's from 0.40 to 2.48.

3 · What the program is for: ranking candidates

A project's questions are mostly comparisons, so an evaluator that orders candidates as the street would is worth having, and that can be measured. Take the twelve distinct detectors of Lesson 2's sixteen programs as 12 candidate systems (the four programs with the naive scene repeat the detector of their label-rule twin). A candidate is a detector with its exam threshold, the one that lets a tenth of the street's pedestrian-free frames alarm, which a project reads off free unlabelled logs (Lesson 10). Run each through Lesson 11's loop (two alarms in three frames brake; 7 m/s, 3 m/s²) on natural step-outs in five worlds: the naive program, the programs of Lessons 3, 4 and 5 (camera, then light, then scene repaired; each with a walking speed of 1.4 m/s) and the street, a lab privilege no project has. Each program meets 1,000 situations per candidate and the street 2,000; within a world every candidate meets the same ones. Rankings are compared by Spearman's ρ (1904), the correlation of the ranks with tied values sharing their mean, and by Kendall's τ (1938), the surplus of concordant over discordant pairs, corrected for ties; their scales differ for the same two rankings (τ is the smaller here), so both are given.

Evaluator (what it costs a project)ρτStreet's collision rate of its best candidateMean absolute gap in collision rate, points
Frame exam (the lab's 3,000 real labelled test frames)0.980.9414.0%—
Naive program (free logs set the thresholds)0.850.7121.0%21.7
+ camera (a bench kit)0.810.6321.0%11.4
+ light (1,000 unlabelled logs)0.990.9714.0%13.3
+ scene (outlines on labelled frames)1.001.0014.0%1.1

On the street the best candidate collides in 14.0% of step-outs, the worst in 34.8%. Three readings. First, the street's thresholds carry the street into every world, so even the naive program orders the candidates at ρ = 0.85 (spread 0.03 when both sides' episodes are resampled); but its best pick collides in 21.0% on the street, a regret of 7.0 points. The camera alone does not lift it (0.81, spread 0.04); the light does (0.99), and the scene completes it (1.00, spread 0.02). Lesson 2's price list is for the detector; for the evaluator the light is what counts. Second, rank agreement is not rate agreement: the light-repaired world ranks at 0.99 and overstates every candidate's collision rate, by 13.3 points on average, because its pedestrians stand nearer (Lesson 5); the scene-repaired world is within 1.1 points on average. Third, the frame exam ranks as well as the loop (0.98): twelve candidates spread from 46.7 to 80.5% of frames missed do not need a loop to be ordered. Codevilla and colleagues (2018) report that offline error is not necessarily correlated with driving quality, and Dauner and colleagues (2023) call short-term planning and long-horizon ego-forecasting "fundamentally misaligned"; Lesson 11 showed where the exam fails, close detectors and the decision rule, which a ranking of well-separated candidates cannot see. Leaving out the one candidate a world's own program trained moves ρ by at most 0.05.

Road not taken · set each candidate's threshold inside the program
It needs no logs: apply the exam's rule inside the world, a threshold that lets a tenth of the world's own pedestrian-free frames alarm. It is Lesson 1's flattering exam again. In the naive world every one of the twelve collides in 2.8 to 4.2% of step-outs, against the street's 14.0 to 34.8%, and their order runs against the street's: ρ = −0.48 (spread 0.16). The camera-repaired world says 3.3 to 5.3%. Thresholds read off free real logs are what carry the street into the loop.

4 · The second bill: is the evaluator right?

An evaluator is trusted when its ranking agrees with the street's, and the street's ranking has to be estimated from real episodes, one campaign per candidate (a real situation cannot be replayed with another candidate). A campaign's estimate of a collision probability has standard error √(p(1 − p)/n), 2.7 points at n = 160, so the estimated ranking is a noisy copy of the street's, and an evaluator that agrees perfectly with the street still shows less than perfect agreement with the copy. Simulate it: draw n of the street's 2,000 stored episodes for each candidate, independently, 400 times, and ask how often the evaluator's agreement with the estimated ranking exceeds 0.9.

Real episodes per candidate20801603206401,280
scene-repaired program: mean agreement0.550.800.880.930.950.97
chance the agreement exceeds 0.90.020.140.420.810.961.00
the same for the naive program0.000.020.060.070.110.15
scene-repaired, ranking by stopping margin0.190.770.920.971.001.00
chance the campaign's best is within a point of the street's best0.460.840.940.991.001.00

The program that agrees with the street at 1.00 shows a mean agreement of 0.55 after 20 real episodes per candidate and 0.95 after 640, and clears "above 0.9, with probability 0.9" at about 450 (a scan of 4,000 campaigns per setting): twelve candidates, 5,400 real step-outs. The naive program never does (0.15 at best): a real campaign rejects it from about 640 episodes, where its mean is 0.84 against 0.95. The stopping margin, a length that Lesson 11 showed to be quieter than a coin, needs only about 140 per candidate (1,680 in all) but must be measured in every real episode, decision time and distance, not only its outcome. Picking one best candidate is cheaper than ordering twelve: 0.94 at 160. What the bill buys is an evaluator that orders new candidates without a real episode, so long as they resemble the twelve; real episodes alone order the twelve at a mean agreement of 0.88 with 160 per candidate and 0.97 with 1,280. UN Regulation No. 157 likewise requires a simulation toolchain to be validated by correlating its outcomes with physical tests, and does not let simulation replace them.

The price of proof, and of trusting an evaluator
Slide n, the real trials (log scale). Top left: the exact 95% bound after n trials with the chosen number of collisions (teal; the others grey; dashed 3/n) and the claims it supports. Top right and bottom left: 400 simulated real campaigns of n episodes per candidate, drawn from the street's stored episodes, against the chosen evaluator's ranking: how its agreement (Spearman) grows with n, and its spread at this n. Bottom right: the twelve candidates, the street's collision rates (bars) and the evaluator's (dots). The bound is computed live; the episodes come from a table stored with the page, so the campaign panels stop at 2,560.
95% bound after n trials
—
rule of three, 3/n
—
real driving at the assumed rate
—
one shuttle around the clock
—
mean agreement with the campaign
—
chance the agreement exceeds 0.9
—
chance the campaign's pick is within a point of the best
—
real episodes in all, twelve candidates
—
agreement with all 2,000 street episodes
—
Show the core JS
L.logBinomCdf = function (f, n, p) {
  var lp = Math.log(p / (1 - p)), lt = n * Math.log1p(-p), s = lt, i;
  for (i = 0; i < f; i++) { lt += Math.log((n - i) / (i + 1)) + lp; s = lt > s ? lt + Math.log1p(Math.exp(s - lt)) : s + Math.log1p(Math.exp(lt - s)); }
  return s;
};
L.cpUpper = function (f, n, alpha) {
  alpha = alpha || L.ALPHA;
  if (f >= n) return 1;
  var lo = 0, hi = 1, mid, i;
  for (i = 0; i < 100; i++) { mid = (lo + hi) / 2; if (L.binomCdf(f, n, mid) > alpha) lo = mid; else hi = mid; }
  return (lo + hi) / 2;
};
L.spearman = function (a, b) { return L.pearson(L.avgRanks(a), L.avgRanks(b)); };

What to try. (1) Leave none and T = 10 and slide to 2,560: the bound is 0.12%, the rule of three says 0.12%, and one in a thousand is not yet claimed (it takes 2,995); at 10,000 it is 0.030%, and the driving is 11.4 years of one shuttle. (2) At 320 trials the bound is 0.93% with no collision, 1.47% with one and 2.41% with three. (3) With the scene-repaired evaluator ranked by collisions, n = 160 gives a mean agreement of 0.88 and a chance of 0.42; at 640 they are 0.95 and 0.96. (4) Switch to the naive program at 640: 0.84 and 0.11, and the dots wander off the bars. (5) Rank by stopping margin at 160: the chance is 0.92, where collisions give 0.42. (6) Camera repaired: the chance never exceeds 0.03, at any n.

5 · Aiming: the program's third use

A program that predicts where collisions fall can aim real trials. Stratify the step-outs by distance, as Lessons 5, 6 and 11 did, and give each stratum its share of the street's step-outs (a million scene draws), the program's predicted collision rate for the natural candidate and the street's:

Step-out distanceShare of step-outs (natural trials)Program saysStreet saysShare of the street's collisionsAimed share of trials
closer than 12 m5.9%98.5%79.8%33%4.2%
12 to 16 m20.6%32.9%30.9%45%44.0%
16 to 22 m45.6%4.1%5.7%18%46.3%
22 m and beyond27.9%0.0%1.9%4%5.6%

The two nearest bands, 27% of step-outs, hold 78% of the collisions. Estimating the street's rate (14.2%) from n trials at the natural shares has variance p(1 − p)/n = 0.1217/n. Allocating in proportion to the strata's shares removes the variance between strata and leaves ΣPhph(1 − ph) = 0.0831, 1.46 times smaller. Neyman's allocation, nh ∝ Ph√(ph(1 − ph)), uses the predicted risks and leaves (ΣPh√(ph(1 − ph)))² = 0.0689 when the risks are the street's, 1.77 times smaller. With the program's risks it asks for no trial beyond 22 m, where the program says zero and the street's rate is 1.9%: the estimate is then blind to 4% of the collisions and its variance is unbounded (Lessons 6 and 11: no weight repairs a case with probability zero). Keeping a fraction λ = 0.2 of the natural shares, the aimed column, leaves a variance 1.44 times smaller and finds 19.6 collisions per 100 trials against 14.2. Simulated at n = 400 from the street's stored episodes, the spread of the estimate is 1.76, 1.44 and 1.48 points for natural, proportional and aimed trials; the formulas say 1.74, 1.44 and 1.45.

The proof is another matter. With no collision seen, the largest rate the trials cannot rule out is the largest ΣPhph whose chance of no failure is still above 5%. At the natural shares that is 1 − 0.051/n, the unstratified bound, and no allocation does better, since a stratum left thin can hide the whole failure budget. With the aimed shares the stratum the program calls safe gets a fifth of its natural share and the bound is 5.0 times looser: the 2,995 trials that prove one in a thousand at the natural shares prove only 0.496%, and the same proof takes 14,850. Aiming lowers the real trials needed to find failures and to estimate a rate; it cannot raise the proof, and a proof wants no thin stratum exactly where the program is confident.

6 · The ledger

The series in one table: what each repair bought, in points of the exam miss rate (Lesson 2's prices, recomputed here from the sixteen programs), and what it took from the street.

LineWhat it boughtWhat it took from the street
Camera (Lesson 3)16.3 pointsa bench kit: 11 grey-card exposures, 1 edge shot and 30 exposure-log frames, 42 frames, one afternoon
Light (Lesson 4)8.5 points1,000 unlabelled logs
Scene (Lesson 5)8.4 pointsoutlines on the first 50 of the labelled frames below
Label rule, rare case, renderer, checks, loop (Lessons 6 to 9, 11)−0.1 points for the label rule; weights, contract, trust and consequencesrenders and computing only
Calibrate, fine-tune, grade (Lesson 10)the detector from 80.3% to 48.8% missed, certified at most 57.3%400 labelled frames: 100 to train, 300 to grade
Validate the evaluator (this lesson)an agreement above 0.9, shown12 candidates × 450 = 5,400 real step-outs
Prove one in a thousand (this lesson)the bound of §12,995 real step-outs

In the order they take from the street: 42 bench frames, 1,000 logs, 400 labelled frames, then 2,995 and 5,400 step-outs. Counting a step-out as one observation, no more than a frame, which understates it (it costs T hours of driving, a frame seconds), step-outs are 8,395 of 9,837 real observations: 85%. The validation dominates, 1.8 times the proof; the proof overtakes it for claims below one in 1,800. At ten hours a step-out the two lines are 83,950 hours, 9.6 years of a shuttle around the clock.

Now make the failure ten times rarer. The proof takes 29,956 step-outs. The validation takes more than ten times its old price: with every candidate's collision probability divided by ten, a campaign clears the same test at about 6,000 episodes per candidate, 13 times as many, 72,000 in all: the gaps between candidates shrink tenfold, so their squares a hundredfold, while the variance of each rate falls by less than ten. The frames, labels and renders do not move: 1,442. The step-outs are then 101,956 of 103,398 observations, 99%, and at ten hours each 116 years of one shuttle. The ledger's total is not a number of renders, which are cheap and unlimited here: it is real step-outs, labelled frames and human attention, and it grows with the rarity of the failure while the frames stay put.

What this lesson did not do
It did not choose ε, the rate of step-outs or the confidence: they are assumptions, and the tables vary them. It did not test the evaluator on candidates unlike the twelve, or on a second street: an agreement belongs to one street and one family of candidates. Real trials are not independent (the same roads, the same pedestrians), which the exact bound ignores and which makes it optimistic. The shuttle's speed, deceleration and latency are Lesson 11's assumptions, so no number here describes a vehicle, and none says the shuttle is safe. It did not say how to spend the bill (the next series does), nor how a model of the street is trained.

Common mistakes / failure modes

"The simulator ran a billion episodes without a collision, so the failure probability is a billionth."
That is psim. The street's p needs real trials, 3/ε of them, and testing the assumption that links the two costs 3.3 times more (§2).
"No collisions in 100 trials means none," or "3/n is the failure rate."
It is an upper limit: 1 − 0.051/100 is 2.95%. One collision makes it 4.74/n, three 7.75/n; one in a thousand takes 2,995 (§1).
"A good evaluator predicts the collision rates."
One that ranks the twelve candidates at 0.99 overstated every rate by 13.3 points: ranking and calibration differ (§3).
"An agreement of 0.9 on 80 real episodes per candidate validates it."
A perfect evaluator shows a mean of 0.80 there and passes with probability 0.14; it takes about 450 (§4).
"Put the trials where the program says failures are."
No trial lands where it says zero, and a proof costs 5.0 times as many (§5).

Checkpoint exercise

Try it
A shuttle runs 600 independent real step-outs. (a) With no collision, what does the exact rule say, and what does 3/n? (b) With one collision, does it support a claim of 1%? (c) At one step-out per 20 hours, how long do 3,000 step-outs take? Answer: (a) 1 − 0.051/600 = 0.498%; the rule says 0.5%. (b) The upper limit is 0.79%, below 1%: yes (with one collision a claim of 1% needs at least 473). (c) 3,000 × 20 = 60,000 hours, 6.8 years of one shuttle.

Where this points next

This lesson priced the proof and the program's place beside it. One collision in a thousand takes 2,995 failure-free real step-outs, 29,950 hours at an assumed one per ten, and no program shortens that; the program ranks twelve candidates as the street does (0.99 and 1.00 once the light and the scene are right) and showing so costs 450 real step-outs per candidate. Of everything the series' ledger asks of the street, 85% are such step-outs, 99% if the failure is ten times rarer, while the frames and labels stay at 1,442. What it cannot do is say how to spend such a budget, nor how to train a model of the street when what is cheap is not what is needed. How should the budget be spent, and how is a model of the street trained when frames are the cheap part and consequences the expensive one?

Takeaway
Proof about the street can only come from real trials, and their number is set by the rarity of the failure: n failure-free trials bound it below about 3/n, so one in a thousand costs 2,995 and ten times rarer ten times more. A program's estimate cannot shorten that, because using it means assuming it agrees with the street, and testing the assumption costs at least as much as the bound it supports. A repaired program can rank candidates as the street does, but showing so takes about 450 real episodes per candidate, and its collision rates can be 13.3 points off while its ranking is right. It can also aim real tests, which finds failures sooner and makes a proof dearer. Added up, the series' ledger is mostly real step-outs, 85% of its real observations and 99% for a failure ten times rarer; none of it is a number of renders.

Interview prompts

Companion reads: Training a Robot Model (the data economics of robot hours, the next series), World Models (a model of consequences, and what it costs to train) and Reinforcement Learning · 19 Offline RL (evidence that lacks the action you need).