What braking changes
Lesson 10 left a detector with an exam number and a job the exam does not describe. The exam grades a frame; the shuttle lives in a loop, seeing ten frames a second, deciding, braking, and ending in a collision or a margin of metres. This lesson builds that loop from the exam's own pieces, shows that a log of human driving cannot say what braking would have changed, and then does what only a program can: replays one step-out with and without the brake from the same random numbers, so the difference is exact for each situation and quiet across a campaign. It then moves one assumption of the program's world to watch the answers move. It cannot say whether the program's pedestrian walks like a real one.
New idea: a program can run both worlds of a decision from one set of random numbers, and the difference between them is the decision's consequence. The natural detector's policy cuts the collisions of the step-outs it meets from 55.4% to 14.6%, and sharing the random numbers lets the difference between two detectors be read from 7.2 times fewer episodes.
Forces next: Because a program can branch, the same step-out can be replayed with and without braking, and with the random numbers shared between the branches the difference between two detectors needs about 7 times fewer runs than two independent campaigns; the stopping margin then separates two detectors that 1,000 episodes of collisions cannot (2.9 standard errors against 1.6), while the exam, which grades frames, cannot see the decision rule that moves one detector's collisions from 9.3% to 20.0%. Every number so far is still a statement about the program, and the program is the thing under test: change only the pedestrian's walking speed, keeping its mean, and the closest band's collision rate moves by 19 points. What would count as proof that the system is safe on the real street, and how much real driving does that proof cost?
1 · From a frame to an episode
The exam asks one question of one frame: is the pedestrian found, at a threshold that lets a tenth of the pedestrian-free frames alarm? For the stored detector of Lesson 6's exact-stage program (the street's scene, light and camera; seed 1) the miss rate is 46.7% over the exam's pedestrians and 80.3% over those who step out from behind the van. The shuttle asks a different question: it drives, the camera records ten frames a second, a rule turns scores into a decision to brake, the brake takes time to act, and the outcome is a collision with a speed, or a stop with a margin in metres. One number per frame cannot say how that ends.
An episode is a step-out followed until the shuttle's front reaches the pedestrian or the shuttle stops. The pedestrian starts hidden at the van's shadow edge, as in Lesson 6's forced step-out, at depth z, and walks toward the lane at speed u. The camera sits at the shuttle's front and moves at speed v; the van and the clutter stand still, so frame k is the street seen from v·tk metres further on, with the pedestrian u·tk metres further left, the light unchanged and a fresh draw of sensor noise. The detector scores each frame, an alarm is a score above the exam's threshold, and the policy brakes once two of the last three frames have alarmed. The rest is mechanics, and these are the lab's assumptions, not facts about any vehicle: v = 7 m/s (25 km/h); a deceleration a = 3 m/s² once the brake acts (gentle: the UK Highway Code's stopping-distance table implies about 6.5 m/s² for a car); a latency τ = 0.4 s from decision to brake; a shuttle 2 m wide; a pedestrian who walks at u = 1.4 m/s. A collision is the front reaching the pedestrian's near surface while the pedestrian is within a metre, plus their radius, of the centre line.
The stopping distance has two terms, the distance covered before the brake acts and the braking distance (from v² = 2·a·s):
d = v·τ + v² / (2a) = 2.80 + 8.17 = 10.97 m
A pedestrian at depth z (near surface z − r, r about 0.25 m) can be stopped short of only if the brake is commanded by the deadline t* = (z − r − d) / v. At 10 m it is −0.17 s: no decision, by any detector, stops the shuttle short of the pedestrian. At 14, 22 and 30 m the deadlines are 0.40, 1.54 and 2.68 s after the step-out. The speed was chosen so that this arithmetic bites where Lesson 6's rare case lives: over the first 400 natural step-outs a shuttle that never brakes collides in 20.3%, 54.5% and 69.8% at 5, 7 and 8 m/s (stopping distances 6.2, 11.0 and 13.9 m), and the natural detector's policy leaves 0.3%, 15.0% and 39.0%. At 7 m/s the stopping distance sits at the edge of Lesson 6's closest step-outs, those nearer than 12 m. Stratify the step-outs by distance in the bands of Lessons 5 and 6 (12, 16 and 22 m), and take the natural detector through 1,000 of them, 250 per band, each band weighted by how often the street draws it (Lesson 6's p/q again):
| Step-out distance | Share of the street's step-outs | Mean deadline | The detector decides at | Collisions with the policy |
|---|---|---|---|---|
| closer than 12 m (case R) | 5.9% | −0.06 s | 0.31 s | 97.6% |
| 12 to 16 m | 20.6% | 0.42 s | 0.41 s | 38.8% |
| 16 to 22 m | 45.6% | 1.13 s | 0.69 s | 2.0% |
| 22 m and beyond | 27.9% | 1.94 s | 1.04 s | 0.0% |
The exam said the detector misses 46.7% of pedestrians, frame by frame. The loop says 14.6% of the step-outs the street produces end in a collision, and shows where. In the closest band the deadline is below zero and nothing can be done; in the next the detector decides at about the deadline (a slack, deadline minus decision, of 0.01 s) and 38.8% still collide; from 16 m on the slack is 0.44 s or more and the collisions fall away. What decides is when the detector decides, which no frame carries.
One choice remains, the threshold. The exam's lets one pedestrian-free frame in ten alarm, an alarm every second at ten frames a second; two alarms among three frames turn that into a decision on 29% of 3.5-second pedestrian-free approaches (300, drawn the same way), 5.9 false brakes a minute, and 15.1% of 1,000 natural step-outs still collide (against 14.6% above: two samples of one rate, whose standard error at 1,000 episodes is 1.1 points). A stricter threshold brakes falsely less and collides more: at 3% false alarms per frame, 2.2 false brakes a minute and 23% collisions; at 1%, 0.3 a minute, once in 208 s, and 38% collisions, against 55% if the shuttle never braked. No setting is a shuttle anyone would ride; the detector of this lab is a teaching detector, and the loop is how that shows. We keep the exam's threshold and the two-of-three rule so that every number of the series stays comparable, and say what they cost. Every detector below gets its own exam threshold, the same false-alarm rate, so none of them buys safety by braking more.
2 · What a log of human driving contains
The tempting route to a consequence is the data a fleet already holds: camera frames of human-driven shuttles, what the driver did, how the encounter ended. Compare the collisions of the logged step-outs in which the driver braked with those in which the driver did not. The lab can build such a log, and so can say what is wrong with it. Its driver has a personal threshold on how big a step-out must look to register, between 6 and 27 px² of silhouette (27 px² is the smallest silhouette of any step-out nearer than 12 m, so everyone sees those). A driver who registers the pedestrian brakes 0.7 s after the first frame in which 6 px² of the pedestrian show, the exam's own rule for a pedestrian who counts (the UK Highway Code's stopping-distance table implies a thinking time of about 0.67 s); one who does not, coasts. Who brakes depends on how big the step-out looks, that on how far away it is, and how far away it is also decides how it ends. That is confounding.
A pencil version first. Near step-outs are 30% of the street, far ones 70%; drivers brake in 90% of the near and 20% of the far. A collision follows with probability 0.5 (braked, near), 0.9 (unbraked, near), 0.02 (braked, far) and 0.10 (unbraked, far). Pooled, braked episodes collide in 33.6% and unbraked ones in 14.1%: braking looks harmful, by +19.5 points. Inside each kind it changes collisions by −40 points near and −8 far, and by −17.6 over the street: the pooled contrast compared near with far.
Now the lab's log: 4,000 natural step-outs driven by this driver, each with its true outcome and, because the program can also run the road not taken, the outcome of the other action.
| Step-out distance | Braked | Collisions: braked | not braked | The log's contrast, points | The program's, points |
|---|---|---|---|---|---|
| closer than 12 m | 100% | 100% | none to compare | none | +2 |
| 12 to 16 m | 84% | 35% | 91% | −56 | −65 |
| 16 to 22 m | 39% | 0% | 59% | −59 | −61 |
| 22 m and beyond | 13% | 0% | 6% | −6 | −7 |
| all, pooled | 44% | 26% | 37% | −10.5 | −42.3 |
The program's column is the effect of everyone braking this way against nobody braking. Three things go wrong with the log. The pooled contrast is 4 times too small (−10.5 points against −42.3): the unbraked episodes are the far ones, where fewer collisions happen whatever anyone does. Stratifying by distance repairs it only where both kinds of episode exist (−56 and −59 points against −65 and −61), and only because the drivers' choices here depend on nothing but how big the step-out looks; a project cannot check that, since a log never says what else the driver was looking at. For the closest step-outs there is nothing to repair. Every driver brakes there, so the log holds no episode without braking, and no adjustment can invent one: the probability of the other action is zero (Rosenbaum and Rubin, 1983, require it strictly between 0 and 1 at every value of the covariates; Lesson 6's point again, no weight repairs a case with probability zero). The program shows what the log cannot: braking changed collisions there by +2 points, since the slower shuttle arrives as some pedestrians reach the lane, and cut the impact speed from 6.8 to 4.3 m/s.
3 · Running both branches
The program holds what the log lacks: the situation. A seed fixes the scene, the light, the pedestrian's speed and the noise of every frame (the determinism of Lesson 9's replay), so the program can run one situation twice, differing in one thing, the action at the decision: the shuttle brakes as the policy decides, or never brakes. The decision is final, so no frame after it is rendered and each branch's outcome is the kinematics of §1 in closed form. Subtract the outcomes and the result is a number for this situation, d = (collision with braking) − (collision without): −1 if braking saved the pedestrian, 0 if it changed nothing, +1 if it caused the collision. Averaged band by band over the street's step-outs it is the effect of braking. For the natural detector:
| Step-out distance | Collisions: never braking | with the policy | Saved | Caused |
|---|---|---|---|---|
| closer than 12 m | 97.6% | 97.6% | 1.6% | 1.6% |
| 12 to 16 m | 95.6% | 38.8% | 58.0% | 1.2% |
| 16 to 22 m | 63.2% | 2.0% | 61.2% | 0.0% |
| 22 m and beyond | 4.0% | 0.0% | 4.0% | 0.0% |
| the street | 55.4% | 14.6% | 41.1% | 0.3% |
Braking prevents 41.1 collisions in a hundred step-outs and causes 0.3. In the closest band it prevents and causes as many (1.6% each): the shuttle that brakes arrives later, so a few pedestrians who would have been struck are clear and a few who would have been clear are in the way. What braking buys there is speed, 6.8 to 3.9 m/s at impact, 65% less kinetic energy.
Lesson 6 left a number open: how much more a miss in the closest band should count than a miss elsewhere (its c, which the exam sets to 1). The loop prices it. In collisions avoided per step-out the policy buys 0.00 in the closest band and a mean of 0.43 in the others, so c = 0.0; in kinetic energy at impact (v²/2 per kilogram, a plain proxy for severity) it buys 16 J/kg against 12, so c = 1.3. The geometry sets a floor: the ideal detector, which decides the moment 6 px² of the pedestrian show, still collides in 89.6% of the closest band and in none beyond (5.3% over the street), and of the 9.4 points between it and the natural detector 85% are in the band from 12 to 16 m, where the mean deadline is 0.42 s and the detector decides at its edge.
4 · Sharing the random numbers
The detectors of Lesson 6 differ in their training sets and nothing else. Does the oversampled one collide less than the natural one? Estimate each rate from n episodes, Xi and Yi being 1 for a collision. The variance of the difference of the two means follows from the definition of variance (the identity behind common random numbers, analysed by Glasserman and Yao, 1992):
Var(X̄ − Ȳ) = [ Var X + Var Y − 2·Cov(X, Y) ] / n
Two separate campaigns, on different situations, have Cov = 0. If both detectors are run on the same situations, a hard step-out is hard for both, the outcomes move together, Cov > 0, and the variance falls by the factor k = (Var X + Var Y) / (Var X + Var Y − 2 Cov), which for equal variances is 1/(1 − ρ), with ρ the correlation of the outcomes. What is shared is the situation (scene, light, walking speed) and the sensor noise of frame k, drawn from a stream keyed by the seed and the frame index, not consumed in order: with one stream per episode, a branch that stops early would shift every later draw and the two branches would be strangers after the first difference. Measured on the 1,000 natural situations:
| Pair, same situations (difference to resolve) | Correlation of outcomes | Variance ratio k | Episodes needed (95%): separate campaigns | shared numbers |
|---|---|---|---|---|
| braking against never braking, natural detector (10 points) | 0.38 | 1.6 | 145 | 93 |
| natural against oversampled (25%) detector (2 points) | 0.86 | 7.2 | 2,417 | 336 |
| natural against naive detector (2 points) | 0.47 | 1.8 | 3,461 | 1,905 |
The episodes needed solve 1.96·√(Var of one episode's difference / n) = the stated difference. Sharing pays in proportion to how alike the arms are. The natural and the oversampled detector end differently in only 3.5% of the situations (ρ = 0.86), so the shared comparison needs 7.2 times fewer episodes (a bootstrap over the situations gives 5.7: k has a sampling error of its own); a detector against a very different one, or an action against none, share little (k = 1.8 and 1.6). The identity also says where sharing fails: it hurts when the arms respond to the shared randomness in opposite directions, Cov < 0, which none of these pairs does.
What to try. (1) Set the distance to 10 m: the detector decides at 0.5 s against a deadline of −0.16 s, the margin is −4.6 m, and braking only cuts the impact from 7.0 to 5.3 m/s (this pedestrian's radius is under §1's 0.25 m). (2) At 14 m, the opening view, it decides at 0.4 s against a deadline of 0.39 s: a margin of −0.1 m and a collision at 0.6 m/s, where never braking hits at 7.0. (3) At 20 m the natural detector decides at 0.1 s and stops 8.1 m short, the naive one at 0.6 s and 4.6 m short; never braking passes clear, since the pedestrian has crossed. (4) Natural against oversampled, box ticked then unticked: the spread of the 20-episode estimate goes from 4.3 to 11.3 points, a variance ratio of 7.0; braking against never braking gives 11.0 and 13.3 (1.5), natural against naive 9.8 and 13.2 (1.8), near §4's k. The dashed line is the value over all 1,000: +0.7, −39.8 and −21.4 points.
5 · The stopping margin, and what the exam cannot see
The stopping margin of a decision at time td is M = z − r − v·td − d = v·(t* − td): the room left, in metres, when the shuttle has stopped, negative if it could not. It does not depend on how the pedestrian walks, so it ranks detectors in a unit the stopping distance understands. Five detectors, each at its own exam threshold, on the same 1,000 situations:
| Detector (trained on) | Exam: frames missed | Decides before the arrival | Mean stopping margin | Step-outs that collide |
|---|---|---|---|---|
| natural (street scene, light, camera) | 46.7% | 98% | 2.8 m | 14.6% |
| oversampled 25% on case R (Lesson 6) | 53.1% | 96% | 2.6 m | 15.7% |
| oversampled 50% | 59.6% | 94% | 1.9 m | 17.5% |
| program scene, street light and camera | 62.1% | 95% | 1.2 m | 20.9% |
| the naive program | 80.5% | 73% | −2.8 m | 35.5% |
Over all eleven detectors (the eight programs of Lesson 2 and the three oversampled ones of Lesson 6) the rank correlation (Spearman, 1904) with the collision rate is 0.97 for the exam, −0.92 for the share of step-outs decided before the arrival and −0.98 for the mean margin. The measures part where detectors are close, and where no frame measure can see.
Close detectors. The oversampled detectors differ from the natural one only in how much of their training went to case R; the exam charges the 25% and 50% ones 6.4 and 12.8 points of frame misses. Paired on the same situations, the 25% detector's collisions differ from the natural detector's by +1.0 ± 0.7 points (1.6 standard errors) and its mean margin by −0.27 ± 0.09 m (−2.9); the 50% one's by +2.9 ± 0.9 points (3.3) and −0.92 ± 0.12 m (−7.5). A thousand episodes of collisions cannot separate the 25% detector from the natural one and the margin does: a collision is a coin, the margin a length, and the same episodes carry more in a length.
The decision rule. The exam is a number per frame, so it cannot depend on the rule that turns frames into a decision. One detector, one threshold, one exam number (46.7%), three rules: brake on one alarm, collisions 9.3% with 10.3 false brakes a minute; two of three, 15.1% and 5.9; three of five, 20.0% and 4.3. The same detector spans 10.7 points of collisions on a choice the exam does not see.
Runs of misses. Over the 11,950 frames of the 1,000 step-outs in which the pedestrian counts (6 px² showing) the detector misses 39.7%. A frame it misses is followed by another miss 72.3% of the time, against 10.0% after a frame it finds; the first three such frames are all misses in 48.3% of the step-outs, where independent misses at the same rate would give 6.3%. What makes a frame hard, how much of the pedestrian is hidden and how far away they are, persists to the next frame, so two alarms among three arrive later than independence promises, and no test set of single frames holds the neighbours that would show it. Codevilla and colleagues (2018) report that offline error is not necessarily correlated with driving quality.
6 · The same detector in two worlds
Every number so far belongs to a world the lab wrote: a pedestrian who walks at exactly 1.4 m/s, a shuttle with the numbers of §1. The street a project will drive is not that world, and only the lab can run both (a lab privilege: a project can observe the street only by driving it). Keep the detector, the shuttle, the scenes, the light, the camera and the seeds, and change one assumption: let the pedestrian's speed vary from step-out to step-out, uniform between 0.6 and 2.2 m/s, which keeps the 1.4 m/s mean. The same 3,000 situations (750 per band), the natural detector, shared numbers:
| Step-out distance | Collisions in the program's world | in the street's | Difference ± standard error, points |
|---|---|---|---|
| closer than 12 m | 98.0% | 78.8% | −19.2 ± 1.5 |
| 12 to 16 m | 38.4% | 29.3% | −9.1 ± 1.7 |
| 16 to 22 m | 4.1% | 4.8% | +0.7 ± 0.9 |
| 22 m and beyond | 0.0% | 2.8% | +2.8 ± 0.6 |
Weighted over the street the two worlds differ by −1.9 ± 0.6 points: small, and made of large parts that cancel. With one walking speed the pedestrian is in the lane in a sharp window; with a spread, the slow ones are still in the lane when a shuttle arrives from far (the farthest band's collisions without braking rise from 4.4% to 33.7%) and the fast ones have left it before one arrives from near (the closest band's fall from 97.5% to 81.3%). The loop was exact, and its answer for the closest band moved by 19.2 points when one assumption changed and the mean did not.
What would it take to see that on the street? Telling 13.6% from 15.6% at 95% confidence takes about 1,238 step-outs; telling the closest band's 78.8% from 98.0% takes 17 step-outs in that band, which at 5.9% of step-outs is about 296 in all. These are counts of step-outs; how much driving that is, no lab can say.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
This lesson gave the program a closed loop whose answers are exact for each situation and quiet across a campaign: 14.6% of the natural step-outs end in a collision with the policy against 55.4% without braking, and a comparison of two detectors that needs 2,417 independent episodes needs 336 with shared numbers. They are answers about the program. The same detector collides in 98.0% of the closest step-outs in the program's world and 78.8% in a street whose pedestrians walk at speeds between 0.6 and 2.2 m/s with the same mean; the lab can run both, a project can run only the street, and a difference of 1.9 points overall would take about 1,238 step-outs to see. What would count as proof that the system is safe on the real street, and how much real driving does that proof cost?
Interview prompts
- A detector misses 47% of pedestrians in a frame-level test. How many step-outs does a shuttle running it hit? (§1 — The exam cannot say: collisions depend on when the decision rule fires against a deadline; here 14.6% of natural step-outs collide.)
- Logged human drives show more collisions among braked episodes than unbraked ones. Does braking hurt? (§2 — Not from that contrast: drivers brake where the pedestrian looks big, which is where collisions are likely anyway; compare within strata and check that both actions occur in each.)
- How do you measure the effect of an action in a simulator, where a log cannot? (§3 — Run the same situation twice from the same random numbers, differing only in the action, and average the differences within strata weighted by the street's frequencies.)
- Two detectors differ by 2 points of collision rate. How many episodes resolve that, and what changes if they share random numbers? (§4 — Var(X̄ − Ȳ) = (Var X + Var Y − 2 Cov)/n; at a correlation of 0.86 a shared comparison needs 336 episodes instead of 2,417.)
- Why can a per-frame exam not rank two detectors for a closed-loop job? (§5 — It sees neither the decision rule nor the runs of misses; the margin, a length not a coin, separates detectors that collision counts cannot.)
- Your simulator says the system collides in 15.6% of step-outs. What may you conclude about the road? (§6 — Only about the simulator: one changed assumption, with the same mean, moves the closest band by 19.2 points, and the street's own rate has to be measured on the street.)
Companion reads: Reinforcement Learning · 19 Offline RL (data that lacks the action you need) and Reinforcement Learning · 12 Importance sampling (the density ratio of Lesson 6, applied to policies).