all_lessons/Synthetic Vision Data/11 · What braking changeslesson 11 / 12

What braking changes

Lesson 10 left a detector with an exam number and a job the exam does not describe. The exam grades a frame; the shuttle lives in a loop, seeing ten frames a second, deciding, braking, and ending in a collision or a margin of metres. This lesson builds that loop from the exam's own pieces, shows that a log of human driving cannot say what braking would have changed, and then does what only a program can: replays one step-out with and without the brake from the same random numbers, so the difference is exact for each situation and quiet across a campaign. It then moves one assumption of the program's world to watch the answers move. It cannot say whether the program's pedestrian walks like a real one.

The thesis, here
The consequence of an action is the difference between two worlds that differ only in the action. A log of human driving holds one of the two worlds for each situation, chosen by the driver's reading of that situation, so it cannot give the difference; a program holds the situation and can build both, and when the two branches share their random numbers the difference is far quieter than the difference of two separate campaigns.
Linear position
Forced by: A real labelled frame is worth 26 frames of the calibrated program, and 400 of them are best spent calibrating and fine-tuning on 100 and grading with the other 300: the detector now misses 49% of the street's pedestrians where it first missed 80%, and the grade certifies at most 57%. That is a rate per frame, and the shuttle meets a pedestrian once per approach: first seen at 16 m, a pedestrian leaves 8 frames before the shuttle can no longer stop, and missing all 8 has probability 49% to the eighth power, 0.3%, if misses are independent and 49% if they repeat, which no set of single frames can say. Nor can any log of human driving say what braking would have changed, because the drivers who braked early are the ones who saw the pedestrian. Can the program say what would have happened had the shuttle braked?
New idea: a program can run both worlds of a decision from one set of random numbers, and the difference between them is the decision's consequence. The natural detector's policy cuts the collisions of the step-outs it meets from 55.4% to 14.6%, and sharing the random numbers lets the difference between two detectors be read from 7.2 times fewer episodes.
Forces next: Because a program can branch, the same step-out can be replayed with and without braking, and with the random numbers shared between the branches the difference between two detectors needs about 7 times fewer runs than two independent campaigns; the stopping margin then separates two detectors that 1,000 episodes of collisions cannot (2.9 standard errors against 1.6), while the exam, which grades frames, cannot see the decision rule that moves one detector's collisions from 9.3% to 20.0%. Every number so far is still a statement about the program, and the program is the thing under test: change only the pedestrian's walking speed, keeping its mean, and the closest band's collision rate moves by 19 points. What would count as proof that the system is safe on the real street, and how much real driving does that proof cost?
The plan
Six moves. (1) Build the loop from the exam's own pieces: an episode, a policy, an outcome. (2) Ask a log of human driving what braking does, and watch it fail. (3) Run both branches of every situation. (4) Share the random numbers. (5) Compare detectors by the stopping margin, and find what the exam cannot see. (6) Run the same detector in two worlds.

1 · From a frame to an episode

The exam asks one question of one frame: is the pedestrian found, at a threshold that lets a tenth of the pedestrian-free frames alarm? For the stored detector of Lesson 6's exact-stage program (the street's scene, light and camera; seed 1) the miss rate is 46.7% over the exam's pedestrians and 80.3% over those who step out from behind the van. The shuttle asks a different question: it drives, the camera records ten frames a second, a rule turns scores into a decision to brake, the brake takes time to act, and the outcome is a collision with a speed, or a stop with a margin in metres. One number per frame cannot say how that ends.

An episode is a step-out followed until the shuttle's front reaches the pedestrian or the shuttle stops. The pedestrian starts hidden at the van's shadow edge, as in Lesson 6's forced step-out, at depth z, and walks toward the lane at speed u. The camera sits at the shuttle's front and moves at speed v; the van and the clutter stand still, so frame k is the street seen from v·tk metres further on, with the pedestrian u·tk metres further left, the light unchanged and a fresh draw of sensor noise. The detector scores each frame, an alarm is a score above the exam's threshold, and the policy brakes once two of the last three frames have alarmed. The rest is mechanics, and these are the lab's assumptions, not facts about any vehicle: v = 7 m/s (25 km/h); a deceleration a = 3 m/s² once the brake acts (gentle: the UK Highway Code's stopping-distance table implies about 6.5 m/s² for a car); a latency τ = 0.4 s from decision to brake; a shuttle 2 m wide; a pedestrian who walks at u = 1.4 m/s. A collision is the front reaching the pedestrian's near surface while the pedestrian is within a metre, plus their radius, of the centre line.

The stopping distance has two terms, the distance covered before the brake acts and the braking distance (from v² = 2·a·s):

d = v·τ + v² / (2a) = 2.80 + 8.17 = 10.97 m

A pedestrian at depth z (near surface z − r, r about 0.25 m) can be stopped short of only if the brake is commanded by the deadline t* = (z − r − d) / v. At 10 m it is −0.17 s: no decision, by any detector, stops the shuttle short of the pedestrian. At 14, 22 and 30 m the deadlines are 0.40, 1.54 and 2.68 s after the step-out. The speed was chosen so that this arithmetic bites where Lesson 6's rare case lives: over the first 400 natural step-outs a shuttle that never brakes collides in 20.3%, 54.5% and 69.8% at 5, 7 and 8 m/s (stopping distances 6.2, 11.0 and 13.9 m), and the natural detector's policy leaves 0.3%, 15.0% and 39.0%. At 7 m/s the stopping distance sits at the edge of Lesson 6's closest step-outs, those nearer than 12 m. Stratify the step-outs by distance in the bands of Lessons 5 and 6 (12, 16 and 22 m), and take the natural detector through 1,000 of them, 250 per band, each band weighted by how often the street draws it (Lesson 6's p/q again):

Step-out distanceShare of the street's step-outsMean deadlineThe detector decides atCollisions with the policy
closer than 12 m (case R)5.9%−0.06 s0.31 s97.6%
12 to 16 m20.6%0.42 s0.41 s38.8%
16 to 22 m45.6%1.13 s0.69 s2.0%
22 m and beyond27.9%1.94 s1.04 s0.0%

The exam said the detector misses 46.7% of pedestrians, frame by frame. The loop says 14.6% of the step-outs the street produces end in a collision, and shows where. In the closest band the deadline is below zero and nothing can be done; in the next the detector decides at about the deadline (a slack, deadline minus decision, of 0.01 s) and 38.8% still collide; from 16 m on the slack is 0.44 s or more and the collisions fall away. What decides is when the detector decides, which no frame carries.

One choice remains, the threshold. The exam's lets one pedestrian-free frame in ten alarm, an alarm every second at ten frames a second; two alarms among three frames turn that into a decision on 29% of 3.5-second pedestrian-free approaches (300, drawn the same way), 5.9 false brakes a minute, and 15.1% of 1,000 natural step-outs still collide (against 14.6% above: two samples of one rate, whose standard error at 1,000 episodes is 1.1 points). A stricter threshold brakes falsely less and collides more: at 3% false alarms per frame, 2.2 false brakes a minute and 23% collisions; at 1%, 0.3 a minute, once in 208 s, and 38% collisions, against 55% if the shuttle never braked. No setting is a shuttle anyone would ride; the detector of this lab is a teaching detector, and the loop is how that shows. We keep the exam's threshold and the two-of-three rule so that every number of the series stays comparable, and say what they cost. Every detector below gets its own exam threshold, the same false-alarm rate, so none of them buys safety by braking more.

2 · What a log of human driving contains

The tempting route to a consequence is the data a fleet already holds: camera frames of human-driven shuttles, what the driver did, how the encounter ended. Compare the collisions of the logged step-outs in which the driver braked with those in which the driver did not. The lab can build such a log, and so can say what is wrong with it. Its driver has a personal threshold on how big a step-out must look to register, between 6 and 27 px² of silhouette (27 px² is the smallest silhouette of any step-out nearer than 12 m, so everyone sees those). A driver who registers the pedestrian brakes 0.7 s after the first frame in which 6 px² of the pedestrian show, the exam's own rule for a pedestrian who counts (the UK Highway Code's stopping-distance table implies a thinking time of about 0.67 s); one who does not, coasts. Who brakes depends on how big the step-out looks, that on how far away it is, and how far away it is also decides how it ends. That is confounding.

A pencil version first. Near step-outs are 30% of the street, far ones 70%; drivers brake in 90% of the near and 20% of the far. A collision follows with probability 0.5 (braked, near), 0.9 (unbraked, near), 0.02 (braked, far) and 0.10 (unbraked, far). Pooled, braked episodes collide in 33.6% and unbraked ones in 14.1%: braking looks harmful, by +19.5 points. Inside each kind it changes collisions by −40 points near and −8 far, and by −17.6 over the street: the pooled contrast compared near with far.

Now the lab's log: 4,000 natural step-outs driven by this driver, each with its true outcome and, because the program can also run the road not taken, the outcome of the other action.

Step-out distanceBrakedCollisions: brakednot brakedThe log's contrast, pointsThe program's, points
closer than 12 m100%100%none to comparenone+2
12 to 16 m84%35%91%−56−65
16 to 22 m39%0%59%−59−61
22 m and beyond13%0%6%−6−7
all, pooled44%26%37%−10.5−42.3

The program's column is the effect of everyone braking this way against nobody braking. Three things go wrong with the log. The pooled contrast is 4 times too small (−10.5 points against −42.3): the unbraked episodes are the far ones, where fewer collisions happen whatever anyone does. Stratifying by distance repairs it only where both kinds of episode exist (−56 and −59 points against −65 and −61), and only because the drivers' choices here depend on nothing but how big the step-out looks; a project cannot check that, since a log never says what else the driver was looking at. For the closest step-outs there is nothing to repair. Every driver brakes there, so the log holds no episode without braking, and no adjustment can invent one: the probability of the other action is zero (Rosenbaum and Rubin, 1983, require it strictly between 0 and 1 at every value of the covariates; Lesson 6's point again, no weight repairs a case with probability zero). The program shows what the log cannot: braking changed collisions there by +2 points, since the slower shuttle arrives as some pedestrians reach the lane, and cut the impact speed from 6.8 to 4.3 m/s.

Road not taken · reweight the log
Weight each logged episode by one over the probability that this driver braked in that kind of situation (Lesson 6's density ratio, applied to an action). It removes the confounding in expectation wherever both actions had a positive chance, which is the stratified answer again; it is undefined where the chance was zero, here the closest band, and where it was small, here the far band with 13% braked, a few episodes carry the estimate. Off-policy evaluation (Precup, Sutton and Singh, 2000; Jiang and Li, 2016) refines the weights, not the missing episode.

3 · Running both branches

The program holds what the log lacks: the situation. A seed fixes the scene, the light, the pedestrian's speed and the noise of every frame (the determinism of Lesson 9's replay), so the program can run one situation twice, differing in one thing, the action at the decision: the shuttle brakes as the policy decides, or never brakes. The decision is final, so no frame after it is rendered and each branch's outcome is the kinematics of §1 in closed form. Subtract the outcomes and the result is a number for this situation, d = (collision with braking) − (collision without): −1 if braking saved the pedestrian, 0 if it changed nothing, +1 if it caused the collision. Averaged band by band over the street's step-outs it is the effect of braking. For the natural detector:

Step-out distanceCollisions: never brakingwith the policySavedCaused
closer than 12 m97.6%97.6%1.6%1.6%
12 to 16 m95.6%38.8%58.0%1.2%
16 to 22 m63.2%2.0%61.2%0.0%
22 m and beyond4.0%0.0%4.0%0.0%
the street55.4%14.6%41.1%0.3%

Braking prevents 41.1 collisions in a hundred step-outs and causes 0.3. In the closest band it prevents and causes as many (1.6% each): the shuttle that brakes arrives later, so a few pedestrians who would have been struck are clear and a few who would have been clear are in the way. What braking buys there is speed, 6.8 to 3.9 m/s at impact, 65% less kinetic energy.

Lesson 6 left a number open: how much more a miss in the closest band should count than a miss elsewhere (its c, which the exam sets to 1). The loop prices it. In collisions avoided per step-out the policy buys 0.00 in the closest band and a mean of 0.43 in the others, so c = 0.0; in kinetic energy at impact (v²/2 per kilogram, a plain proxy for severity) it buys 16 J/kg against 12, so c = 1.3. The geometry sets a floor: the ideal detector, which decides the moment 6 px² of the pedestrian show, still collides in 89.6% of the closest band and in none beyond (5.3% over the street), and of the 9.4 points between it and the natural detector 85% are in the band from 12 to 16 m, where the mean deadline is 0.42 s and the detector decides at its edge.

4 · Sharing the random numbers

The detectors of Lesson 6 differ in their training sets and nothing else. Does the oversampled one collide less than the natural one? Estimate each rate from n episodes, Xi and Yi being 1 for a collision. The variance of the difference of the two means follows from the definition of variance (the identity behind common random numbers, analysed by Glasserman and Yao, 1992):

Var(X̄ − Ȳ) = [ Var X + Var Y − 2·Cov(X, Y) ] / n

Two separate campaigns, on different situations, have Cov = 0. If both detectors are run on the same situations, a hard step-out is hard for both, the outcomes move together, Cov > 0, and the variance falls by the factor k = (Var X + Var Y) / (Var X + Var Y − 2 Cov), which for equal variances is 1/(1 − ρ), with ρ the correlation of the outcomes. What is shared is the situation (scene, light, walking speed) and the sensor noise of frame k, drawn from a stream keyed by the seed and the frame index, not consumed in order: with one stream per episode, a branch that stops early would shift every later draw and the two branches would be strangers after the first difference. Measured on the 1,000 natural situations:

Pair, same situations (difference to resolve)Correlation of outcomesVariance ratio kEpisodes needed (95%): separate campaignsshared numbers
braking against never braking, natural detector (10 points)0.381.614593
natural against oversampled (25%) detector (2 points)0.867.22,417336
natural against naive detector (2 points)0.471.83,4611,905

The episodes needed solve 1.96·√(Var of one episode's difference / n) = the stated difference. Sharing pays in proportion to how alike the arms are. The natural and the oversampled detector end differently in only 3.5% of the situations (ρ = 0.86), so the shared comparison needs 7.2 times fewer episodes (a bootstrap over the situations gives 5.7: k has a sampling error of its own); a detector against a very different one, or an action against none, share little (k = 1.8 and 1.6). The identity also says where sharing fails: it hurts when the arms respond to the shared randomness in opposite directions, Cov < 0, which none of these pairs does.

One step-out, two branches
Slide the step-out distance; the situation is computed live. Left: distance against time. The amber bar is when the pedestrian is in the lane, at their depth (the dashed line); the grey dashed curve is the shuttle that never brakes, the teal curve the policy's; a red dot is a collision, a green dot a clear pass, the violet line the deadline, the red line the decision. Right: six frames (red border: an alarm) and every frame's score against the threshold. Below: the same 20-episode estimate of the chosen pair, redrawn 400 times from the 1,000 stored natural situations (decisions from the builder, outcomes recomputed here): tick the box for shared random numbers, untick it for two separate campaigns.
the detector decides at
—
deadline for a decision
—
stopping margin
—
with the policy
—
never braking
—
spread of the 20-episode estimate
—
variance ratio, shared against separate
—
the pair over all 1,000
—
Show the core JS
L.dStop = function (sh) { return sh.v0 * sh.tau + sh.v0 * sh.v0 / (2 * sh.a); };
L.timeAt = function (s, v0, a, tb) { if (s <= v0 * tb) return s / v0; var disc = v0 * v0 - 2 * a * (s - v0 * tb); return disc < 0 ? Infinity : tb + (v0 - Math.sqrt(disc)) / a; };
// brake from the onset tb (Infinity: never): a stop short of the pedestrian, a collision, or a pass
L.outcome = function (sit, tb, sh) {
  var zc = sit.z0 - sit.r, tf = L.timeAt(zc, sh.v0, sh.a, tb), margin = tb === Infinity ? null : zc - L.travel(Infinity, sh.v0, sh.a, tb);
  if (tf === Infinity) return { outcome: 'stop', collision: false, vImpact: 0, margin: margin, tf: Infinity };
  var xp = sit.x0 - sit.u * tf, hit = Math.abs(xp) < sh.half + sit.r;
  return { outcome: hit ? 'collision' : 'pass', collision: hit, vImpact: hit ? L.speedAt(tf, sh.v0, sh.a, tb) : 0, margin: margin, tf: tf };
};
L.windowCount = function (alarms, k, n) { var c = 0, j; for (j = Math.max(0, k - n + 1); j <= k; j++) if (alarms[j]) c++; return c; };
L.decide = function (alarms, m, n) { for (var k = 0; k < alarms.length; k++) if (L.windowCount(alarms, k, n) >= m) return k; return -1; };
L.branch = function (ep, action) {
  var tb = action === 'coast' ? Infinity : action === 'brake' ? ep.onset : action;
  var o = L.outcome(ep.sit, tb, ep.sh); o.onset = tb; return o;
};
L.noise = function (seed, k) { return SV.stream(seed, 'sensor' + k); };

What to try. (1) Set the distance to 10 m: the detector decides at 0.5 s against a deadline of −0.16 s, the margin is −4.6 m, and braking only cuts the impact from 7.0 to 5.3 m/s (this pedestrian's radius is under §1's 0.25 m). (2) At 14 m, the opening view, it decides at 0.4 s against a deadline of 0.39 s: a margin of −0.1 m and a collision at 0.6 m/s, where never braking hits at 7.0. (3) At 20 m the natural detector decides at 0.1 s and stops 8.1 m short, the naive one at 0.6 s and 4.6 m short; never braking passes clear, since the pedestrian has crossed. (4) Natural against oversampled, box ticked then unticked: the spread of the 20-episode estimate goes from 4.3 to 11.3 points, a variance ratio of 7.0; braking against never braking gives 11.0 and 13.3 (1.5), natural against naive 9.8 and 13.2 (1.8), near §4's k. The dashed line is the value over all 1,000: +0.7, −39.8 and −21.4 points.

5 · The stopping margin, and what the exam cannot see

The stopping margin of a decision at time td is M = z − r − v·td − d = v·(t* − td): the room left, in metres, when the shuttle has stopped, negative if it could not. It does not depend on how the pedestrian walks, so it ranks detectors in a unit the stopping distance understands. Five detectors, each at its own exam threshold, on the same 1,000 situations:

Detector (trained on)Exam: frames missedDecides before the arrivalMean stopping marginStep-outs that collide
natural (street scene, light, camera)46.7%98%2.8 m14.6%
oversampled 25% on case R (Lesson 6)53.1%96%2.6 m15.7%
oversampled 50%59.6%94%1.9 m17.5%
program scene, street light and camera62.1%95%1.2 m20.9%
the naive program80.5%73%−2.8 m35.5%

Over all eleven detectors (the eight programs of Lesson 2 and the three oversampled ones of Lesson 6) the rank correlation (Spearman, 1904) with the collision rate is 0.97 for the exam, −0.92 for the share of step-outs decided before the arrival and −0.98 for the mean margin. The measures part where detectors are close, and where no frame measure can see.

Close detectors. The oversampled detectors differ from the natural one only in how much of their training went to case R; the exam charges the 25% and 50% ones 6.4 and 12.8 points of frame misses. Paired on the same situations, the 25% detector's collisions differ from the natural detector's by +1.0 ± 0.7 points (1.6 standard errors) and its mean margin by −0.27 ± 0.09 m (−2.9); the 50% one's by +2.9 ± 0.9 points (3.3) and −0.92 ± 0.12 m (−7.5). A thousand episodes of collisions cannot separate the 25% detector from the natural one and the margin does: a collision is a coin, the margin a length, and the same episodes carry more in a length.

The decision rule. The exam is a number per frame, so it cannot depend on the rule that turns frames into a decision. One detector, one threshold, one exam number (46.7%), three rules: brake on one alarm, collisions 9.3% with 10.3 false brakes a minute; two of three, 15.1% and 5.9; three of five, 20.0% and 4.3. The same detector spans 10.7 points of collisions on a choice the exam does not see.

Runs of misses. Over the 11,950 frames of the 1,000 step-outs in which the pedestrian counts (6 px² showing) the detector misses 39.7%. A frame it misses is followed by another miss 72.3% of the time, against 10.0% after a frame it finds; the first three such frames are all misses in 48.3% of the step-outs, where independent misses at the same rate would give 6.3%. What makes a frame hard, how much of the pedestrian is hidden and how far away they are, persists to the next frame, so two alarms among three arrive later than independence promises, and no test set of single frames holds the neighbours that would show it. Codevilla and colleagues (2018) report that offline error is not necessarily correlated with driving quality.

Road not taken · count the step-outs it decides in time
The share of step-outs in which the detector decides before the arrival is a per-episode measure in the shuttle's own unit, and needs no kinematics. It saturates: the program-scene detector decides in 95% of the step-outs and the 50% oversampled one in 94%, which puts the first ahead by +1.0 ± 0.7 points; yet their collisions differ by +3.4 ± 0.9 points against the first (3.9 standard errors) and their margins by −0.74 ± 0.13 m (−5.6). The margin measures how early, not whether.

6 · The same detector in two worlds

Every number so far belongs to a world the lab wrote: a pedestrian who walks at exactly 1.4 m/s, a shuttle with the numbers of §1. The street a project will drive is not that world, and only the lab can run both (a lab privilege: a project can observe the street only by driving it). Keep the detector, the shuttle, the scenes, the light, the camera and the seeds, and change one assumption: let the pedestrian's speed vary from step-out to step-out, uniform between 0.6 and 2.2 m/s, which keeps the 1.4 m/s mean. The same 3,000 situations (750 per band), the natural detector, shared numbers:

Step-out distanceCollisions in the program's worldin the street'sDifference ± standard error, points
closer than 12 m98.0%78.8%−19.2 ± 1.5
12 to 16 m38.4%29.3%−9.1 ± 1.7
16 to 22 m4.1%4.8%+0.7 ± 0.9
22 m and beyond0.0%2.8%+2.8 ± 0.6

Weighted over the street the two worlds differ by −1.9 ± 0.6 points: small, and made of large parts that cancel. With one walking speed the pedestrian is in the lane in a sharp window; with a spread, the slow ones are still in the lane when a shuttle arrives from far (the farthest band's collisions without braking rise from 4.4% to 33.7%) and the fast ones have left it before one arrives from near (the closest band's fall from 97.5% to 81.3%). The loop was exact, and its answer for the closest band moved by 19.2 points when one assumption changed and the mean did not.

What would it take to see that on the street? Telling 13.6% from 15.6% at 95% confidence takes about 1,238 step-outs; telling the closest band's 78.8% from 98.0% takes 17 step-outs in that band, which at 5.9% of step-outs is about 296 in all. These are counts of step-outs; how much driving that is, no lab can say.

What this lesson did not do
It did not make the pedestrian real: the figure is a cylinder that walks at one speed, or in §6 at a speed from a range the lab chose, and whether real pedestrians step out like that is no question a lab can answer (Lesson 12). The shuttle's speed, deceleration and latency are assumptions; change them and every collision rate changes. No policy is learned and nothing is planned: the rule is fixed, because the lesson measures a decision's consequence rather than choosing one.

Common mistakes / failure modes

"47% of frames missed, so 47% of step-outs end badly."
The exam counts frames; the outcome depends on when two alarms land against a deadline: 14.6% collide, and the closest step-outs end badly whatever the detector does (§1).
"The braked episodes collided more, so braking hurts."
Drivers brake where the step-out looks big, which is near, where collisions are likely anyway: the toy's pooled +19.5 points hide −40 and −8 (§2).
"Common random numbers always cut the variance."
By 1/(1 − ρ) when the arms move together (7.2 here, 1.6 for braking against not braking), and not at all if the response flips sign; key the noise by (seed, frame) (§4).
"A better exam score is a safer shuttle."
The exam charges oversampling 6.4 and 12.8 points that the loop bills at 1.0 and 2.9, and cannot see the decision rule (§5).
"The simulator's collision rate is the street's."
One assumption moved, the mean kept: the closest band moved 19.2 points (§6).

Checkpoint exercise

Try it
A shuttle drives at 10 m/s, brakes at 4 m/s² and needs 0.5 s from the decision to the brake acting. A pedestrian stands in its lane 25 m ahead (near surface 24.75 m). (a) What are its stopping distance and its deadline? (b) The detector decides 0.9 s after the pedestrian appears: is there a collision, and at what speed? (c) The outcomes of two detectors across situations correlate at 0.9 and have equal variances: how many times fewer episodes does a shared comparison need? Answer: (a) d = 10·0.5 + 10²/8 = 17.5 m, and the deadline is (24.75 − 17.5)/10 = 0.725 s. (b) The decision is 0.175 s late: the margin is 24.75 − 10·0.9 − 17.5 = −1.75 m, so there is a collision, at the speed left after braking over the last 10.75 m: √(100 − 2·4·10.75) = 3.7 m/s. (c) 1/(1 − 0.9) = 10 times.

Where this points next

This lesson gave the program a closed loop whose answers are exact for each situation and quiet across a campaign: 14.6% of the natural step-outs end in a collision with the policy against 55.4% without braking, and a comparison of two detectors that needs 2,417 independent episodes needs 336 with shared numbers. They are answers about the program. The same detector collides in 98.0% of the closest step-outs in the program's world and 78.8% in a street whose pedestrians walk at speeds between 0.6 and 2.2 m/s with the same mean; the lab can run both, a project can run only the street, and a difference of 1.9 points overall would take about 1,238 step-outs to see. What would count as proof that the system is safe on the real street, and how much real driving does that proof cost?

Takeaway
An episode turns the exam's frame scores into the shuttle's outcome, a collision with a speed or a stop with a margin in metres, decided by when the detector decides against a deadline set by speed, deceleration and latency. A log of human driving cannot say what braking changes: drivers brake where collisions are likely anyway, and for the closest step-outs nobody did not brake. A program can run both branches of one situation exactly, and sharing the random numbers between branches or detectors turns a difference of noisy rates into a quiet one, by 1/(1 − ρ) for equal variances (7.2 here). The frame exam does not see the decision rule or the order of the frames; the stopping margin does. All of it holds inside the program's world, and moving one assumption of that world moved the closest band by 19.2 points.

Interview prompts

Companion reads: Reinforcement Learning · 19 Offline RL (data that lacks the action you need) and Reinforcement Learning · 12 Importance sampling (the density ratio of Lesson 6, applied to policies).