A little reality
Lesson 9 ended with a number we can trust to describe the program, and a program that is wrong in ways nothing inside it can see. Real labelled frames can see them, and they are scarce: the exam this series has used since Lesson 1 holds 4,500 of them. This lesson prices a real frame by what it is spent on, in one currency, the frames of real-only training that give the same exam miss rate. Calibrating the program on 50 frames takes the detector from 80% to 56%; fine-tuning the calibrated program on the same 50 reaches 50%, which real frames alone need 158 to match, and past a few hundred frames real frames alone do better. Grading takes the largest share: ±5 points of miss rate costs 869 frames. And a miss rate per frame is not yet a probability per approach.
New idea: a real frame is worth what it is turned into, measured in the frames of real-only training that give the same exam miss rate. A frame that fine-tunes the calibrated program is worth 3 of them at 50 frames, and the program's 1,600 frames are worth 62, so one real frame is worth 26 of the program's.
Forces next: A real labelled frame is worth 26 frames of the calibrated program, and 400 of them are best spent calibrating and fine-tuning on 100 and grading with the other 300: the detector now misses 49% of the street's pedestrians where it first missed 80%, and the grade certifies at most 57%. That is a rate per frame, and the shuttle meets a pedestrian once per approach: first seen at 16 m, a pedestrian leaves 8 frames before the shuttle can no longer stop, and missing all 8 has probability 49% to the eighth power, 0.3%, if misses are independent and 49% if they repeat, which no set of single frames can say. Nor can any log of human driving say what braking would have changed, because the drivers who braked early are the ones who saw the pedestrian. Can the program say what would have happened had the shuttle braked?
1 · What a real frame can buy
The exam holds 3,000 test frames and 1,500 validation frames, 4,500 real labelled frames in all. The repair of Lessons 3 to 5 used none for the camera (a bench kit and thirty logs), none for the light (a thousand unlabelled logs) and at most 200 outlined frames for the scene, so knowing that it worked cost at least 22.5 times what making it work did. A project has an M, and a few hundred is a good one. A labelled frame can be spent three ways:
| A real frame is spent on… | What comes out | What it costs |
|---|---|---|
| grading | a number and its interval: the detector's miss rate on the street, from frames it has never seen | the bill of §2 |
| calibrating | a program: its stages' settings, read from outlines (scene), unlabelled logs (light) and a bench kit (camera) | no labelled frame for camera and light, about 50 for the scene (Lesson 5) |
| fine-tuning | a detector: the real frames join the synthetic ones in training | each frame, once |
They are not rivals in the same way. Calibrating reads a few numbers off the frames and leaves them where they were, so a frame that calibrated can still train; a frame that trained is no longer a fair test. The real competition is between training and grading. To compare uses we need one currency, and the exam gives only miss rates, so the currency is what a frame is worth: the frames of real training that reach the same miss rate. The yardstick is the real-only learning curve R(M), the fixed detector trained on the first M frames of the street's labelled stream and graded by the exam. Every number below averages afternoons, each one a stream of labelled frames with its own logs and kit: three for every strategy, six for the yardstick, whose trainings are small. The free repairs come first. The camera read from a kit and the brightness read from 1,000 logs, with no labelled frame, take the naive program's detector from 80.3% to 62.5%, while a detector trained on 1,600 real frames misses 46.0%. The first 18 points cost nothing; the other 17 are what a labelled frame must buy, and the question is which use buys them cheapest.
2 · Grading: the use nothing replaces, and the largest bill
A miss rate is a count: of n pedestrians, k are missed, and p̂ = k/n has standard error √(p(1 − p)/n). What counts is pedestrians, not frames: only 44.1% of the street's frames hold one the exam counts (1,323 of 3,000), so a labelled frame brings 0.441 of them and 25 frames bring about 11. At 11 pedestrians and an 80% miss rate the interval p̂ ± 1.96 √(p̂(1 − p̂)/n) covers the true rate in 90% of samples, not 95%, and it collapses at a count of zero. Wilson's interval (1927) keeps every p within 1.96 standard errors of the data, a quadratic, covers 95% there, and survives zero misses (0 of 20 gives an upper end of 16%):
p ∈ [ ( p̂ + z²/2n ± z √( p̂(1 − p̂)/n + z²/4n² ) ) / ( 1 + z²/n ) ], z = 1.96
Inverting the normal interval gives the pedestrians a half-width h needs, n = 1.96² p(1 − p)/h², and dividing by 0.441 gives frames, for a detector that misses about half (the series' detectors) and for the naive one:
| Half-width of the 95% interval | Pedestrians (miss rate 47%) | Frames (47%) | Frames (80%) |
|---|---|---|---|
| ±5 points | 383 | 869 | 558 |
| ±3 points | 1,064 | 2,413 | 1,549 |
| ±2 points | 2,393 | 5,427 | 3,486 |
The exam's 1,323 pedestrians read a 47% miss rate to ±2.7 points, the best case a project will see. A difference of two detectors needs more care. Graded on the same pedestrians, two detectors are right and wrong about many of the same ones, and the variance of the difference of their miss rates is (b + c)/n² − (b − c)²/n³, where b pedestrians are found by the first and missed by the second and c the reverse; independent samples would add the two binomial variances. Lesson 4's detectors with the exact camera, and with the exact camera and light, differ by 4.7 points on the exam. They disagree on 19.5% of the pedestrians, so the paired standard deviation is 1.2 points and the unpaired one 1.9: the same difference is 3.9 standard deviations from zero paired and 2.5 unpaired. A 3-point improvement needs 829 pedestrians on the same ones, and 1,952 on different ones.
The threshold is a measurement too. The exam asks for the detector that lets 10% of pedestrian-free frames alarm, a quantile estimated from M₀ such frames, so the false-alarm rate it delivers has standard deviation √(0.09/M₀). At the exact-stage detector's operating point the miss rate is 61.8, 48.3 and 39.2% at 5, 10 and 15% false alarms, 2.3 points of miss rate per point of false alarms. Redrawing the threshold 500 times:
| Pedestrian-free frames M₀ | 50 | 150 | 500 | 1,500 |
|---|---|---|---|---|
| sd of the false-alarm rate, points | 4.4 | 2.5 | 1.3 | 0.8 |
| sd of the miss rate, points | 9.3 | 5.1 | 2.8 | 1.7 |
The exam's own threshold comes from 798 pedestrian-free validation frames and carries ±2.4 points of miss rate, more than the ±1.4 of its pedestrians; for this detector it let 10.6% of fresh pedestrian-free frames alarm, worth 1.6 points in its favour, inside that error. What the threshold needs are pedestrian-free frames, and 98% of the unlabelled logs are: the 90th percentile of the scores on 2,000 logs lets 9.1% of pedestrian-free frames alarm, 0.9 points short of 10 because the logs' own pedestrians use up part of the 10% ((0.1 − 0.02·TP)/0.98 gives 9.2%), with no labelled frame at all. The threshold costs frames that are free; the pedestrians cost frames that are not.
So grading is the largest bill and no other use can pay it: a calibrated program cannot certify itself, because its own frames are what is under test. The exam spent 4,500 frames to read ±2.7 points; a project with M = 400 reads ±7.4 at best. Everything below was graded on the lab's big sets, and §5 prices what a project would have to spend to know it.
3 · Four ways to spend M on a detector
| Strategy | The detector's training set | Evidence it reads |
|---|---|---|
| R, real only | the M real frames | M labelled frames |
| F, fine-tune the naive program | 1,600 frames of the naive program and the M real frames, each real frame weighing 1,600/M | M labelled frames |
| C, calibrate | 1,600 frames of the program rebuilt from evidence | a bench kit, 1,000 logs, and M labelled frames for the scene only |
| CF, calibrate then fine-tune | 1,600 frames of the calibrated program and the M real frames, weighted as in F | all of the above |
The weight is decided before the lab speaks: the real frames weigh as much in total as the synthetic ones, treating the two sources as equally informative per unit of weight; §4 measures what that costs. The calibration is Lessons 3 to 5 composed. A kit of size one (11 exposures of a grey card, one edge shot, 30 exposure-log frames) gives the camera and 1,000 logs give the brightness range; the first M labelled frames give the pedestrians' distance range and step-out probability through Lesson 5's ruler. The sun, palettes, clutter, van and label rule stay the naive program's: no estimator reaches them. What the evidence reads, three afternoons (± is the standard deviation over them), against the street's own settings, which only the lab knows:
| Read from the first… frames | 25 | 50 | 200 | 800 | The street |
|---|---|---|---|---|---|
| nearest pedestrian, m | 10.1 ± 1.4 | 8.8 ± 1.0 | 8.0 ± 0.0 | 8.0 ± 0.0 | 8 |
| farthest pedestrian, m | 23.2 ± 2.5 | 24.6 ± 0.0 | 24.6 ± 0.0 | 24.6 ± 0.0 | 24 |
| step-out probability | 0.30 ± 0.20 | 0.27 ± 0.10 | 0.25 ± 0.02 | 0.28 ± 0.02 | 0.28 |
The kit reads the full well as 4,014 ± 13 electrons (the street: 4,000) and the logs read the brightness range as 0.11 to 0.95 (the street: 0.12 to 1.00), neither with a labelled frame.
What to try. Slide to 25 frames: the same frames feed all four boxes, and R misses 68% against CF's 55%. Slide to 100 and compare none with the best split (50 · 50): the certificate box fills, at most 71% for a detector that misses 50%. At 400 the best split is 100 · 300 and certifies at most 57.3% (±8.5); half certifies 58.3% and a quarter 62.7%, because the interval widens faster than the miss rate falls. Tick the street's settings at 25 and then at 200 frames and watch the teal range settle on the red ticks. At 1,600, R (46%) has dropped below CF (49%).
4 · The exchange rate
Mean exam miss rate, %, over the afternoons, every entry a trained detector:
| Frames M | 0 | 25 | 50 | 100 | 200 | 400 | 800 | 1,600 |
|---|---|---|---|---|---|---|---|---|
| R, real only | — | 68.1 | 59.1 | 51.4 | 48.8 | 47.4 | 45.9 | 46.0 |
| F, naive + real | — | 60.6 | 60.4 | 55.6 | 54.3 | 54.1 | 54.0 | 54.1 |
| C, calibrated | 62.5 | 59.7 | 55.6 | 55.9 | 55.5 | 56.0 | 56.0 | 56.1 |
| CF, calibrated + real | — | 55.0 | 49.9 | 48.8 | 47.9 | 48.8 | 49.2 | 48.6 |
C is flat from 50 frames on: three numbers are all the labelled frames set, and 50 frames set them. It stops 3.7 points short of the program with the street's own camera, light and pedestrian settings (Lesson 5: 52.2%), because the light's sun and palettes and the clutter come from no estimator. R catches C at 68 frames and CF at 264. CF misses 14.8, 9.9 and 2.5 points less than R at 25, 50 and 100 frames (in all three afternoons), 1.0 less at 200, and 1.5, 3.0 and 3.0 points more at 400, 800 and 1,600: half of its loss is the program's frames, which hold it near 49%, while real frames alone keep improving. F beats R only at 25 frames and is 8 points behind at 1,600. Inside the mixture the calibration is worth 6, 10, 7 and 6 points at 25, 50, 100 and 200 frames. The afternoons differ (CF at 25 frames ranges over 15 points), so gaps under about 3 points are inside the spread.
For the conversion, fit the real-only curve R(M) = a + b M−c: a floor a = 44.8%, c = 0.83, residual 3.0 points, the afternoon spread. The real-only size that matches a miss rate m is E(m) = ((m − a)/b)−1/c; a synthetic frame is worth E/N real ones and a real frame spent on a use E/M:
| What the frames are | Real frames spent | Exam miss | Real-only frames that match it | Per frame spent |
|---|---|---|---|---|
| the calibrated program, scene from 50 frames (C) | 50 | 55.6% | 64 | 1.3 |
| calibrate, then fine-tune on the same 50 (CF) | 50 | 49.9% | 158 | 3.2 |
| … on 100 | 100 | 48.8% | 217 | 2.2 |
| … on 200 | 200 | 47.9% | 289 | 1.4 |
| … on 400 | 400 | 48.8% | 215 | 0.5 |
The calibrated program's 1,600 frames are worth 62 real frames, 0.039 each, so one real frame is worth 26 of them; the naive program's are worth 0.009 each (15 real frames in all) and the free evidence alone 35. The ceiling is the street's own scene, light and camera, which no project can build: in Lesson 2's table its 1,600 frames equal 1,600 real ones (0.3 points apart), worth 1 each. A frame that fine-tunes the calibrated program returns 3 real frames up to 100 frames, 1.4 at 200 and under one at 400, where CF falls behind R. The rate falls with the budget because the program's contribution is a head start, worth at most about 289 real frames, and a head start matters less the farther the real frames have run.
The weight is the exchange rate in disguise: a real frame of weight w counts as w synthetic ones in the loss. At 200 frames (one afternoon) F misses 62.8, 57.4, 54.9, 52.1 and 48.8% at w = 1, 4, 8 (the rule), 32 and 128, a spread of 14 points; CF misses 49.5, 50.2, 49.7, 49.8 and 48.7%, a spread of 1.5. A naive frame is worth 0.009 real ones, about a hundred to one, and F is best at the swept weight nearest that, 128. The calibrated program is right at any weight and the naive one only at the weight its worth implies: what calibration buys inside the mixture is a weight nobody had to tune.
5 · Spending a budget
A frame can calibrate and train at once and cannot train and grade, so a budget M splits into Mt frames that calibrate and fine-tune and G = M − Mt held out. The detector trained on Mt misses p(Mt), the best of the four curves, and the held-out frames hold 0.441G pedestrians, so the strictest claim they support, the upper end of the 95% interval (planned with the normal interval), is
bound(Mt) = p(Mt) + 1.96 √( p(1 − p) / (0.441 (M − Mt)) )
Training more lowers the first term and widens the second; the best split is where a frame moved from grading to training lowers the miss rate exactly as much as it widens the interval. Over the measured training budgets:
| Budget M | Train (calibrate and fine-tune) | Grade | Detector misses | Certified: at most | Price of certainty |
|---|---|---|---|---|---|
| 200 | 50 | 150 | 49.9% | 62.0% | 12.0 |
| 400 | 100 | 300 | 48.8% | 57.3% | 8.5 |
| 800 | 200 | 600 | 47.9% | 53.9% | 6.0 |
| 1,600 | 800 | 800 | 45.9% | 51.1% | 5.2 |
| 4,500 (the exam) | 800 | 3,700 | 45.9% | 48.3% | 2.4 |
The rule has three parts, each with its break-even. Calibrate first: it costs no frame it cannot give back, and below 68 frames it beats training outright. Fine-tune on what the curve still falls for: with the calibrated program the curve is flat from about 100 frames (48.8, 47.9, 48.8%), and past 264 frames real frames alone beat the mixture. Grade with the rest: from 400 frames on, more than half the budget goes to the grade, because a trained frame now buys under a point and a held-out one narrows the interval. The price of certainty, how far the certified bound sits above the miss rate, is 12 points at 200 frames and 2.4 at the exam's 4,500: a project cannot show what it has built. Turned around, certifying a miss rate of at most 55% takes 200 frames of training and 434 of grade, 634 in all, and at most 52% takes 1,378.
6 · A rate per frame is not a probability per approach
Every price above is in the exam's currency, a miss rate per frame, and the detector the best split trains at 400 frames misses 49% where the naive one missed 80%. The shuttle does not live per frame. Take the next lesson's shuttle, the lab's assumptions and not facts about any vehicle: it drives at 7 m/s (25 km/h), decelerates at 3 m/s² once it has decided, and takes 0.4 s from decision to deceleration, so it stops after dstop = vτ + v²/2a = 2.8 + 8.2 = 11.0 m. A pedestrian first seen at depth z₀ is safe only if the detector fires while the shuttle is still farther away than dstop. The camera records ten frames a second, the shuttle moves 0.7 m between frames, and the frames that can still help number ⌊(z₀ − dstop)/0.7⌋ + 1. If the detector misses a frame with probability p = 49%:
| Pedestrian first seen at | 12 m | 14 m | 16 m | 20 m |
|---|---|---|---|---|
| frames before the last useful moment, k | 2 | 5 | 8 | 13 |
| no detection, misses independent: pk | 23.8% | 2.8% | 0.3% | 0.01% |
| no detection, pedestrians differ (correlation 0.5 between a pedestrian's frames) | 36% | 24% | 18% | 15% |
| no detection, misses repeat in every frame: p | 49% | 49% | 49% | 49% |
The middle row gives every pedestrian their own miss probability (some are easy, some hard, the same in every frame of one approach) with mean p and correlation 0.5 between two frames of the same pedestrian: the chance that all k frames miss is E[qk] for q beta-distributed, a product of k terms. At 16 m the rows differ by a factor of 153, and each describes some detector with a 49% miss rate per frame. The exam cannot say which describes this one: its test set holds single frames, and the correlation is a property of pairs of frames of one approach. This detector misses partly pedestrians who are hard to see and stay hard as the shuttle nears (misses that repeat) and partly noise (misses that do not), and nothing in a set of single frames weighs the two. The number the shuttle needs exists only in the program's closed loop.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
A real frame now has a price in every currency the project uses: 26 frames of the calibrated program, about 3 frames of real training when it fine-tunes that program, and nothing else when it grades, which at 400 frames certifies 57% for a detector that misses 49%. But that is a rate per frame, and the shuttle meets a pedestrian once per approach: first seen at 16 m, a pedestrian leaves 8 frames, and missing all of them has probability 0.3% if misses are independent and 49% if they repeat, which a set of single frames cannot say. And braking is an action whose consequences no log of human driving contains, because the drivers who braked early are the ones who saw the pedestrian. Can the program say what would have happened had the shuttle braked?
Interview prompts
- You have 300 labelled images and a simulator. How do you spend them? (§5 — calibrate the simulator's few parameters, and the camera and light from free evidence, first; fine-tune on what the curve still falls for; grade with the rest, and expect the certificate to sit well above the miss rate.)
- What is a real image worth in synthetic ones? (§4 — it depends on the use: 26 frames of the calibrated program, about a hundred of the naive one, and nothing replaces it when it grades.)
- How many test pedestrians does it take to see a 3-point improvement? (§2 — 1,952 if the two systems are graded on different pedestrians, 829 on the same ones; pairing removes the variance the two share.)
- Your threshold was set on 50 frames. What does that do? (§2 — its false-alarm rate has standard deviation √(0.09/50) = 4.4 points, which moves the miss rate by ±9; unlabelled logs can set it instead, with a known bias.)
- When does adding a simulator to your real data stop helping? (§4 — when real frames alone have caught up with the program's head start, near 264 frames here, because the program's frames stay in the loss.)
- A detector misses 49% of pedestrians per frame. What is the chance it misses one approach? (§6 — from 0.3% to 49% at 8 frames to spare, depending on how misses repeat, which a set of single frames cannot show.)
Companion reads: Training a Robot Model · 16 Exchange rates (the same unit for hours of robot data), Training a Robot Model · 15 Evaluation (what grading costs there) and Computer Vision · 19 Evaluation, deployment, and CV system design.