Where the gap lives
A gap of 33 points between a detector trained on the program and one trained on the street says that the program is wrong, not where. A program is a pipeline of four stages: the scene, which decides what exists; the light that falls on it; the camera that records it; and the rule that names what is in it. This lesson swaps the stages one at a time between the program and the street and trains sixteen detectors. It prices each stage in points of miss rate, finds that the repairs compound, and asks what evidence a project could collect about each stage without labelling the street. The camera turns out to be the largest gap and the cheapest to measure, which fixes where the track goes next.
New idea: the program is four stages, and swapping them one at a time prices each. The camera is worth 16 points of miss rate, the light and the scene about 8 each, the label nothing; the repairs compound, so a stage is worth more once the others are right; and the alarm rate on unlabelled frames measures the gap without a label. The price and the cost of the evidence together say where to start.
Forces next: Swapping the stages one at a time prices them: of the 33 points of miss rate between a detector trained on the program (80%) and one trained on the street itself (47%), the camera accounts for 16, the light and the scene for about 8 each, and the label for nothing, and the repairs compound, so the camera is worth 14.5 points while the other stages are wrong and 21 once they are right. The camera is therefore the place to start, because it is both the most expensive gap and the cheapest to measure: a few frames of a grey card pin down its noise without a single label. How does a camera turn light into numbers, and how do we measure it well enough to simulate it?
1 · The program is four stages
To make a frame together with its answer, a program decides four things, and it decides them in a fixed order, each stage reading the output of the one before.
| Stage | The question it answers | The program SIM0 | The street |
|---|---|---|---|
| Scene | What exists, and where? | one white van 8–12 m away; one pedestrian in half the frames, always in the open, 8–16 m out; a few poles and trees | vans of any size and place, sometimes none; pedestrians out to 24 m or more, about a quarter of them stepping out from behind the van; more clutter |
| Light | How is it lit and coloured? | noon sun behind the camera, one brightness, white van, bright shirts and dark trousers, blue sky | sun between 12° and 65° from any side, brightness varying eightfold, vans and shirts of any colour, blue, grey or warm sky, windows on the façade |
| Camera | How does light become numbers? | ideal: a fixed exposure, no noise, no blur | photon and read noise, a lens that blurs by 0.7 px, auto-exposure whose gain rises in the dark, gamma, 8 bits |
| Label | What counts as a pedestrian? | any pixel showing | at least 15% of the silhouette showing |
Nothing here is specific to a street: any data-generating program, in any renderer, has the same four joints, and in a production pipeline each is a different team's code. That is why a gap between a program and the world should be divided up before it is repaired.
2 · Sixteen programs
Name a program by where each stage comes from, four letters in the order scene, light, camera, label: a for the setting of SIM0, b for the street's. Then aaaa is SIM0 and bbbb is the street itself, and aabb is a program with the naive scene and light and the street's camera and label. There are 2⁴ = 16 of them. Train the fixed detector on 1,600 frames of each, three seeds, and grade it on the street with the exam of Lesson 1.
The lab can do this because the street is a program. A project cannot: it has no way to build a hybrid of its program and the real street, and §5 says what it can do instead. The table shows the eight programs that differ in scene, light and camera, with the label left as the program's rule (making it the street's changes the numbers by less than a point, as explained below the table).
| Scene | Light | Camera | Miss rate on the street | AUC | Alarm rate on unlabelled frames |
|---|---|---|---|---|---|
| program | program | program | 80.3% | 0.596 | 100% |
| street | program | program | 74.8% | 0.645 | 100% |
| program | street | program | 73.3% | 0.666 | 100% |
| program | program | street | 65.8% | 0.690 | 93.7% |
| street | street | program | 68.3% | 0.694 | 100% |
| street | program | street | 60.4% | 0.730 | 81.8% |
| program | street | street | 61.6% | 0.749 | 25.5% |
| street | street | street | 46.8% | 0.792 | 10.7% |
Read the first column of numbers down. Making one stage the street's moves the miss rate from 80.3% to 74.8% (scene), 73.3% (light) or 65.8% (camera); making all three the street's reaches 46.8%. No stage is the whole story, and every stage is some of it. The label does not appear: aaab reproduces aaaa to the last digit and bbbb differs from bbba by 0.3 points, because the two label rules disagree only about a pedestrian who is partly hidden, and a program whose scene never hides one never meets the disagreement. The label is not irrelevant, but its cost is not paid in this exam; Lesson 7 shows where it is.
3 · The price list, and why repairs compound
How much is a stage worth? The tempting answer, the gain from repairing it alone, depends on what else is already repaired, so there is no single answer; there is one for each order of repair. Take the average over all of them. For stage i and a set S of stages already repaired, let v(S) be the hit rate (one minus the miss rate) of the program in which exactly the stages of S are the street's. The price of stage i is its gain averaged over every order in which the four stages could be repaired:
φi = ΣS ⊆ stages∖{i} [ |S|! (3 − |S|)! / 4! ] · ( v(S ∪ {i}) − v(S) )
This is the Shapley value of cooperative game theory, and it has the property that matters here: the four prices add up exactly to the whole gain, v(all) − v(none), however the stages interact. From the sixteen programs the prices are:
| Stage | Price in points of miss rate | Share of the 33.2 points |
|---|---|---|
| Camera | 16.3 | 49% |
| Light | 8.5 | 26% |
| Scene | 8.4 | 25% |
| Label | −0.1 | 0% |
Each cell is a mean of three seeds and wobbles by about 1 point between them, so the camera is clearly the largest price and the light and the scene are a tie. Behind each price is a table of what the stage buys given the state of the others, and it shows the second fact: repairs compound. The camera is worth 14.5 points when light and scene are both wrong and 21.5 when both are right. The scene is worth 5.5 points while the light and the camera are wrong, 5.4 and 5.0 if only one of them is repaired, and 14.8 once both are. The light behaves the same way (7.0 points at worst, 13.5 at best). Taken together, light and scene are worth 12.0 points with the camera wrong and 18.9 with it right.
One way to see why. A detector learns in the program's pixels, and what it learns reaches the street only through whatever the two sets of pixels have in common. If the camera model is wrong, a new case or a better range of lighting is learned in pixels that no street camera produces, and part of it does not transfer. A pedestrian is found when the case was seen and the light was familiar and the noise was tolerable, a conjunction, and a repair of one conjunct pays little until the others hold; the pixel stages, light and camera, gate everything else the program can teach. The picture is only a picture. If the stages cost independent fractions of the pedestrians, the losses would multiply, and the naive program would find 17.1% of them; it finds 19.7%, because the same pedestrians are lost for several reasons at once. The table is the measurement and the picture only a way to remember its shape.
What to try. Start as the page opens, nothing repaired: the program is SIM0, the miss rate is 80.3%, and all 24 unlabelled street frames are red. Slide to camera: 65.8%, the first repair buys 14.5 points, and the meter hardly stirs, 93.7% of unlabelled frames still alarm. Slide to light: 61.6%, only 4.2 points more on the exam, yet the alarm rate drops from 93.7% to 25.5%: most thumbnails go quiet. Slide to scene: 46.8%, 14.8 points, the repair that was worth 5.5 when the camera and the light were wrong. The last notch, label, buys nothing. Now take the other order with the boxes: tick scene alone (74.8%), then light (68.3%), then camera: 46.8%, 21.5 points from the last repair, and the first two bought 5.5 and 6.4. The end point is the same; the road to it is not. In the bars on the left every one of the sixteen programs sits somewhere between the first bar and the last, and the label dots move a bar by less than a point.
4 · A meter that needs no label
The exam needs labelled street frames, and labelled street frames are what a project is trying to avoid buying. But there is a measurement that needs none. A pedestrian is rare in ordinary driving (about one frame in fifty here), so a stream of unlabelled street frames is, to a good approximation, a stream of pedestrian-free frames. Take the threshold the program chose for itself, the one that makes one in ten of its own pedestrian-free frames alarm, and apply it to those street frames. If the program matched the street the alarm rate would be about ten per cent, because the street's frames would score as the program's do. Any excess is the gap, seen without a label.
The column on the right of the table above is that alarm rate over 1,000 unlabelled frames. The naive program scores 100%: every street frame, nearly, scores above a threshold that a tenth of the program's own frames exceed, because the street's frames carry clutter, grain and light the detector has never seen and treats as evidence. The matched program scores 10.7%. A finer version of the same meter compares the two distributions of scores directly: the probability that a street frame outscores a pedestrian-free frame of the program, 1.00 for the naive program and 0.50 for the matched one, where 0.5 means the two cannot be told apart. Across all sixteen programs this meter and the real miss rate agree in rank to a Spearman correlation of 0.98: a program that looks wrong on unlabelled frames is wrong on the exam.
Two cautions keep it honest. The meter says that the gap exists and how large it is, not which stage it lives in: it was blind to the camera repair alone (100% to 93.7%) and woke only when the light was repaired as well. And it saturates: at 100% it cannot tell a bad program from a hopeless one. It is a cheap reading to take at every step of the work, not a replacement for the exam.
5 · What evidence answers to which stage, and what to repair first
Each stage leaves a different trace in the data a project can get, and the traces differ in cost.
| Stage (price) | Evidence that sees it | What it costs | What it cannot tell you |
|---|---|---|---|
| Camera (16.3) | flat frames of a grey card at several light levels; a sharp edge; the exposure and gain the camera logs | an afternoon with the camera; no labels, no street | anything about what stands in the street |
| Light (8.5) | unlabelled street frames: pixels whose content is known (sky, road) and ratios that exposure control cancels | free, from logs the fleet already collects | what the scene holds; how a rare pedestrian is dressed |
| Scene (8.4) | a few hundred labelled frames and the ground plane, which turns a mask into a distance; logs, for how often things happen | people labelling frames | how the pictures are lit or recorded |
| Label (−0.1) | the benchmark's written definitions | reading them | whether a detector will meet the disagreement |
Price says what a repair is worth; the table says what it costs to learn what to repair. The camera is first on the first count and first on the second. Its repair is also the one that most raises the value of the others: light and scene together are worth 12.0 points before it and 18.9 after. And it gives the clearest feedback, because it is the repair whose effect on the exam is visible at once. The scene stage, whose evidence is the most expensive, can wait until its payoff is large enough to see.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
The swap table prices the stages, and it is a laboratory privilege: a project holds the street's frames and cannot hold the street's program. What it can hold is the traces of §5, and they are not equally cheap. The camera is the most expensive gap, 16.3 of the 33.2 points, and the one whose evidence is cheapest, because a grey card in front of the lens answers a question about the camera and nothing else. Repairing it raises what the rest are worth, from 12.0 points to 18.9, and the repair shows on the exam at once. The camera is also the stage the program treats most casually: it writes the renderer's light straight into the image as if it were the number the sensor reports. How does a camera turn light into numbers, and how do we measure it well enough to simulate it?
Interview prompts
- How would you decide which part of a simulator to improve first? (§3, §5 — price each stage by swapping it between the program and the world, average the gain over orders of repair, and weigh the price against the cost of the evidence that measures the stage.)
- Why can a stage be worth five points in one setting and fourteen in another? (§3 — repairs compound: a detector learns in the program's pixels, so a repair of what exists or how it is lit pays only where the pixels already resemble the street's.)
- What is a Shapley value and why use it here? (§3 — the average gain of a stage over every order of repair; unlike fix-one or break-one it does not depend on the order, and the prices sum to the total gain.)
- How do you tell, without any real labels, that synthetic training data is off? (§4 — the alarm rate on unlabelled real frames at a threshold set on the synthetic data; pedestrians are rare, so almost every frame is pedestrian-free and the rate should be the chosen one.)
- What are the limits of that test? (§4 — it says that the gap exists, not which stage; it saturates; it was blind to the camera repair alone.)
- Why start with the camera, the stage that looks least interesting? (§5 — the largest price, the cheapest evidence, and the stage whose repair raises the value of the others.)
- A team repairs its simulator in pipeline order and sees no gain for two quarters. What happened? (§5, road not taken — the early stages cost the most to measure and pay the least while the pixels are wrong; the first repair that moves the exam is the camera's.)
Companion reads: Computer Vision · 04 Cameras and projection geometry (the camera as a function, which Lesson 3 turns into a measurement), and Computer Graphics, from first principles (the stages the program is made of, run forwards).