all_lessons/Synthetic Vision Data/02 · Where the gap liveslesson 2 / 12

Where the gap lives

A gap of 33 points between a detector trained on the program and one trained on the street says that the program is wrong, not where. A program is a pipeline of four stages: the scene, which decides what exists; the light that falls on it; the camera that records it; and the rule that names what is in it. This lesson swaps the stages one at a time between the program and the street and trains sixteen detectors. It prices each stage in points of miss rate, finds that the repairs compound, and asks what evidence a project could collect about each stage without labelling the street. The camera turns out to be the largest gap and the cheapest to measure, which fixes where the track goes next.

The thesis, here
A gap between two programs is a list of gaps between their stages, and each stage answers to a different kind of evidence. Pricing the stages in one currency, points of the same exam, turns "make the simulator better" into a price list, and the price list together with what each kind of evidence costs chooses what to repair first.
Linear position
Forced by: Free labels are exact and unlimited, and a detector trained on them stops making mistakes on the program's own pictures after about forty frames; on pictures of the real street it still misses 80% of the pedestrians, whether it saw forty frames or twenty-five hundred. The gap is not noise that more data averages away: it is a difference between the program and the world, and the program has four places to be wrong: how the camera records the scene, how the scene is lit, what exists in it, and what the label means. Which of them costs the most, and how could we tell without a real label for every frame?
New idea: the program is four stages, and swapping them one at a time prices each. The camera is worth 16 points of miss rate, the light and the scene about 8 each, the label nothing; the repairs compound, so a stage is worth more once the others are right; and the alarm rate on unlabelled frames measures the gap without a label. The price and the cost of the evidence together say where to start.
Forces next: Swapping the stages one at a time prices them: of the 33 points of miss rate between a detector trained on the program (80%) and one trained on the street itself (47%), the camera accounts for 16, the light and the scene for about 8 each, and the label for nothing, and the repairs compound, so the camera is worth 14.5 points while the other stages are wrong and 21 once they are right. The camera is therefore the place to start, because it is both the most expensive gap and the cheapest to measure: a few frames of a grey card pin down its noise without a single label. How does a camera turn light into numbers, and how do we measure it well enough to simulate it?
The plan
Five moves. (1) Open the program: four stages, in the order they run. (2) Swap them: sixteen programs, each a mixture of the naive program and the street, each trained and graded the same way. (3) Price each stage by the average of what it buys over every order of repair, and see why the repairs compound. (4) Find a meter of the gap that needs no label. (5) List what evidence each stage answers to and what it costs, and choose what to repair first.

1 · The program is four stages

To make a frame together with its answer, a program decides four things, and it decides them in a fixed order, each stage reading the output of the one before.

StageThe question it answersThe program SIM0The street
SceneWhat exists, and where?one white van 8–12 m away; one pedestrian in half the frames, always in the open, 8–16 m out; a few poles and treesvans of any size and place, sometimes none; pedestrians out to 24 m or more, about a quarter of them stepping out from behind the van; more clutter
LightHow is it lit and coloured?noon sun behind the camera, one brightness, white van, bright shirts and dark trousers, blue skysun between 12° and 65° from any side, brightness varying eightfold, vans and shirts of any colour, blue, grey or warm sky, windows on the façade
CameraHow does light become numbers?ideal: a fixed exposure, no noise, no blurphoton and read noise, a lens that blurs by 0.7 px, auto-exposure whose gain rises in the dark, gamma, 8 bits
LabelWhat counts as a pedestrian?any pixel showingat least 15% of the silhouette showing

Nothing here is specific to a street: any data-generating program, in any renderer, has the same four joints, and in a production pipeline each is a different team's code. That is why a gap between a program and the world should be divided up before it is repaired.

2 · Sixteen programs

Name a program by where each stage comes from, four letters in the order scene, light, camera, label: a for the setting of SIM0, b for the street's. Then aaaa is SIM0 and bbbb is the street itself, and aabb is a program with the naive scene and light and the street's camera and label. There are 2⁴ = 16 of them. Train the fixed detector on 1,600 frames of each, three seeds, and grade it on the street with the exam of Lesson 1.

The lab can do this because the street is a program. A project cannot: it has no way to build a hybrid of its program and the real street, and §5 says what it can do instead. The table shows the eight programs that differ in scene, light and camera, with the label left as the program's rule (making it the street's changes the numbers by less than a point, as explained below the table).

SceneLightCameraMiss rate on the streetAUCAlarm rate on unlabelled frames
programprogramprogram80.3%0.596100%
streetprogramprogram74.8%0.645100%
programstreetprogram73.3%0.666100%
programprogramstreet65.8%0.69093.7%
streetstreetprogram68.3%0.694100%
streetprogramstreet60.4%0.73081.8%
programstreetstreet61.6%0.74925.5%
streetstreetstreet46.8%0.79210.7%

Read the first column of numbers down. Making one stage the street's moves the miss rate from 80.3% to 74.8% (scene), 73.3% (light) or 65.8% (camera); making all three the street's reaches 46.8%. No stage is the whole story, and every stage is some of it. The label does not appear: aaab reproduces aaaa to the last digit and bbbb differs from bbba by 0.3 points, because the two label rules disagree only about a pedestrian who is partly hidden, and a program whose scene never hides one never meets the disagreement. The label is not irrelevant, but its cost is not paid in this exam; Lesson 7 shows where it is.

3 · The price list, and why repairs compound

How much is a stage worth? The tempting answer, the gain from repairing it alone, depends on what else is already repaired, so there is no single answer; there is one for each order of repair. Take the average over all of them. For stage i and a set S of stages already repaired, let v(S) be the hit rate (one minus the miss rate) of the program in which exactly the stages of S are the street's. The price of stage i is its gain averaged over every order in which the four stages could be repaired:

φi = ΣS ⊆ stages∖{i} [ |S|! (3 − |S|)! / 4! ] · ( v(S ∪ {i}) − v(S) )

This is the Shapley value of cooperative game theory, and it has the property that matters here: the four prices add up exactly to the whole gain, v(all) − v(none), however the stages interact. From the sixteen programs the prices are:

StagePrice in points of miss rateShare of the 33.2 points
Camera16.349%
Light8.526%
Scene8.425%
Label−0.10%

Each cell is a mean of three seeds and wobbles by about 1 point between them, so the camera is clearly the largest price and the light and the scene are a tie. Behind each price is a table of what the stage buys given the state of the others, and it shows the second fact: repairs compound. The camera is worth 14.5 points when light and scene are both wrong and 21.5 when both are right. The scene is worth 5.5 points while the light and the camera are wrong, 5.4 and 5.0 if only one of them is repaired, and 14.8 once both are. The light behaves the same way (7.0 points at worst, 13.5 at best). Taken together, light and scene are worth 12.0 points with the camera wrong and 18.9 with it right.

One way to see why. A detector learns in the program's pixels, and what it learns reaches the street only through whatever the two sets of pixels have in common. If the camera model is wrong, a new case or a better range of lighting is learned in pixels that no street camera produces, and part of it does not transfer. A pedestrian is found when the case was seen and the light was familiar and the noise was tolerable, a conjunction, and a repair of one conjunct pays little until the others hold; the pixel stages, light and camera, gate everything else the program can teach. The picture is only a picture. If the stages cost independent fractions of the pedestrians, the losses would multiply, and the naive program would find 17.1% of them; it finds 19.7%, because the same pedestrians are lost for several reasons at once. The table is the measurement and the picture only a way to remember its shape.

The switchboard
Choose which stages of the program come from the street (or slide along the order this track repairs them in: camera, light, scene, label). The bars show the miss rate of all sixteen programs, sorted, with the chosen one outlined; the dots under them say which stages are the street's. The panel on the right is the label-free meter: 24 unlabelled street frames, red where the detector of the chosen program alarms at the threshold set on its own program's pedestrian-free frames. The prices are computed live from the table.
miss rate on the street
—
AUC on the street
—
alarm rate, unlabelled frames
—
score gap meter (AUC)
—
Show the core JS
  function code() { return el.cb.map(function (c) { return c.checked ? 'b' : 'a'; }).join(''); }
  function v(c) { return 1 - mean(T.swap[c].realMiss); }
  function shapley() {
    var n = 4, fact = [1, 1, 2, 6, 24], phi = [0, 0, 0, 0];
    for (var i = 0; i < n; i++) for (var m = 0; m < 16; m++) {
      if (m & (1 << i)) continue;
      var s = 0; for (var b = 0; b < n; b++) if (m & (1 << b)) s++;
      phi[i] += fact[s] * fact[n - s - 1] / fact[n] * (v(codeOf(m | (1 << i))) - v(codeOf(m)));
    }
    return phi;
  }

What to try. Start as the page opens, nothing repaired: the program is SIM0, the miss rate is 80.3%, and all 24 unlabelled street frames are red. Slide to camera: 65.8%, the first repair buys 14.5 points, and the meter hardly stirs, 93.7% of unlabelled frames still alarm. Slide to light: 61.6%, only 4.2 points more on the exam, yet the alarm rate drops from 93.7% to 25.5%: most thumbnails go quiet. Slide to scene: 46.8%, 14.8 points, the repair that was worth 5.5 when the camera and the light were wrong. The last notch, label, buys nothing. Now take the other order with the boxes: tick scene alone (74.8%), then light (68.3%), then camera: 46.8%, 21.5 points from the last repair, and the first two bought 5.5 and 6.4. The end point is the same; the road to it is not. In the bars on the left every one of the sixteen programs sits somewhere between the first bar and the last, and the label dots move a bar by less than a point.

4 · A meter that needs no label

The exam needs labelled street frames, and labelled street frames are what a project is trying to avoid buying. But there is a measurement that needs none. A pedestrian is rare in ordinary driving (about one frame in fifty here), so a stream of unlabelled street frames is, to a good approximation, a stream of pedestrian-free frames. Take the threshold the program chose for itself, the one that makes one in ten of its own pedestrian-free frames alarm, and apply it to those street frames. If the program matched the street the alarm rate would be about ten per cent, because the street's frames would score as the program's do. Any excess is the gap, seen without a label.

The column on the right of the table above is that alarm rate over 1,000 unlabelled frames. The naive program scores 100%: every street frame, nearly, scores above a threshold that a tenth of the program's own frames exceed, because the street's frames carry clutter, grain and light the detector has never seen and treats as evidence. The matched program scores 10.7%. A finer version of the same meter compares the two distributions of scores directly: the probability that a street frame outscores a pedestrian-free frame of the program, 1.00 for the naive program and 0.50 for the matched one, where 0.5 means the two cannot be told apart. Across all sixteen programs this meter and the real miss rate agree in rank to a Spearman correlation of 0.98: a program that looks wrong on unlabelled frames is wrong on the exam.

Two cautions keep it honest. The meter says that the gap exists and how large it is, not which stage it lives in: it was blind to the camera repair alone (100% to 93.7%) and woke only when the light was repaired as well. And it saturates: at 100% it cannot tell a bad program from a hopeless one. It is a cheap reading to take at every step of the work, not a replacement for the exam.

5 · What evidence answers to which stage, and what to repair first

Each stage leaves a different trace in the data a project can get, and the traces differ in cost.

Stage (price)Evidence that sees itWhat it costsWhat it cannot tell you
Camera (16.3)flat frames of a grey card at several light levels; a sharp edge; the exposure and gain the camera logsan afternoon with the camera; no labels, no streetanything about what stands in the street
Light (8.5)unlabelled street frames: pixels whose content is known (sky, road) and ratios that exposure control cancelsfree, from logs the fleet already collectswhat the scene holds; how a rare pedestrian is dressed
Scene (8.4)a few hundred labelled frames and the ground plane, which turns a mask into a distance; logs, for how often things happenpeople labelling frameshow the pictures are lit or recorded
Label (−0.1)the benchmark's written definitionsreading themwhether a detector will meet the disagreement

Price says what a repair is worth; the table says what it costs to learn what to repair. The camera is first on the first count and first on the second. Its repair is also the one that most raises the value of the others: light and scene together are worth 12.0 points before it and 18.9 after. And it gives the clearest feedback, because it is the repair whose effect on the exam is visible at once. The scene stage, whose evidence is the most expensive, can wait until its payoff is large enough to see.

Road not taken · repair in the order the program runs
The obvious order is the program's own: scene, then light, then camera, each repair feeding the next. Measured, that road buys 5.5 points, then 6.4, then 21.5. A team on it spends its labelling budget on the scene first and sees the exam move by about five points. The price-ordered road buys 14.5, 4.2 and 14.8, and the first step needs no labels. Both roads arrive at the same 46.8%, which is why the order is a matter of cost and feedback rather than of correctness. Where a case is missing altogether no repair downstream can supply it, so the scene stage cannot be skipped; it can only wait, and Lessons 5 and 6 return to it.
What this lesson did not do
It repaired nothing: every program above is the program or the street, with no stage in between. The prices are for this learner, this exam and this street, averaged over orders of repair that nobody would follow, and they say where the effort pays, not how much any one parameter is worth inside a stage. The sixteen-way swap is a laboratory privilege, and the evidence table is only a list: Lessons 3 and 4 do the first two rows, Lessons 5 and 6 the third, Lesson 7 the fourth. The meter was never calibrated against a probability of anything; it is a ranking.

Common mistakes / failure modes

"Fix whatever looks least realistic"
A person judging pictures cannot see what a detector leans on. The camera is the stage nobody notices in a thumbnail, and it costs the most (§3).
"A repair that moves nothing was wasted"
Repairs compound. The scene repair moved the exam by 5.5 points while the pixels were wrong and by 14.8 once they were right; judge a repair at the state where it can pay (§3).
"The price of a stage is what it buys alone"
That is the price for one order of repair. The Shapley price averages over all orders and adds up to the whole gain (§3).
"The label stage does not matter"
It moves no detector here, but it changes the number the exam reports, and an exam that counts a different thing is a different exam (§2; Lesson 7).
"The alarm rate is a miss-rate estimate"
It is a gap detector. It saturates at 100%, it was blind to the camera repair alone, and it says nothing about which stage is wrong (§4).
"Repair in pipeline order, since each stage feeds the next"
Pipeline order is correct for building the program and a poor order for repairing it: the early repairs pay little and cost the most (§5).

Checkpoint exercise

Try it
A program has two stages, A and B, and the hit rates of the four programs are: neither repaired 0.20, only A 0.30, only B 0.25, both 0.50. What is the Shapley price of each stage, and what does the table say about how to order the repairs? Answer: with two stages the price of A is the average of its gain when repaired first (0.30 − 0.20 = 0.10) and last (0.50 − 0.25 = 0.25), 0.175; the price of B is the average of 0.05 and 0.20, 0.125; together 0.30, which is the whole gain 0.50 − 0.20. Each stage is worth more once the other is repaired (A: 0.10 against 0.25), so the repairs compound; A has the higher price and goes first unless the evidence for B is much cheaper to collect.

Where this points next

The swap table prices the stages, and it is a laboratory privilege: a project holds the street's frames and cannot hold the street's program. What it can hold is the traces of §5, and they are not equally cheap. The camera is the most expensive gap, 16.3 of the 33.2 points, and the one whose evidence is cheapest, because a grey card in front of the lens answers a question about the camera and nothing else. Repairing it raises what the rest are worth, from 12.0 points to 18.9, and the repair shows on the exam at once. The camera is also the stage the program treats most casually: it writes the renderer's light straight into the image as if it were the number the sensor reports. How does a camera turn light into numbers, and how do we measure it well enough to simulate it?

Takeaway
The program is four stages, scene, light, camera and label, run in that order, and the gap between a detector trained on the program and one trained on the street, 33.2 points of miss rate at 10% false alarms, divides among them. Swapping stages between the program and the street gives sixteen programs and a price list: the camera 16.3 points, the light 8.5, the scene 8.4, the label nothing. The repairs compound, because a detector learns in the program's pixels and what it learns reaches the street only where the pixels agree, so the pixel stages gate the value of everything else. A meter exists that needs no label, the alarm rate on unlabelled street frames at a threshold set on the program, and it tracks the gap by rank but cannot say which stage. Price and the cost of evidence put the camera first: it is the largest gap and the cheapest to measure.

Interview prompts

Companion reads: Computer Vision · 04 Cameras and projection geometry (the camera as a function, which Lesson 3 turns into a measurement), and Computer Graphics, from first principles (the stages the program is made of, run forwards).