all_lessons/Synthetic Vision Data/01 · Free labels, and the walllesson 1 / 12

Free labels, and the wall

Computer graphics taught the renderer as a function: a scene goes in, an image comes out. A program that draws a picture also knows what is in it, so it can write the answer beside the picture, exactly and in any number, for nothing, where vision has always paid people to write it. This lesson builds the smallest honest laboratory for that bargain: a street, a shuttle's camera, a detector that never changes and an exam. It trains the detector on the program's pictures and gets two curves that start together and then part. The error on the program's own pictures falls to nothing within forty frames; the error on the real street settles at a floor, and sixty-four times more frames do not lower it. The rest of the track is about what that floor is made of.

The thesis, here
A model learns the pattern in its data and nothing else. The data of a program is a pattern its author chose, so what a model trained on it is worth on the street is decided by how well that choice stands in for the street. The track asks, in the order the evidence allows, what the program must get right, what each part costs, how to measure it without labelling the street, and how to know afterwards. This lesson shows that the question is real: free labels do not make it go away, and no number of them does.
Linear position
Forced by: Computer graphics taught the renderer as a function: a scene goes in, an image comes out, and the program that draws the image knows exactly what is in it. A vision model needs the opposite service, images that arrive with the answer, by the thousand, for nothing, and a renderer provides exactly that. What does a model learn from pictures a program drew, and how would we find out whether it learned the street or only the program?
New idea: a renderer returns the answer with the picture, so labels are free and exact, but exact for the program's own world. Train on them and the error on the program's pictures goes to nothing; the error on real pictures does not, because the gap between the program and the street is bias, and more data removes variance, not bias.
Forces next: Free labels are exact and unlimited, and a detector trained on them stops making mistakes on the program's own pictures after about forty frames; on pictures of the real street it still misses 80% of the pedestrians, whether it saw forty frames or twenty-five hundred. The gap is not noise that more data averages away: it is a difference between the program and the world, and the program has four places to be wrong: how the camera records the scene, how the scene is lit, what exists in it, and what the label means. Which of them costs the most, and how could we tell without a real label for every frame?
The plan
Five moves. (1) One render, five answers: see what a program hands over with a picture. (2) Fix the laboratory for the whole track: the Street, the program, a detector that never changes, and the exam. (3) Train on the program's pictures and grade on two worlds: two curves. (4) Split the gap into what data removes (variance) and what it does not (bias), and find what the same detector reaches when it trains on the street itself. (5) Say what the wall is not.

1 · One render, five answers

A real camera turns a scene into numbers and keeps nothing else. A renderer is the same map written as a program, and the program has the scene in hand while it draws. Every pixel is made by tracing a ray from the camera until it hits something, so at the moment the colour is computed the program also knows what was hit, how far away it was, and what lies behind it. It has only to write those facts down.

The world of this track is the Street: a plane of vertical objects seen from a camera that a shuttle carries 1.4 m above the road. The objects are a parked van (a box), poles and trees (cylinders), a far wall as a backdrop, and a pedestrian made of three stacked cylinders, legs, torso and head, who either stands in the open or steps out from behind the van. The picture is 96 × 24 pixels with three colour channels, taken by a pinhole of focal length f = 83.1 px (60° across) with the horizon on row 9. A ray through a pixel points in the direction (tan φ, ρ, 1), with tan φ = (u + ½ − 48)/f across and ρ = (9 − v − ½)/f up. Depth z is the distance along the optical axis; range r is the straight-line distance to the point, and the two differ by the length of that direction vector:

r = z · √(1 + tan²φ + ρ²)

On the axis the factor is 1; at the edge of the field of view, where φ = 30°, it is 1.155. Nobody has to guess any of this: for every pixel the renderer solves the intersection of that ray with the objects, and keeps the answer. One render therefore yields five things:

AnswerWhat the renderer writesWhat it takes to get it from a photograph
the imagethe colour of every pixel, after the camera modeltaking the photograph
the maskfor every pixel, the fraction of its area the pedestrian covers (four rays per pixel)a person outlining the pedestrian
the depthz at the first hit along the pixel's centre raya ranging sensor, calibrated to the camera and aligned in time
the ranger at the same hitthe same, plus a convention for which of the two distances you mean
the silhouettethe mask of the pedestrian alone, as if nothing stood in frontnobody can see it; an annotator guesses the hidden outline

The word exact needs care. These labels are exact for the program's own definitions: coverage is counted with four rays per pixel, depth is read at the pixel's centre, a pedestrian "counts" once enough of the silhouette shows. For a pedestrian a few pixels wide four rays are a coarse count: the partly hidden pedestrian of street frame 1 covers 15 px² at four rays per pixel and 19 at sixty-four, and the rule that decides whether the pedestrian counts is applied to the first number. A different renderer, or a person with a polygon tool, may define each of these a little differently, and Lessons 7 and 8 measure how much that matters. For the purposes of this lesson the labels are simply right: nothing a detector could do with them is limited by their noise.

2 · The laboratory: a program, a street, a detector, an exam

The program and the street. The first program a team writes, here called SIM0, is naive in the ordinary way. The sun is at noon behind the camera, the van is white and 8–12 m away, shirts are bright and trousers dark, the sky is blue, the camera is ideal (no grain, no blur), every pedestrian stands in the open between 8 and 16 m, and any visible pixel of a pedestrian makes the pedestrian count. The street is a second, richer program that the lessons may only sample, never read: low sun and glare, vans of every colour, shirts of any colour, grey and warm skies, a camera with photon noise, blur and automatic exposure, pedestrians out to 24 m or more, 28% of them stepping out from behind the van, and a rule that a pedestrian counts only once 15% of the silhouette shows. It is a program too, and that is deliberate: a program lets us swap one of its parts for the program's own, which no project can do with a real street. Every claim about what a project can know is held to what sampling the street allows: a labelled test set, a small labelled budget, and unlabelled frames.

The detector. It is fixed for the whole track. Fifty-four filters, vertical bars and steps at three widths and three heights, are applied to luminance and to two colour-difference channels and read in both polarities; with the row of the cell and its square, every cell of a stride-2 grid gets 110 numbers. A logistic head turns them into one score per cell, and the score of a frame is the largest cell score in it. The head is trained on the exact masks (a cell is positive where its pixel is at least half pedestrian) and on one round of its own false alarms. It has 111 numbers to learn and trains in seconds. Nothing about it changes from lesson to lesson, so whatever moves a number later is something that happened to the data.

The exam. A score means nothing without a threshold, and a miss rate means nothing without the threshold that produced it. So the exam fixes the operating point first: take frames of the world being graded that contain no pedestrian, set the threshold so that 10% of them raise an alarm, and then count the pedestrians whose frames score below it. The miss rate is that fraction, taken over pedestrians with at least 6 visible pixels of area, 1,323 of them in the 3,000 street frames the lessons use. The threshold is chosen on separate validation frames of the same world. It is the most generous honest choice there is: a detector gets the operating point that suits the world it is graded on.

3 · Two curves

The experiment is the obvious one. Draw N frames from the program and train the detector. Grade it twice: on fresh frames from the program, and on frames from the street. Repeat for N from 10 to 2,560 and for three random seeds, and average. For comparison, train the same detector on N frames of the street and grade it on the street. The widget plots all three, and lets you look at what a detector trained on N frames does to a single frame of each world.

The same detector, two worlds
Top row: a frame from the program (left) and one from the street (right), each with the detector's score map beneath it; the boxed cell is the largest score, which is the frame's score, outlined cells are above the threshold, and the dashed teal box is where the pedestrian really is. Middle row: the free labels of either frame. Bottom: the learning curves, with the number of training frames marked. The curves are measured (three seeds each); the score maps are computed live from the trained weights.
miss, program frames
—
miss, street frames
—
street-trained, street
—
found, 24 + 24 frames
—
Show the core JS
SV.thrAtFPR = function (negScores, alpha) { var s = negScores.slice().sort(function (a, b) { return a - b; }); return s[Math.min(s.length - 1, Math.floor((1 - alpha) * s.length))]; };
SV.rate = function (scores, thr) { var k = 0; for (var i = 0; i < scores.length; i++) if (scores[i] > thr) k++; return k / scores.length; };
SV.exam = function (set, scores, o) {
  var pos = [], neg = [], i;
  for (i = 0; i < set.length; i++) { if (set[i].y) pos.push(i); else neg.push(i); }
  var negS = neg.map(function (j) { return scores[j]; });
  if (thr === undefined) thr = SV.thrAtFPR(negS, fa);
  pos.forEach(function (j) {
    var a = set[j].area;
    if (a >= minA) { n++; if (scores[j] > thr) hit++; }
  });
  return { thr: thr, miss: 1 - hit / n, ... };
};
  function mean(a) { var s = 0; for (var k = 0; k < a.length; k++) s += a[k]; return s / a.length; }
  var found = function (arr, thr) { var k = 0; arr.forEach(function (v) { if (v > thr) k++; }); return k; };

What to try. Start as the page opens: N = 160 program frames, frame 0, the program-trained detector. On the program's frame the score map has one strong cell on the pedestrian and the frame is found; across the 24 program frames 24 of 24 are found. Look at the street's frame beside it: the same detector's map is busy everywhere and the pedestrian is not where the largest score is, and across the 24 street frames only 1 of 24 is found, even though the threshold has been set on street frames to give it every chance. Slide N from 10 to 2,560. On the program's own frames the miss rate is 0.9% at N = 10 and 0.0% from N = 40 on. On the street it starts at 84.6%, reaches 80.2% at 40 frames and is still 80.2% at 2,560, sixty-four times more. Now switch the detector to the one trained on the street. It is already better at 10 frames, 74.5% against 84.6%, and the gap only widens: 66.8% at 40 frames and 47.7% at 2,560. Last, switch free labels of to the street's frame and step through the frames: when the van hides part of a pedestrian the silhouette is larger than the mask (frame 1 has 15 pixels of visible area and a silhouette of 41), a label no photograph can give.

Road not taken · make the program photorealistic
The obvious reply to the street's frame above is that the program should simply look more like it. That is not one knob: a picture depends on what is in the scene, how it is lit, how the camera records it, and what the label calls a pedestrian, and a person judging a picture by eye cannot tell which of these a detector is leaning on. Lesson 2 splits the gap into those four parts and prices each, and Lessons 3 to 7 repair them in the order the prices and the evidence suggest. Lesson 8 then puts the program's contract into a production renderer, which does look photographic, and checks what that changes.

4 · Why the second curve does not move: variance and bias

Write riskP(h) for the miss rate of a detector h on world P under the exam. Let ĥN be what training returns after N frames of the program S, and let h*S and h*R be the best detectors of this kind for the program and for the street, the limits of training on endless frames of each. Then on the street

riskR(ĥN) − riskR(h*R) = [riskR(ĥN) − riskR(h*S)] + [riskR(h*S) − riskR(h*R)]

The first bracket is variance: how far a detector trained on N frames is from the best one that program could ever produce. It shrinks as N grows, because the program is a fixed source and more frames pin its pattern down. The second is bias: how much worse the best program-trained detector is, on the street, than the best street-trained one. It has no N in it. The curves measure both. Take 1,600 frames, the size the rest of the track uses, as "enough". The program-trained detector's street miss rate falls by only 4.3 points between N = 10 and N = 1,600, which is all the variance there was to remove, and what stays is a bias of 33.2 points above what the street-trained detector reaches at the same size (it has all but stopped improving by 1,280 frames, so it stands in for h*R).

The street-trained detector has the opposite problem. Between 10 and 1,600 frames its miss rate falls by 27.4 points, almost all variance, and it is still falling at 40 frames, where it stands at 66.8%. So the two kinds of data differ in which half of the error they leave: real labels are expensive, so there are few of them, and their error is variance; free labels are plentiful, so their variance is gone, and what is left is bias. Free labels buy you out of exactly one half of the error.

The matched ceiling is also not zero. Trained on the street itself the detector misses 47.1% of the pedestrians who count at 1,600 frames, 40.2% of those who stand in the open and 80.4% of those who step out from behind the van, who are 17% of the pedestrians who count. The program-trained detector at the same size misses 78.9% and 87.0%. So the street has its own difficulty, mostly the partly hidden pedestrian, and the program's pictures differ from the street's by much more than that.

5 · What the wall is not

There are four obvious suspects, and each has an alibi in the curves.

What is left is that the program's pictures and the street's are not drawn from the same distribution, and a detector learns what it is shown. The labels are exact answers to the program's questions. The pedestrian of the street is a different object in every way a filter can feel: what they wear, how they are lit, how a camera grains them, and how often a van hides half of them.

The same fact is visible without a single label. Choose the threshold that makes one in ten of the program's pedestrian-free frames raise an alarm and carry it, unchanged, to pedestrian-free frames of the street. For the program-trained detector 100% of them alarm; for the street-trained detector 9.1% do, which is what the threshold was meant to give. No human labelled anything to produce that comparison, and Lesson 2 turns it into a meter.

What this lesson did not do
It did not say which part of the program is wrong: the gap of 33.2 points is a sum of causes, and the curves cannot separate them. It treated the street as a program we can sample but not read, and gave no way of measuring anything about it except by grading a detector. It did not price the labels or the frames: the street-trained detector needed real labelled frames, and how many a project can afford, and what the cases that matter cost in them, is where Lessons 5, 6 and 10 return. Lesson 2 opens the program.

Common mistakes / failure modes

"A high synthetic score means a good detector"
It means the detector learned the program: 0.0% miss on the program's frames sits next to 80% on the street (§3).
"More frames will close the gap"
More frames shrink variance. The street miss rate of the program-trained detector moves by 4.3 points between 10 and 1,600 frames; the bias is 33.2 (§4).
"Exact labels mean perfect data"
Exact for the program's definitions and its pictures. The labels answer questions the program asked; they cannot say whether the program asked the street's questions (§1, §5).
"If it looks right to a person, it will transfer"
A detector leans on whatever separates classes in its data, which a person judging by eye cannot see. Realism is judged by the exam, not by taste (Road not taken).
"The street-trained detector shows the ceiling is zero error"
It misses 47.1% at 1,600 frames, 80.4% of the pedestrians who step out from behind the van. The street has its own difficulty (§4).
"A miss rate is a property of a detector"
It is a property of a detector, a world and a threshold. Quote one without the other two and the number means nothing (§2).

Checkpoint exercise

Try it
A team trains on 20,000 synthetic frames and reports that, at 10% false alarms, the detector misses 0.1% of pedestrians in a held-out synthetic set. On 300 labelled real frames, with the threshold set to 10% false alarms on those frames, it misses 62%. They propose to render 200,000 more frames. What will that buy, and which two measurements would tell the team how much of the 62% is bias? Answer: almost nothing. At 20,000 frames the variance term is already gone, as it was by 40 frames in the lab (§3), so ten times more frames moves the real miss rate by a point or two at most, as 64 times more moved it by 0.0 points here. The 62% is bias, and two measurements size it. One is the real miss rate of a detector trained on real frames, however few: the street-trained detector at 40 frames already stood at 66.8%, 13.4 points better than the program-trained one at 2,560, which makes the bias at least that large, and more as the real-trained curve is followed up. The other is the alarm rate on unlabelled real frames at the threshold chosen on synthetic ones (§5), which needs no labels at all.

Where this points next

The wall is a difference between two programs, the one we wrote and the one the street runs, and a difference between programs is a list of differences between their parts. The program has four: the scene decides what exists and where, the light decides what colour it is, the camera turns light into numbers, and the label says what counts. A project cannot swap one of these between its program and the street to see which matters, because the street is not a program it can read. The lab can, and the next lesson does it: sixteen programs, each a different mixture of the naive one and the street, each trained and graded the same way. Free labels have removed the variance and left the 33.2 points of bias where they were, and those points have to be divided among the stages before any stage can be repaired. Which of them costs the most, and how could we tell without a real label for every frame?

Takeaway
A renderer writes the answer next to the picture: mask, depth, range and silhouette are free and exact, for the program's own definitions. Trained on them, the fixed detector is perfect on the program's frames after about 40 and misses 80% of the street's pedestrians at every training size, because the error of a model has a variance half that data removes and a bias half, the difference between the program and the street, that it does not. Free labels buy out the first half only; the street-trained detector, which is variance-limited, reaches 47% and is still improving at 40 frames. The wall is not the learner, the data size, the label noise or overfitting: it is that the program's pictures are not the street's. That is visible without any label, in the alarm rate on unlabelled frames at a threshold set on the program.

Interview prompts

Companion reads: Computer Graphics, from first principles (the renderer run forwards, which this lesson uses as a labelling machine), Computer Vision · 04 Cameras and projection geometry (the pinhole model the Street uses) and 3D Vision · 01 A pixel is a ray (what depth and range mean for a pixel).