Free labels, and the wall
Computer graphics taught the renderer as a function: a scene goes in, an image comes out. A program that draws a picture also knows what is in it, so it can write the answer beside the picture, exactly and in any number, for nothing, where vision has always paid people to write it. This lesson builds the smallest honest laboratory for that bargain: a street, a shuttle's camera, a detector that never changes and an exam. It trains the detector on the program's pictures and gets two curves that start together and then part. The error on the program's own pictures falls to nothing within forty frames; the error on the real street settles at a floor, and sixty-four times more frames do not lower it. The rest of the track is about what that floor is made of.
New idea: a renderer returns the answer with the picture, so labels are free and exact, but exact for the program's own world. Train on them and the error on the program's pictures goes to nothing; the error on real pictures does not, because the gap between the program and the street is bias, and more data removes variance, not bias.
Forces next: Free labels are exact and unlimited, and a detector trained on them stops making mistakes on the program's own pictures after about forty frames; on pictures of the real street it still misses 80% of the pedestrians, whether it saw forty frames or twenty-five hundred. The gap is not noise that more data averages away: it is a difference between the program and the world, and the program has four places to be wrong: how the camera records the scene, how the scene is lit, what exists in it, and what the label means. Which of them costs the most, and how could we tell without a real label for every frame?
1 · One render, five answers
A real camera turns a scene into numbers and keeps nothing else. A renderer is the same map written as a program, and the program has the scene in hand while it draws. Every pixel is made by tracing a ray from the camera until it hits something, so at the moment the colour is computed the program also knows what was hit, how far away it was, and what lies behind it. It has only to write those facts down.
The world of this track is the Street: a plane of vertical objects seen from a camera that a shuttle carries 1.4 m above the road. The objects are a parked van (a box), poles and trees (cylinders), a far wall as a backdrop, and a pedestrian made of three stacked cylinders, legs, torso and head, who either stands in the open or steps out from behind the van. The picture is 96 × 24 pixels with three colour channels, taken by a pinhole of focal length f = 83.1 px (60° across) with the horizon on row 9. A ray through a pixel points in the direction (tan φ, ρ, 1), with tan φ = (u + ½ − 48)/f across and ρ = (9 − v − ½)/f up. Depth z is the distance along the optical axis; range r is the straight-line distance to the point, and the two differ by the length of that direction vector:
r = z · √(1 + tan²φ + ρ²)
On the axis the factor is 1; at the edge of the field of view, where φ = 30°, it is 1.155. Nobody has to guess any of this: for every pixel the renderer solves the intersection of that ray with the objects, and keeps the answer. One render therefore yields five things:
| Answer | What the renderer writes | What it takes to get it from a photograph |
|---|---|---|
| the image | the colour of every pixel, after the camera model | taking the photograph |
| the mask | for every pixel, the fraction of its area the pedestrian covers (four rays per pixel) | a person outlining the pedestrian |
| the depth | z at the first hit along the pixel's centre ray | a ranging sensor, calibrated to the camera and aligned in time |
| the range | r at the same hit | the same, plus a convention for which of the two distances you mean |
| the silhouette | the mask of the pedestrian alone, as if nothing stood in front | nobody can see it; an annotator guesses the hidden outline |
The word exact needs care. These labels are exact for the program's own definitions: coverage is counted with four rays per pixel, depth is read at the pixel's centre, a pedestrian "counts" once enough of the silhouette shows. For a pedestrian a few pixels wide four rays are a coarse count: the partly hidden pedestrian of street frame 1 covers 15 px² at four rays per pixel and 19 at sixty-four, and the rule that decides whether the pedestrian counts is applied to the first number. A different renderer, or a person with a polygon tool, may define each of these a little differently, and Lessons 7 and 8 measure how much that matters. For the purposes of this lesson the labels are simply right: nothing a detector could do with them is limited by their noise.
2 · The laboratory: a program, a street, a detector, an exam
The program and the street. The first program a team writes, here called SIM0, is naive in the ordinary way. The sun is at noon behind the camera, the van is white and 8–12 m away, shirts are bright and trousers dark, the sky is blue, the camera is ideal (no grain, no blur), every pedestrian stands in the open between 8 and 16 m, and any visible pixel of a pedestrian makes the pedestrian count. The street is a second, richer program that the lessons may only sample, never read: low sun and glare, vans of every colour, shirts of any colour, grey and warm skies, a camera with photon noise, blur and automatic exposure, pedestrians out to 24 m or more, 28% of them stepping out from behind the van, and a rule that a pedestrian counts only once 15% of the silhouette shows. It is a program too, and that is deliberate: a program lets us swap one of its parts for the program's own, which no project can do with a real street. Every claim about what a project can know is held to what sampling the street allows: a labelled test set, a small labelled budget, and unlabelled frames.
The detector. It is fixed for the whole track. Fifty-four filters, vertical bars and steps at three widths and three heights, are applied to luminance and to two colour-difference channels and read in both polarities; with the row of the cell and its square, every cell of a stride-2 grid gets 110 numbers. A logistic head turns them into one score per cell, and the score of a frame is the largest cell score in it. The head is trained on the exact masks (a cell is positive where its pixel is at least half pedestrian) and on one round of its own false alarms. It has 111 numbers to learn and trains in seconds. Nothing about it changes from lesson to lesson, so whatever moves a number later is something that happened to the data.
The exam. A score means nothing without a threshold, and a miss rate means nothing without the threshold that produced it. So the exam fixes the operating point first: take frames of the world being graded that contain no pedestrian, set the threshold so that 10% of them raise an alarm, and then count the pedestrians whose frames score below it. The miss rate is that fraction, taken over pedestrians with at least 6 visible pixels of area, 1,323 of them in the 3,000 street frames the lessons use. The threshold is chosen on separate validation frames of the same world. It is the most generous honest choice there is: a detector gets the operating point that suits the world it is graded on.
3 · Two curves
The experiment is the obvious one. Draw N frames from the program and train the detector. Grade it twice: on fresh frames from the program, and on frames from the street. Repeat for N from 10 to 2,560 and for three random seeds, and average. For comparison, train the same detector on N frames of the street and grade it on the street. The widget plots all three, and lets you look at what a detector trained on N frames does to a single frame of each world.
What to try. Start as the page opens: N = 160 program frames, frame 0, the program-trained detector. On the program's frame the score map has one strong cell on the pedestrian and the frame is found; across the 24 program frames 24 of 24 are found. Look at the street's frame beside it: the same detector's map is busy everywhere and the pedestrian is not where the largest score is, and across the 24 street frames only 1 of 24 is found, even though the threshold has been set on street frames to give it every chance. Slide N from 10 to 2,560. On the program's own frames the miss rate is 0.9% at N = 10 and 0.0% from N = 40 on. On the street it starts at 84.6%, reaches 80.2% at 40 frames and is still 80.2% at 2,560, sixty-four times more. Now switch the detector to the one trained on the street. It is already better at 10 frames, 74.5% against 84.6%, and the gap only widens: 66.8% at 40 frames and 47.7% at 2,560. Last, switch free labels of to the street's frame and step through the frames: when the van hides part of a pedestrian the silhouette is larger than the mask (frame 1 has 15 pixels of visible area and a silhouette of 41), a label no photograph can give.
4 · Why the second curve does not move: variance and bias
Write riskP(h) for the miss rate of a detector h on world P under the exam. Let ĥN be what training returns after N frames of the program S, and let h*S and h*R be the best detectors of this kind for the program and for the street, the limits of training on endless frames of each. Then on the street
riskR(ĥN) − riskR(h*R) = [riskR(ĥN) − riskR(h*S)] + [riskR(h*S) − riskR(h*R)]
The first bracket is variance: how far a detector trained on N frames is from the best one that program could ever produce. It shrinks as N grows, because the program is a fixed source and more frames pin its pattern down. The second is bias: how much worse the best program-trained detector is, on the street, than the best street-trained one. It has no N in it. The curves measure both. Take 1,600 frames, the size the rest of the track uses, as "enough". The program-trained detector's street miss rate falls by only 4.3 points between N = 10 and N = 1,600, which is all the variance there was to remove, and what stays is a bias of 33.2 points above what the street-trained detector reaches at the same size (it has all but stopped improving by 1,280 frames, so it stands in for h*R).
The street-trained detector has the opposite problem. Between 10 and 1,600 frames its miss rate falls by 27.4 points, almost all variance, and it is still falling at 40 frames, where it stands at 66.8%. So the two kinds of data differ in which half of the error they leave: real labels are expensive, so there are few of them, and their error is variance; free labels are plentiful, so their variance is gone, and what is left is bias. Free labels buy you out of exactly one half of the error.
The matched ceiling is also not zero. Trained on the street itself the detector misses 47.1% of the pedestrians who count at 1,600 frames, 40.2% of those who stand in the open and 80.4% of those who step out from behind the van, who are 17% of the pedestrians who count. The program-trained detector at the same size misses 78.9% and 87.0%. So the street has its own difficulty, mostly the partly hidden pedestrian, and the program's pictures differ from the street's by much more than that.
5 · What the wall is not
There are four obvious suspects, and each has an alibi in the curves.
- The learner. The same detector reaches 0.0% on the program and 47.7% on the street. It can represent both patterns.
- Too few frames. The street miss rate is flat from 40 frames to 2,560.
- Noisy labels. They are exact. Every mask pixel is a pixel of the pedestrian by construction.
- Overfitting. On held-out frames of the program the detector misses nothing. It generalises perfectly, to the program.
What is left is that the program's pictures and the street's are not drawn from the same distribution, and a detector learns what it is shown. The labels are exact answers to the program's questions. The pedestrian of the street is a different object in every way a filter can feel: what they wear, how they are lit, how a camera grains them, and how often a van hides half of them.
The same fact is visible without a single label. Choose the threshold that makes one in ten of the program's pedestrian-free frames raise an alarm and carry it, unchanged, to pedestrian-free frames of the street. For the program-trained detector 100% of them alarm; for the street-trained detector 9.1% do, which is what the threshold was meant to give. No human labelled anything to produce that comparison, and Lesson 2 turns it into a meter.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
The wall is a difference between two programs, the one we wrote and the one the street runs, and a difference between programs is a list of differences between their parts. The program has four: the scene decides what exists and where, the light decides what colour it is, the camera turns light into numbers, and the label says what counts. A project cannot swap one of these between its program and the street to see which matters, because the street is not a program it can read. The lab can, and the next lesson does it: sixteen programs, each a different mixture of the naive one and the street, each trained and graded the same way. Free labels have removed the variance and left the 33.2 points of bias where they were, and those points have to be divided among the stages before any stage can be repaired. Which of them costs the most, and how could we tell without a real label for every frame?
Interview prompts
- What does a renderer give you that a camera does not, and in what sense is it "exact"? (§1 — the program has the scene in hand, so mask, depth, range and silhouette are read off the intersection, exact for the program's own definitions of coverage, depth and "counts".)
- A model trained on synthetic data scores 99.9% on synthetic validation and 40% on real data. Is it overfitting? (§4, §5 — no: the held-out synthetic score is the proof it generalises, to the program; the gap is bias, a difference between distributions, not variance.)
- Decompose the excess error on the real distribution of a model trained on synthetic data. (§4 — variance: the distance from the best program-trained detector, shrinking with N; bias: the distance between the best program-trained and best real-trained detector, independent of N.)
- Why does a learning curve that is flat on a log axis tell you something about the next batch of data? (§3, §4 — the variance term is already spent, so more frames of the same program cannot lower the error.)
- Why is the real-trained learning curve steep where the synthetic one is flat? (§4 — the real-trained model has few frames, so its error is mostly variance; the synthetic one has plenty, so what is left is bias.)
- Why quote a miss rate with its threshold? (§2 — it is a property of detector, world and operating point; the same detector can be made to miss anything or nothing by moving the threshold.)
- How can you detect that a synthetic training set does not match reality without labelling any real data? (§5 — the alarm rate on unlabelled real frames at a threshold chosen on synthetic ones: pedestrians are rare, so almost every frame is pedestrian-free and should alarm at the chosen rate.)
Companion reads: Computer Graphics, from first principles (the renderer run forwards, which this lesson uses as a labelling machine), Computer Vision · 04 Cameras and projection geometry (the pinhole model the Street uses) and 3D Vision · 01 A pixel is a ray (what depth and range mean for a pixel).