Cover the cases
A detector is right only where its data has pedestrians, and the scene decides which pedestrians its data has. This lesson turns "which cases does the scene never draw?" into a count: a few hundred outlined frames and the ground plane give every pedestrian a distance and a hiding test, and the same ruler on the program's own frames gives the other side of the table. The program draws cells that hold 42.0% of the street's pedestrians. The count says which kinds of case are missing, not how often the rare ones occur.
New idea: a case is a cell of the situations the exam asks about, and a detector is right only in cells its data fills: count the street's pedestrians per cell with a ruler built from outlines and the ground plane, and count the program's with the same ruler. The program fills cells that hold 42.0% of them.
Forces next: Counting cases on a small labelled sample, with each pedestrian's distance read off the foot row of its outline and a nearer van touching it as the sign of a step-out, shows that the scene program draws cells holding only 42% of the street's pedestrians: 51% of them stand beyond the 16 m it ever drew and 17% step out from behind a van it never lets hide anyone. Setting the distance range and the step-out probability from 200 outlined frames takes the miss rate from 62% to about 52%, as far as the street's own pedestrian settings do, and the clutter, once the pedestrians are right, takes it to 47%. But a cell is a box, not a frequency: the step-outs nearest the shuttle, which leave the least time to stop, occur in one frame in 122, so a training set of 1,600 frames holds about 13 of them, and the detector trained with the street's scene misses 57% of them against 10% of the pedestrians as near in the open. How should a generator spend its renders on cases that rarely happen, and what does that do to the statistics the model learns?
1 · A detector is right where its data has pedestrians
The entry baton compares two prices of one repair: the street's scene bought 5.5 points while the camera and light were wrong and buys 14.8 now (with the street's camera and light, abba misses 61.6%; with the street's scene as well, bbba, 46.8%). It also records one thing about the program's scene: every pedestrian it drew stood in the open and within 16 m. Sort the 1,323 pedestrians the exam counts in the street's 3,000 test frames the way that sentence does, and grade two detectors, trained on 1,600 frames (three seeds), on each group.
| The pedestrian… | Share of the street's | Share the program draws | Missed by the detector trained on the program (abba) | Missed by the one trained with the street's scene (bbba) |
|---|---|---|---|---|
| stands in the open, within 16 m | 40.7% | all of them | 30.7% | 15.6% |
| stands in the open, beyond 16 m | 42.0% | none | 80.9% | 63.4% |
| steps out from behind a van | 17.2% | none | 87.4% | 80.4% |
The groups use the street's own record of each pedestrian's distance and of whether the pedestrian steps out, which only the lab has; §2 builds what a project can use. Both detectors are best where the program has pedestrians and worst where it has none, but the second, which has all three kinds, shows the same order at lower levels, so part of the order is difficulty: a pedestrian at 20 m is a few pixels wide, a hidden one a strip. What the data decides is better read from a repair than from a comparison, and §5 adds the street's settings one at a time.
The data also decides something no later stage can undo. Lesson 2 said that a case the scene never draws cannot be supplied downstream, and the table cells show it: repairing the camera and the light took the open pedestrians' miss rate from 78.9% (aaaa) to 56.2% (abba) and left the step-outs' at 87.0% and 87.4%.
Call a case a cell of the situations the exam asks about. Benchmarks cut their exams this way already: Caltech (Dollár et al., 2012) reports pedestrians by pixel height, near at 80 px or more, medium 30 to 80 and far under 30, and KITTI (Geiger et al., 2012) reports easy, moderate and hard objects, defined by box height, occlusion and truncation. Here a cell is a distance band and whether the pedestrian is hidden. The coverage is the share of the street's pedestrians whose cell the program draws at all. To count it, every pedestrian needs a coordinate, and a project does not have the lab's.
2 · A ruler for cases, from outlines and the ground plane
What a project holds. A labelled sample: in each of n frames a person outlines every pedestrian the written rule counts (at least 15% of the silhouette showing) and every vehicle, and nothing more: no distance, no depth, no flag for hidden. The cost is the number of outlines, 1.3 a frame on average (a van in 81% of the frames, a counted pedestrian in 43%), so 100 frames are about 128 outlines; Lesson 10 prices a labelled frame. An outline is binary: a pixel belongs to the pedestrian when at least half of it does. The program's own frames have them for nothing.
Distance from the foot row. The camera is hc = 1.4 m above the road, its focal length is f = 83.1 px and the horizon lies on row V0 = 9, rows counting down. A road point at depth z (metres along the axis) lies hc/z below the horizon in slope, so it falls on row
v = V0 + f·hc / z ⇔ z = f·hc / (v − V0)
with f·hc = 116.3 px·m. The lowest row of a pedestrian's outline, the foot row, is where the feet meet the road. A foot in row j has v between j + ¼ and j + 1¼ (a pixel is outlined once the foot passes its first sample row: the renderer takes two per pixel, at ¼ and ¾), so the reading takes v = j + ¾: row 16 reads 16.75, 15.0 m. The uncertainty is half a row, Δv = ½. Differentiating the law, |dz| = z²·dv/(f·hc), so
Δz / z = z·Δv / (f·hc) = 3.4% at 8 m, 6.9% at 16 m, 10.3% at 24 m
The ruler loses resolution in proportion to distance. A brute-force check (a pedestrian placed every 0.01 m and drawn by the renderer) reproduces these intervals to within 0.15 m:
| Foot row | Reading z | The foot is really between | Half-width, % of the reading |
|---|---|---|---|
| 20 | 9.9 m | 9.5 and 10.3 m | 4.3% |
| 18 | 11.9 m | 11.4 and 12.6 m | 5.1% |
| 16 | 15.0 m | 14.1 and 16.0 m | 6.5% |
| 14 | 20.2 m | 18.6 and 22.2 m | 8.8% |
Against the true position of the pedestrian's axis a reading is off by 0.3, 0.5 and 0.9 m (standard deviation) in the first three bands of §3, and 0.2 m short at close range, because the foot row sees the front of the legs. Another outlining convention shifts every reading by up to half a row, 2.5 m at 24 m: one reason Lesson 7 treats conventions as part of the label. The same ruler on the program's own frames, where the truth is free, measures the offset; §6 uses it.
Hidden or not. Only a van can hide a pedestrian in this street. One who steps out from behind it shows a strip of the pedestrian beside the van's silhouette, and the van is nearer. Both facts are in the outlines: the pedestrian's touches the van's, and the van's reaches lower in the picture than the pedestrian's feet, since nearer things stand lower. That test needs no distance. On the exam pedestrians of the test set it finds 228 of the 228 who step out and calls 0 of the others. How much of the silhouette shows comes from area, not width: at 16 m a whole pedestrian is 2.6 px wide, an outline is 2, 3 or 4 columns, and width says almost nothing, while area adds up the pixels. The outline's area against the area a whole pedestrian has at that distance, A0·(f/z)², with A0 = 0.69 m² the median of area·(z/f)² over the open pedestrians of the sample, reads 1.01 (spread 0.15: bodies differ in size) for open pedestrians and 0.80 for step-outs, and follows the true visible fraction of the step-outs with correlation 0.83.
3 · Counting: the street's cases and the program's
A case needs two coordinates. The first is the distance band, with its edges where the foot row changes: 12.58, 16.05 and 22.2 m. The 16 m edge falls on a row boundary, so the program's own limit is read exactly: its pedestrians, between 8 and 16 m, read on row 16 or nearer the bottom of the picture and never in the bands beyond. The second is whether the pedestrian steps out. Four bands and two classes make eight cells. Run the ruler over a pool of 40,000 street frames, 17,260 counted pedestrians, which no project has labelled but the lab can draw, and over 12,000 of the program's frames:
| Distance band | Open | Stepping out | ||
|---|---|---|---|---|
| street | program | street | program | |
| 8 to 12.6 m | 24.4% | 60.5% | 2.2% | 0.0% |
| 12.6 to 16 m | 17.6% | 39.5% | 4.4% | 0.0% |
| 16 to 22 m | 31.5% | 0.0% | 7.9% | 0.0% |
| 22 m and beyond | 9.2% | 0.0% | 2.8% | 0.0% |
Call a cell drawn when the program puts at least 1% of its pedestrians there; every cell is either far above the floor or exactly empty. The coverage is the street's share in the drawn cells, 42.0%. Of the street's pedestrians 51.3% stand beyond the 16 m the program draws and 17.3% step out. The street's own record, which only the lab has, puts 40.8% in the open within 16 m and the same 17.3% stepping out: the ruler reads the coverage 1.2 points high, because pedestrians between 16.0 and 16.25 m put their feet in the last row the program uses and cannot be told from its farthest. That is the ruler's resolution, not a bias to correct. No project can count 40,000 frames. What do a few hundred give?
4 · How many labelled frames does the count need?
A cell holding a share s of the street's pedestrians appears in a frame with probability p = s·0.43, a frame holding 0.43 counted pedestrians on average. It is absent from n independent frames with probability (1 − p)n ≈ e−np. To see it at least once with 95% confidence, (1 − p)n = 0.05, so n = ln 0.05 / ln(1 − p), and since np ≈ ln 20 = 3.0, about 3/p frames: the rule of three. Frames of one drive are correlated, so this is a best case. The commonest cell (16 to 22 m, open: 31.5% of the pedestrians, 13.6% of the frames) needs 21 frames. The rarest of the eight (the nearest band, stepping out: 2.2% of the pedestrians, one frame in 105) needs 312; 100 labelled frames lack it with probability 38%, 400 with 2%.
The coverage is a proportion of the m = 0.43n pedestrians a sample holds, so its standard deviation is √(c(1−c)/m). Drawing n frames from the pool over and over gives the same:
| Labelled frames n | 25 | 100 | 400 | 800 |
|---|---|---|---|---|
| Standard deviation of the coverage, from the formula | 15.0 points | 7.5 | 3.8 | 2.7 |
| Standard deviation over repeated draws of the pool | 15.4 | 7.6 | 3.5 | 2.5 |
The distance range the repair needs is set by the two extreme pedestrians, the least stable statistic of a sample. With 50 frames the farthest open pedestrian reads 20.3 m or less, a row short of the 24.6 m of the full sample, in 16% of the draws; with 25 frames the nearest reads 9.2 m or more in half of them, not 8; from 200 frames on the range is the same every time.
What to try. The widget opens at 100 labelled frames with the naive scene. The sample holds 46 counted pedestrians and the table reads a coverage of 30% (whiskers 22 to 37), 57% beyond 16 m and 26% stepping out, against 42.0, 51.3 and 17.3 in the pool. The whiskers hold the pool's coverage in 79% of 100-frame samples, and this one is among the rest. Slide n to 800: 39%, whiskers 36 to 43; down to 25: 21%, whiskers 7 to 36. It also holds 0 nearest-band step-outs where 1.0 is expected, which the red curve says happens 38% of the time. Pick frame 3: the foot is on row 14, so z = 83.1 × 1.4 / 5.75 = 20.2 m, between 18.6 and 22.2; tick the lab truth and the axis is at 20.3 m. Frame 5 is a step-out: the pedestrian's outline touches the van's and the visible fraction reads 0.76. Switch the program's scene: the coverage becomes 74% with the distance range, 57% with the step-outs, 100% with both and 100% with the street's scene. The 100% with both is the sample's: the one cell the program still lacks (the farthest step-outs, 0.5% of its pedestrians against 2.8% of the street's) holds none of these 100 frames' pedestrians, and the pool says 97%.
5 · The two halves of the scene
The street's scene differs from the program's in five settings: the distance range of the open pedestrians (8 to 24 m against 8 to 16), the share who step out (28% against none), the pedestrian's size, the van (anywhere from 5 to 18 m, sometimes absent) and the clutter of poles and trees (twice as much, a tenth of the poles red). Adding the street's value of one setting at a time to abba prices them inside the stage, as Lesson 2 priced the stages, and settles what §1 left open: whether a group improves because its own data arrives. Each row is three seeds at 1,600 frames; the last column is the label-free meter of Lesson 2, the alarm rate on 1,000 unlabelled street frames at the threshold that makes a tenth of the program's own pedestrian-free frames alarm.
Added to abba | Miss rate | Gain, points | Open beyond 16 m missed | Step-outs missed | Alarm rate, unlabelled |
|---|---|---|---|---|---|
| nothing | 61.6% | — | 80.9% | 87.4% | 25.5% |
| the distance range, 8 to 24 m | 56.6% | 5.0 | 74.5% | 86.5% | 25.6% |
| the step-outs, 28% | 59.9% | 1.7 | 79.6% | 85.4% | 23.3% |
| both | 52.2% | 9.4 | 69.6% | 82.9% | 25.3% |
| the clutter | 60.2% | 1.4 | 80.5% | 87.3% | 14.7% |
all five (bbba) | 46.8% | 14.8 | 63.4% | 80.4% | 10.7% |
Read the groups first: the distance range moves the far open pedestrians and hardly the step-outs, the step-outs move the step-outs and not the far ones. A group improves when its own data arrives, which §1 could not show. The van and the pedestrian's size, left out of the table, move the exam by −0.1 and 0.4 points. Then the last column. The distance range, the step-outs and the size leave the meter within three points of its 25.5%; the clutter and the van each take it to 14.7%. An unlabelled frame holds a pedestrian one time in fifty, so the program's pedestrians cannot show in it. The exam sees the pedestrian cases, the meter sees the background, and a repair guided by one is blind to half the stage.
The settings compound once more. Distance and step-outs are worth 5.0 and 1.7 points alone and 9.4 together, 2.7 more than their sum. The clutter is worth 1.4 on the program and 4.9 on the program whose pedestrians are right (52.2% to 47.3%), and the van adds nothing once the clutter is there (−0.5). A pedestrian in more places gives the detector more places to mistake a pole for the pedestrian, and the clutter is where it learns not to. A repair judged where it was cheapest to try, on the program as it stood, would have been dropped.
How far should the background be pushed? With the pedestrian cases repaired, clutter at 0.5, 0.75, 1.0, 1.5 and 2.0 poles and trees a frame (the street has 1.0; one seed each) gives alarm rates of 25.7, 18.0, 16.4, 13.1 and 9.0% and miss rates of 54.1, 49.9, 49.1, 48.3 and 46.5%. The meter keeps falling past the street's own value and crosses its nominal 10% near twice the street's clutter, because the van, wrong as well, moves it as far as the clutter does and the sweep charges that gap to the one knob it was given. It says that the background is incomplete and by how much, not in which setting: a repair with no stopping rule, which §6 leaves alone.
6 · Repair from the evidence, and price it
The labelled sample sets two settings. The distance range is the nearest and the farthest open pedestrian the ruler reads, each corrected by the ruler's offset, −0.09 m, measured on the program's own frames run with a deliberately wide range (8 to 26 m), where the truth is free. The step-out probability needs care. The sample shows the share of step-outs among the pedestrians counted, ŝ = 17.3%, and the program must show the same share through the same ruler. Counting favours the open pedestrian (a strip has to reach 6 px²), so the odds of a counted step-out are a fraction κ of the odds of drawing one, and κ is measured on the program's own frames: 0.56. Matching odds(ŝ) = κ·e/(1 − e) gives e = odds(ŝ)/(odds(ŝ) + κ) = 0.27, within a point of the street's 0.28, which the sample never saw. The van, the clutter and the size stay the program's. Each training seed gets its own labelled sample, so the evidence's variation is in the spread.
| Program, 3 seeds | Nearest and farthest open pedestrian set, m | Step-out probability | Miss rate | Gap to the exact pedestrian settings, points |
|---|---|---|---|---|
abba | 8 to 16 | 0 | 61.6% | — |
| ruler, 50 labelled frames | 9.0 to 24.6 | 0.31 | 49.6% | −2.6 |
| ruler, 200 labelled frames | 8.2 to 24.6 | 0.23 | 51.5% | −0.7 |
| ruler, 800 labelled frames | 8.0 to 24.6 | 0.26 | 51.5% | −0.7 |
| exact pedestrian settings (lab) | 8 to 24 | 0.28 | 52.2% | 0 |
| exact pedestrian settings and the clutter (lab) | 8 to 24 | 0.28 | 47.3% | — |
the street's scene, bbba (lab) | 8 to 24 | 0.28 | 46.8% | — |
The ruler's two numbers cost nothing the exam can see: the programs set from 50, 200 and 800 labelled frames miss 49.6%, 51.5% and 51.5%, against 52.2% for the exact settings, inside the spread of the exact program's own three seeds (5.1 points). The estimates are noisy at small n, the step-out probability ranging from 0.25 to 0.39 across the three 50-frame samples, and the exam cannot see it. The count can: a 50-frame sample stops short at 20.3 m or less in 16% of draws (§4), and its program would never draw the farthest 4 m; the three used here all reached 24.6. What the ruler cannot supply is the rest of the way to the exact stage, 52.2% to 47.3% with the clutter, which §5 gave to the meter.
7 · A kind is not a frequency
A cell drawn is the least a program owes it. A cell is a box several metres deep, and the case the shuttle has least time for sits at the near end of one. Take the pedestrian who steps out from behind a van and whose axis is nearer than 12 m. The street is a program, so the lab can count this without rendering a frame: an exact integral of its scene law gives 0.822% of the frames, a frame in 122, and 1,000,000 scene draws agree, 8,136 of them holding one. A training set of 1,600 frames should hold 13.2 of them, and the three training sets of the street-scene program hold 11, 13 and 16. The table's nearest-band cell (a reading under 12.6 m, counted by the exam) is a little wider: one frame in 105.
To grade a detector on the case itself the lab draws it on purpose: scenes of the street's program with a step-out nearer than 12 m, kept by rejection at the scene-draw level (a scene costs microseconds, one in 17 is kept, and only the kept ones are rendered). A project cannot do this. It gives 1,200 frames, 970 with a pedestrian the exam counts. The detector trained with the street's scene misses 56.6% of them and 9.8% of the pedestrians who stand in the open nearer than 12 m (1,200 frames drawn the same way); the program-trained detector misses 76.5% and 17.7%. The program has been told about the case and the detector has seen about thirteen of it: the cell is covered and the case is not learned. Waiting does not help. To hold 100 of them a program that draws as the street does needs 12.2 thousand frames, and almost all of those frames show pedestrians the detector already finds.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
The ruler turns "which cases does the scene never draw?" into a table of eight cells and a coverage of 42.0%; a few hundred labelled frames read it to 3.5 points (400 frames), and a program whose pedestrians are set from 200 of them misses 51.5%, level with the 52.2% of the exact settings. That tells a program which cells to fill, not how often. The case the shuttle has least time for, a step-out nearer than 12 m, occurs in one frame in 122; a training set of 1,600 frames holds about 13 of them; and the detector trained with the street's scene misses 57% of them, against 10% of the pedestrians who stand as near in the open. A program that draws as the street does needs 12.2 thousand frames to hold a hundred of them, and nothing in the street lets it ask for fewer. The program can ask: it can draw the case on purpose. How should a generator spend its renders on cases that rarely happen, and what does that do to the statistics the model learns?
Interview prompts
- How would you find out which situations a synthetic data generator never produces? (§2, §3 — outline a few hundred real frames, give each pedestrian a coordinate from the outlines and the geometry, run the same ruler on the generator's frames, and compare the share of each cell.)
- How far away is a pedestrian in a monocular frame, and how accurate is the reading? (§2 — z = f·hc/(v − V0) from the foot row on the ground plane; a half-row error grows linearly with distance.)
- Why can't a better camera model recover a case the scene never draws? (§1 — repairs downstream change the pixels of the cases that exist; the step-out miss stayed put while the open pedestrians' fell by 22.7 points.)
- How many labelled samples make you 95% sure of seeing a case of probability p? (§4 — (1 − p)n = 0.05 gives n = ln 0.05/ln(1 − p) ≈ 3/p, a best case: frames of one drive are correlated.)
- Why can a label-free domain-gap meter miss the most important gap? (§5 — pedestrians are rare in unlabelled frames, so the alarm rate saw the clutter and the van and stayed put when the pedestrian settings were repaired.)
- Two repairs each help a little and help more together. Why? (§5 — a repair pays most where the others have made it matter: the clutter, worth 1.4 alone, is worth 4.9 once the pedestrians are right.)
- The generator draws every kind of case and the model still fails on one. What do you check? (§7 — how many examples of it the training set holds: the nearest step-outs are one frame in 122, about 13 in 1,600 frames.)
Companion reads: Computer Vision · 04 Cameras and projection geometry (the ground-plane law the ruler uses), and 3D Vision, from first principles (distance from a single image).