What the label means
Lesson 6 ended on a number that depends on a definition: the same detector misses 63%, 58% or 44% of the rare case depending on which pedestrians count, and Lesson 2 had priced the label stage at nothing. Both are true. A label enters a project twice, as the target a detector is taught and as the rule that decides whom a miss rate is averaged over, and different clauses move the two: how much of a pedestrian must show moves the reported number by up to 25 points and the detector by about one, while where the mask lies and what it includes move the detector and no frame-level number. A label is a definition with a clause for each question a mask leaves open. This lesson measures each clause on both doors and writes them down as a contract. It cannot say whether a renderer we did not write keeps it.
New idea: a label is a definition, written as clauses, and the clauses that move a reported number are not the ones that move a trained detector. The visibility rule moves the rare case's miss rate by 19.2 points when it scores the detector and by at most 2.4 when it teaches it; a mask half a pixel off costs the detector 2.3 points on the exam and moves no frame-level number.
Forces next: The labels are definitions, and the places where two definitions disagree are now written down and measured: visible mask or whole silhouette, what counts as a pedestrian at all (the same detector scores 44.7% or 51.8% on the exam and 40.1% or 65.1% on the rare case, depending on the rule), depth along the axis or range along the ray (17% of the close step-outs by depth are farther than 12 m by range), a mask half a pixel off. The clauses that move a reported number are not the ones that move a detector: the visibility rule moves what is reported by up to 25 points and what is learned by about one, while a mask half a pixel off costs 2.3 points of training and the hidden part of a silhouette 1.6. All of it was settled in a renderer we wrote ourselves, where we chose every convention. A production pipeline uses a production renderer whose conventions we did not choose, and its defaults decide, silently, how a sliver is counted. How do we make a renderer we did not write obey the same contract?
1 · One detector, seven rules
Take the detector that the exact-stage program bbba trains (the street's scene, light and camera, the program's own label), three seeds, and the exam of Lesson 1: 3,000 real test frames, a threshold that lets a tenth of the pedestrian-free validation frames alarm. 1,517 of the frames show a pedestrian, and the exam counts 1,323 of them. A benchmark could have written another sentence about who counts. Each sentence below is a rule that scores the same scores differently, and none touches a weight; the rare case is Lesson 6's slice of 3,000 frames of pedestrians who step out from behind the van closer than 12 m, which a project cannot buy (a lab privilege).
| Rule: a pedestrian counts when | Counted of 1,517 | Miss, the exam's frames | Miss, the rare case |
|---|---|---|---|
| at least 1 px² of the silhouette shows (the program's own label) | 1,460 | 50.7% | 63.4% |
| at least 6 px² show | 1,323 | 46.8% | 59.2% |
| at least 15% of the silhouette shows (the street's label) | 1,438 | 50.0% | 58.9% |
| both of the last two (the exam) | 1,323 | 46.8% | 58.4% |
| at least half of the silhouette shows | 1,327 | 46.7% | 44.2% |
| at most 35% of the silhouette is hidden (the visibility clause of Caltech's "reasonable" subset, Dollár et al., 2012) | 1,263 | 44.7% | 40.1% |
| the whole silhouette, hidden part included, is at least 6 px² (an amodal count) | 1,508 | 51.8% | 65.1% |
The reported miss rate spans 7.1 points over the exam's frames and 25.0 over the rare case (the three rules of Lesson 6, any pixel, the exam's and half, span 19.2). A rule is a cut on the curve of Lesson 6's last section, the miss rate as a function of the visible fraction, and the number it reports is the average of the curve above the cut. In the rare case 49% of the pedestrians show less than half of themselves, against 13% of the exam's, so a cut there moves the average far more.
Lesson 2 priced the label stage at −0.1 points: aaab reproduced aaaa to the last digit, and the street's rule in place of the program's moved bbba to bbbb by 0.3 points. A stage that moves a reported number by 25 points cannot be worth nothing. Either the price was measured on the wrong thing or the label does two jobs, and only one of them was in the price.
2 · Two doors
A definition enters a project in two places. As a scoring rule it chooses whom a miss rate is averaged over; §1 changed that and nothing else. As a training label it chooses what the detector is taught, and that is the only door Lesson 2 swapped. The head is taught per cell of a stride-2 grid: a cell is positive when the pixel at its centre is at least half pedestrian and the frame counts as holding one under the rule, lab = c >= 0.5 && y. The rule enters through y alone, so it can change a cell only in a frame whose pedestrian is partly hidden. Count such frames in the program's scene and in the street's, for the rules of the program (any pixel), the street (15% of the silhouette) and the exam:
| Scene drawn by | Frames | With a pedestrian | Label differs: program's rule against the street's | against the exam's |
|---|---|---|---|---|
| the program | 40,000 | 19,974 | 0 | 0 |
| the street | 40,000 | 20,048 | 428 (2.1%) | 1,948 (9.7%) |
The program's scene never hides anyone, so its rule is never asked: that is why aaab equals aaaa to the last digit. Even in the street's scene the training rule decides about one pedestrian frame in fifty, and a frame decides only its covered cells: of the 5,579 positive cells in the 819 pedestrian frames of the seed-1 training set, the street's rule changes 23 (0.4%), a rule asking for half the silhouette 176 (3.2%). The scoring rule touches no cell; it changes the population. The exam counts 1,323 pedestrians, any pixel 1,460, at most 35% hidden 1,263, and the average runs over a curve that falls from 92% missed when under a quarter of the silhouette shows to 36% when over four fifths do (seed 1, the rare case).
Teach under three rules and score under three (the street's rule is the cell bbbb; the 50% detector is new; three seeds each, the series' protocol):
| Taught with | Miss, the exam's frames, scored by | Miss, the rare case, scored by | ||||
|---|---|---|---|---|---|---|
| any pixel | the exam's rule | half | any pixel | the exam's rule | half | |
| any pixel (the program) | 50.7 | 46.8 | 46.7 | 63.4 | 58.4 | 44.2 |
| 15% of the silhouette (the street) | 51.0 | 47.1 | 47.0 | 63.7 | 58.7 | 44.5 |
| half of the silhouette | 51.3 | 47.5 | 47.4 | 65.0 | 60.2 | 46.6 |
Down a column the detector changes and the rule stays: the numbers differ by at most 0.7 points over the exam's frames and 2.4 over the rare case. Along a row the detector stays and the rule changes: 4.0 and 19.2 points, 6 and 8 times as far. Lesson 2's price was the price of one door. The other is paid every time two numbers are compared.
3 · What a mask leaves open: six clauses
The label stage turns what a renderer returns, coverage, a centre-ray depth and range, an id, into a frame label and cell labels. A mask of a pedestrian answers "which pixels" and leaves six questions on which two honest implementations disagree. Each is a clause of the definition.
| Clause | The question | The lab's setting | A setting in use elsewhere |
|---|---|---|---|
| 1 · which pixels | only the visible ones, or all of the silhouette? | modal: the visible mask; the silhouette is rendered alone to say how much is hidden | amodal masks, with a depth order (Zhu et al., 2017; on KITTI images, Qi et al., 2019); Caltech matches the full box and keeps the visible box beside it |
| 2 · how much counts | how much must show, for the label and for the miss rate? | any pixel to teach; 15% of the silhouette in the street's label; 15% and 6 px² in the exam | "reasonable": at most 35% occluded and at least 50 px tall; KITTI: three levels by height, occlusion and truncation |
| 3 · how far | distance along the axis, or along the ray? | depth, which the ground-plane law z = f·hc/(v − V0) turns into a row | range, which a LiDAR measures; in Blender the Depth pass is along the axis, the Mist pass Euclidean |
| 4 · aligned how | is pixel (0, 0) a corner or a centre? | centres at +0.5, the graphics convention | calibration code in the OpenCV tradition: centres at integers, ucv = u − 0.5; boxes: COCO [x, y, width, height], KITTI [left, top, right, bottom] |
| 5 · counted how finely | how many rays decide a pixel, and what is summed? | 2 × 2 rays; area is the sum of coverage | one ray through the centre (an object-index pass), 8 × 8 rays, alpha after a pixel filter |
| 6 · which instant | which moment does the label belong to? | one: picture and mask are the same moment | a label timestamp and an exposure window are two clocks |
How far. A pixel with tan φ across and ρ up looks along (tan φ, ρ, 1), so a surface at depth z sits at z·(tan φ, ρ, 1) and its range is
r = z·√(1 + tan²φ + ρ²)
In the ground-plane law of the table f = 83.1 px, hc = 1.4 m and V0 = 9 are Lesson 1's, and v counts rows down. The factor is 1.155 on the horizon at the edge of the 60° field and 1.165 at the corner pixel; pedestrians stand near the axis, and among the exam's the largest is 1.06. Small as that is, it moves pedestrians across the 12 m line of Lesson 6's case R: by depth 0.822% of the street's frames are close step-outs, by range 0.683%, 17% fewer, and every weight p/q of that lesson inherits the difference.
Aligned how. The same point has coordinate u where pixel i spans [i, i + 1) and u − 0.5 where centres are integers; a mask projected with one convention onto a grid that assumes the other is displaced by half a pixel in both directions, and a pedestrian at 16 m is about 2.5 pixels wide. Counted how finely. Coverage is an area integral. With 2 × 2 rays a pedestrian narrower than the ray spacing slips between them: 57 pedestrians of the test set have zero visible area at 2 × 2 rays and 14 at 8 × 8, and "at least 6 px²" gives another answer for 33 of the 310 whose area lies between 3 and 10 px², a threshold effect that matters only near the line. Which instant. A pedestrian at 1.5 m/s, 12 m away, in a 20 ms exposure moves f·v·Δt/z = 0.21 px; the lab has one instant, but a pipeline that stamps a label at one time and exposes the picture over another has two, and the offset between them is the motion during the exposure.
4 · Each clause, both doors
Measure every clause twice. Taught: retrain the detector of bbba on the same 1,600 frames and seeds with only the label changed. Scored: keep the detector and change the clause in the rule. The dial does the second live, over the stored decisions of nine detectors, one per label.
What to try. As the page opens, the program's detector on the rare case under the exam's rule: 2,404 pedestrians counted, 58.1% missed (seed 1; the tables' three-seed mean is 58.4%). Cut to 0 with the 6 px² box unticked: any pixel, 2,824 counted, 63.2% missed; cut at 50%: 1,518 counted, 43.7%. Switch to the exam's frames and repeat: 50.5% to 46.6%, because most pedestrians clear the cut. Back on the rare case, step through the nine detectors at the exam's rule: best to worst is 9.6 points, where the cut alone moves the program's detector by 19.5. Read the area with one ray per pixel, what an object-index pass delivers: 2,376 counted, 57.8% missed; after the default filter: 2,043 counted, 52.1%, six points from a reader with the detector unchanged. Choose range: of the 2,398 counted, all are closer than 12 m by depth and 2,022 by range. Tick half a pixel off on the exam's frames: violet marks where the displaced mask lands, and its overlap with the true mask is 57%.
The other column comes from nine retrained detectors. Six seeds, the differences paired by seed (the same 1,600 frames, only the label differs), because several differences are of the size of the seed wobble:
| Clause: the label taught | Cells changed (seed 1) | Taught: exam's miss against the program's label | Scored: the same detector, the clause changed in the rule |
|---|---|---|---|
| 1 · whole silhouette (amodal) | 564 added (10.1%), all hidden | +1.6 ± 0.2; rare case −2.8; false alarms within 2 px of the van 38% → 47% | count the silhouette: 1,508 counted, 51.8% missed against 46.8 |
| 2 · 15% / 50% of the silhouette | 23 / 176 | +0.2 ± 0.1 / +0.1 ± 0.5 | any pixel 50.7, half 46.7, 35% hidden 44.7; the rare case up to 25.0 apart |
| 3 · range for depth | none: the detector never sees a distance | — | 3.9% of the exam's pedestrians change distance bin (12, 16, 20 m); closer than 12 m: 289 by depth, 266 by range |
| 4 · mask half a pixel / one pixel off | 2,200 / 4,395 (39% / 79%) | +2.3 ± 0.4 / +5.8 ± 0.5 | overlap with the true mask 57% / 32%, below 50% for 25% / 89% of pedestrians; the foot row moves for 51%, the distance read from it by 8% |
| 5 · 8×8 rays / one ray / default filter | 908 / 1,096 / 960 | −0.4 ± 0.4 / −0.1 ± 0.3 / −1.8 ± 0.5 | the exam's decision flips for 2.5% / 5.2% / 4.1% of its pedestrians (§6) |
| 6 · which instant | not measured | ||
Three things stand out. Reporting effects are large and cost nothing: no detector was retrained to get the right-hand column. The clauses that move the detector are not the ones that move the number. The visibility rule, argued about most, decides 0.4 to 3.2% of the positive cells and moves the detector by less than its wobble; so do 8×8 rays and one ray, which change 16 to 20%. A mask half a pixel off moves 39% of the cells and costs 2.3 points in all six seeds, though no frame-level number of the exam depends on where a mask lies; the whole silhouette adds cells that show only van and costs 1.6. The sign of a clause is not known in advance. The softened mask of the default filter removes 18% of the positive cells and adds none: an edge pixel that is exactly half covered, which the lab's rule counts, reads just under one half once blurred, and thin parts are mostly edge, so 32% of the head's cells go, 24% of the feet's, and 75% of all that go lie on the outer columns. The detector taught on it is −1.8 points better on the exam, in all six seeds.
And a rule never reorders detectors. Across the 16 programs of Lesson 2 (seed 1, exam miss from 46.7 to 80.5%) and the 11 designs of Lesson 6, the order under each of the seven rules is the order under the exam's (rank correlation at least 0.998), and any pair that swaps places is within 0.6 points under one of the two rules. What reorders two detectors is the population. The amodal-taught detector misses −2.8 points fewer of the rare case than the program's and 1.6 more of the exam's pedestrians, in all six seeds; the default filter's does the opposite, −1.8 better on the exam and 1.2 worse on the rare case. Whose misses are averaged is part of the definition too.
5 · The contract
A clause that is a habit of one author's code is invisible until two pipelines meet. The evidence for the label stage is what the benchmark wrote down, an afternoon of reading, and what a renderer returns for scenes whose answers are known in closed form: a few scenes, no labels and no street. So write each clause as a setting and a test.
| Clause | The lab's setting | Test | A violation looks like |
|---|---|---|---|
| 1 · which pixels | modal labels, silhouette rendered alone | the visible mask lies inside the silhouette | false alarms on the van (47% against 38%) |
| 2 · how much counts | 15% and 6 px² for the exam | a sliver of known size: the decision equals the closed form | one detector at 44.7% or 51.8% |
| 3 · how far | depth along the axis | on the ground, depth equals f·hc/(v + 0.5 − V0) at every pixel | a range pass is off by up to 0.165 |
| 4 · aligned how | centres at +0.5 | a sliver at a known sub-pixel place: its mask's centroid is there | overlap 57%, +2.3 points |
| 5 · counted how finely | 2 × 2 rays, sum of coverage | a 6.0 px² sliver at sixteen sub-pixel offsets: the counted area stays within a stated error | worst error 4.0 px² with one ray, 1.5 with 2 × 2, 0.25 with 8 × 8 |
| 6 · which instant | one instant | a target at a known speed: the label is where the target was at the exposure's midpoint | an offset of f·v·Δt/z pixels |
The listing under the dial holds the counting rules and two of these tests. The coverage test places a 1.2 × 5 pixel sliver, whose area is exactly 6.0 px², at sixteen sub-pixel offsets and reports the worst error of the counted area; the distance test reads a frame's depth pass on the ground and reports its largest deviation from the ground-plane law. The lab passes it to 0.0000, and a range map in its place fails it by 0.165. At two rays a 6 px² sliver is counted within 1.5 px², 25% of its area, so "at least 6 px²" is a coin toss for it: the rule is applied to a number only that accurate, and the contract says so.
6 · A renderer we did not write
All six clauses were settled in street.js by one author, who chose every convention. A production pipeline settles them with a renderer's defaults. In Blender 5.2 LTS, whose defaults were checked by running it, the Depth pass is the distance along the camera axis and the Mist pass the Euclidean distance; the Object Index pass is point-sampled at the pixel centre, a hard label with no coverage, where Cryptomatte stores coverage-weighted ids; and the default pixel filter, Blackman-Harris 1.5 px wide, lets a step edge that lies on a pixel boundary leak 12% into the next pixel, where the box filter gives exact area coverage. Each is a setting of a clause above, chosen by someone else.
The filter is the worked example. A Gaussian blur of standard deviation σ, read at a pixel centre half a pixel from the edge, leaks Φ(−0.5/σ) (Φ is the standard normal distribution function), so a leak of 12% means σ = 0.43 px. The lab's own pixel is a box of standard deviation 0.29 px; variances add, so the default filter's extra blur is 0.31 px. Read the exam's pedestrians' area through each of these, same detector (seed 1), the exam's rule:
| Area is read as | Counted | Decisions that differ from the exam's | Miss, exam's frames | Miss, rare case | Taught on that mask: exam's miss against the program's |
|---|---|---|---|---|---|
| the lab's: sum of 2 × 2 coverage | 1,323 | 0 | 46.7% | 58.1% | 0 |
| sum of 8 × 8 coverage | 1,320 | 33 (2.5%) | 46.4% | 58.0% | −0.4 |
| one ray per pixel centre (an index pass) | 1,358 | 69 (5.2%) | 47.8% | 57.8% | −0.1 |
| pixels at least 0.5 after the default filter | 1,269 | 54 (4.1%) | 44.8% | 52.1% | −1.8 |
None of these is wrong. Each is a different definition chosen by a default, and each changes a number a team will report, 54 pedestrians of the exam's frames and six points of the rare case for the filter alone, or a detector it will train, in a direction nobody chose.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
The clauses of a label are written down and measured. The visibility rule moves the number a team reports by up to 25 points and the detector by about one: the same detector scores 44.7% or 51.8% on the exam depending on the rule. What moves a trained detector is a clause nobody argues about, a mask half a pixel off, the hidden part of a silhouette, a softened edge, by 2.3 and 1.6 points worse and −1.8 better. All of it was settled in a renderer we wrote ourselves, where we chose every convention, and a production renderer's defaults decide, silently, how a sliver is counted: an exam's pedestrians change status in 5.2% of cases when their area is read from an index pass and in 4.1% through the default filter. A production pipeline uses a production renderer whose conventions we did not choose. How do we make a renderer we did not write obey the same contract?
Interview prompts
- Two papers report miss rates for the same detector five points apart. What do you check first? (§1 — the counting rule: visible fraction, minimum area, whether hidden pedestrians count; seven rules moved one detector's scores 7.1 points.)
- A synthetic label stage changed and no model moved. Is the stage irrelevant? (§2 — it acts through the training door only where partly hidden pedestrians exist; the scoring door was not tested, and it moves what is reported.)
- Would you teach an amodal mask to a detector scored on visible pedestrians? (§4 — it contradicts the pixels: false alarms move to the van and the exam loses 1.6 points, though the rare case gains.)
- What is the difference between depth and range, and when does it matter? (§3 — along the axis against along the ray, r = z√(1 + tan²φ + ρ²); it matters at any distance gate, case table or weight: 17% of close step-outs.)
- How do you test a mask-building convention without a reference renderer? (§5 — scenes with closed-form answers: a sliver of known area at sub-pixel offsets, a ground plane for the depth law, a target at a known place.)
- Why does a coverage threshold at 6 px² need a tolerance? (§5 — at 2 × 2 rays the counted area of a 6 px² sliver errs by up to 1.5 px²; only near the line do decisions flip.)
Companion reads: Computer Graphics · 06 Sampling and antialiasing (coverage, rays per pixel and pixel filters), Computer Vision · 09 Object detection (overlap and the matching rule) and Computer Vision · 04 Cameras and projection geometry (pixel centres and the pinhole).