all_lessons/Synthetic Vision Data/07 · What the label meanslesson 7 / 12

What the label means

Lesson 6 ended on a number that depends on a definition: the same detector misses 63%, 58% or 44% of the rare case depending on which pedestrians count, and Lesson 2 had priced the label stage at nothing. Both are true. A label enters a project twice, as the target a detector is taught and as the rule that decides whom a miss rate is averaged over, and different clauses move the two: how much of a pedestrian must show moves the reported number by up to 25 points and the detector by about one, while where the mask lies and what it includes move the detector and no frame-level number. A label is a definition with a clause for each question a mask leaves open. This lesson measures each clause on both doors and writes them down as a contract. It cannot say whether a renderer we did not write keeps it.

The thesis, here
A label is a definition, not a measurement: a function of what the renderer returns, with a clause for each question that a mask leaves open. A definition acts through two doors, the target a detector is taught and the rule that selects whom a number is averaged over, and the clauses that move one door are not the ones that move the other: how much of a pedestrian must show moves what is reported and hardly the detector, where the mask lies moves the detector and no frame-level number. Two reported numbers are comparable only if their definitions agree clause by clause.
Linear position
Forced by: Drawing the rare case on purpose, 12 times as often as it occurs, gives the detector 160 examples of it instead of 13 for the same 1,600 renders, and what it makes of them depends on the weight each carries: as drawn they say the case is common, and it misses 8 points fewer of the rare case and 1 point more of everything else; weighted by p/q they restore the street, and the detector is the natural one at an effective sample size of 1,463 frames. The pictures, the cases and their frequencies are now right, and the training data and the exam agree about what is in a frame. They may still disagree about what counts: a pedestrian with two pixels showing is a pedestrian to the program and not to an exam that asks for a visible fraction, and the miss rate on the rare case moves by 19 points with the rule. What does a label mean, and where can the program's definition differ from the exam's?
New idea: a label is a definition, written as clauses, and the clauses that move a reported number are not the ones that move a trained detector. The visibility rule moves the rare case's miss rate by 19.2 points when it scores the detector and by at most 2.4 when it teaches it; a mask half a pixel off costs the detector 2.3 points on the exam and moves no frame-level number.
Forces next: The labels are definitions, and the places where two definitions disagree are now written down and measured: visible mask or whole silhouette, what counts as a pedestrian at all (the same detector scores 44.7% or 51.8% on the exam and 40.1% or 65.1% on the rare case, depending on the rule), depth along the axis or range along the ray (17% of the close step-outs by depth are farther than 12 m by range), a mask half a pixel off. The clauses that move a reported number are not the ones that move a detector: the visibility rule moves what is reported by up to 25 points and what is learned by about one, while a mask half a pixel off costs 2.3 points of training and the hidden part of a silhouette 1.6. All of it was settled in a renderer we wrote ourselves, where we chose every convention. A production pipeline uses a production renderer whose conventions we did not choose, and its defaults decide, silently, how a sliver is counted. How do we make a renderer we did not write obey the same contract?
The plan
Six moves. (1) Score one detector under seven rules and watch the number move. (2) Find the two doors a definition enters by and count what each changes. (3) Derive the clauses from what a mask leaves open. (4) Measure each clause on both doors. (5) Write the clauses as a contract with a test each. (6) Check a renderer's defaults against it.

1 · One detector, seven rules

Take the detector that the exact-stage program bbba trains (the street's scene, light and camera, the program's own label), three seeds, and the exam of Lesson 1: 3,000 real test frames, a threshold that lets a tenth of the pedestrian-free validation frames alarm. 1,517 of the frames show a pedestrian, and the exam counts 1,323 of them. A benchmark could have written another sentence about who counts. Each sentence below is a rule that scores the same scores differently, and none touches a weight; the rare case is Lesson 6's slice of 3,000 frames of pedestrians who step out from behind the van closer than 12 m, which a project cannot buy (a lab privilege).

Rule: a pedestrian counts whenCounted of 1,517Miss, the exam's framesMiss, the rare case
at least 1 px² of the silhouette shows (the program's own label)1,46050.7%63.4%
at least 6 px² show1,32346.8%59.2%
at least 15% of the silhouette shows (the street's label)1,43850.0%58.9%
both of the last two (the exam)1,32346.8%58.4%
at least half of the silhouette shows1,32746.7%44.2%
at most 35% of the silhouette is hidden (the visibility clause of Caltech's "reasonable" subset, Dollár et al., 2012)1,26344.7%40.1%
the whole silhouette, hidden part included, is at least 6 px² (an amodal count)1,50851.8%65.1%

The reported miss rate spans 7.1 points over the exam's frames and 25.0 over the rare case (the three rules of Lesson 6, any pixel, the exam's and half, span 19.2). A rule is a cut on the curve of Lesson 6's last section, the miss rate as a function of the visible fraction, and the number it reports is the average of the curve above the cut. In the rare case 49% of the pedestrians show less than half of themselves, against 13% of the exam's, so a cut there moves the average far more.

Lesson 2 priced the label stage at −0.1 points: aaab reproduced aaaa to the last digit, and the street's rule in place of the program's moved bbba to bbbb by 0.3 points. A stage that moves a reported number by 25 points cannot be worth nothing. Either the price was measured on the wrong thing or the label does two jobs, and only one of them was in the price.

2 · Two doors

A definition enters a project in two places. As a scoring rule it chooses whom a miss rate is averaged over; §1 changed that and nothing else. As a training label it chooses what the detector is taught, and that is the only door Lesson 2 swapped. The head is taught per cell of a stride-2 grid: a cell is positive when the pixel at its centre is at least half pedestrian and the frame counts as holding one under the rule, lab = c >= 0.5 && y. The rule enters through y alone, so it can change a cell only in a frame whose pedestrian is partly hidden. Count such frames in the program's scene and in the street's, for the rules of the program (any pixel), the street (15% of the silhouette) and the exam:

Scene drawn byFramesWith a pedestrianLabel differs: program's rule against the street'sagainst the exam's
the program40,00019,97400
the street40,00020,048428 (2.1%)1,948 (9.7%)

The program's scene never hides anyone, so its rule is never asked: that is why aaab equals aaaa to the last digit. Even in the street's scene the training rule decides about one pedestrian frame in fifty, and a frame decides only its covered cells: of the 5,579 positive cells in the 819 pedestrian frames of the seed-1 training set, the street's rule changes 23 (0.4%), a rule asking for half the silhouette 176 (3.2%). The scoring rule touches no cell; it changes the population. The exam counts 1,323 pedestrians, any pixel 1,460, at most 35% hidden 1,263, and the average runs over a curve that falls from 92% missed when under a quarter of the silhouette shows to 36% when over four fifths do (seed 1, the rare case).

Teach under three rules and score under three (the street's rule is the cell bbbb; the 50% detector is new; three seeds each, the series' protocol):

Taught withMiss, the exam's frames, scored byMiss, the rare case, scored by
any pixelthe exam's rulehalfany pixelthe exam's rulehalf
any pixel (the program)50.746.846.763.458.444.2
15% of the silhouette (the street)51.047.147.063.758.744.5
half of the silhouette51.347.547.465.060.246.6

Down a column the detector changes and the rule stays: the numbers differ by at most 0.7 points over the exam's frames and 2.4 over the rare case. Along a row the detector stays and the rule changes: 4.0 and 19.2 points, 6 and 8 times as far. Lesson 2's price was the price of one door. The other is paid every time two numbers are compared.

Road not taken · teach under the rule you will be scored by
If the benchmark counts a pedestrian only when half of the silhouette shows, teach the detector that. The lab measures it, and it backfires: the detector taught at 50% misses 46.6% of the rare case under the half rule where the program's detector misses 44.2%, 2.4 points more, and 47.4% against 46.7% over the exam's frames. A rule that calls a pedestrian with 40% showing a non-pedestrian makes the pedestrian's cells negatives, and the detector is taught to ignore them. A scoring rule says whom to count, not whom to detect; the team that is scored by the half rule should report under it and keep teaching the program's way.

3 · What a mask leaves open: six clauses

The label stage turns what a renderer returns, coverage, a centre-ray depth and range, an id, into a frame label and cell labels. A mask of a pedestrian answers "which pixels" and leaves six questions on which two honest implementations disagree. Each is a clause of the definition.

ClauseThe questionThe lab's settingA setting in use elsewhere
1 · which pixelsonly the visible ones, or all of the silhouette?modal: the visible mask; the silhouette is rendered alone to say how much is hiddenamodal masks, with a depth order (Zhu et al., 2017; on KITTI images, Qi et al., 2019); Caltech matches the full box and keeps the visible box beside it
2 · how much countshow much must show, for the label and for the miss rate?any pixel to teach; 15% of the silhouette in the street's label; 15% and 6 px² in the exam"reasonable": at most 35% occluded and at least 50 px tall; KITTI: three levels by height, occlusion and truncation
3 · how fardistance along the axis, or along the ray?depth, which the ground-plane law z = f·hc/(v − V0) turns into a rowrange, which a LiDAR measures; in Blender the Depth pass is along the axis, the Mist pass Euclidean
4 · aligned howis pixel (0, 0) a corner or a centre?centres at +0.5, the graphics conventioncalibration code in the OpenCV tradition: centres at integers, ucv = u − 0.5; boxes: COCO [x, y, width, height], KITTI [left, top, right, bottom]
5 · counted how finelyhow many rays decide a pixel, and what is summed?2 × 2 rays; area is the sum of coverageone ray through the centre (an object-index pass), 8 × 8 rays, alpha after a pixel filter
6 · which instantwhich moment does the label belong to?one: picture and mask are the same momenta label timestamp and an exposure window are two clocks

How far. A pixel with tan φ across and ρ up looks along (tan φ, ρ, 1), so a surface at depth z sits at z·(tan φ, ρ, 1) and its range is

r = z·√(1 + tan²φ + ρ²)

In the ground-plane law of the table f = 83.1 px, hc = 1.4 m and V0 = 9 are Lesson 1's, and v counts rows down. The factor is 1.155 on the horizon at the edge of the 60° field and 1.165 at the corner pixel; pedestrians stand near the axis, and among the exam's the largest is 1.06. Small as that is, it moves pedestrians across the 12 m line of Lesson 6's case R: by depth 0.822% of the street's frames are close step-outs, by range 0.683%, 17% fewer, and every weight p/q of that lesson inherits the difference.

Aligned how. The same point has coordinate u where pixel i spans [i, i + 1) and u − 0.5 where centres are integers; a mask projected with one convention onto a grid that assumes the other is displaced by half a pixel in both directions, and a pedestrian at 16 m is about 2.5 pixels wide. Counted how finely. Coverage is an area integral. With 2 × 2 rays a pedestrian narrower than the ray spacing slips between them: 57 pedestrians of the test set have zero visible area at 2 × 2 rays and 14 at 8 × 8, and "at least 6 px²" gives another answer for 33 of the 310 whose area lies between 3 and 10 px², a threshold effect that matters only near the line. Which instant. A pedestrian at 1.5 m/s, 12 m away, in a 20 ms exposure moves f·v·Δt/z = 0.21 px; the lab has one instant, but a pipeline that stamps a label at one time and exposes the picture over another has two, and the offset between them is the motion during the exposure.

4 · Each clause, both doors

Measure every clause twice. Taught: retrain the detector of bbba on the same 1,600 frames and seeds with only the label changed. Scored: keep the detector and change the clause in the rule. The dial does the second live, over the stored decisions of nine detectors, one per label.

The definition dial
Choose whom the miss rate is averaged over and how a pedestrian's area is read, then move the cut. Curves: the reported miss rate of nine detectors (seed 1, each taught with a different label) at every cut, the chosen one bold. Strip: three pedestrians just under the cut and three just over, the visible mask in teal, the hidden part in amber. Counts and rates are computed live from stored decisions; the pictures are the real frames, re-rendered.
pedestrians counted
—
miss rate, chosen detector
—
against the exam's rule
—
best to worst of the nine
—
closer than 12 m: counted · miss
—
overlap if the mask is half a pixel off
—
Show the core JS
L.counts = function (r, rule) {
    case 'pixel': return r.area >= 1;
    case 'exam': return r.area >= 6 && r.area >= 0.15 * r.full;
    case 'half': return r.area > 0 && r.area >= 0.5 * r.full;
    case 'amodal': return r.full >= 6;
  for (i = 0; i < rows.length; i++) if (L.counts(rows[i], rule)) { n++; if (scores[i] > thr) hit++; }
  return { n: n, miss: n ? 1 - hit / n : null };
L.sliverArea = function (z, w, h) { return SV.CAM.f * SV.CAM.f * w * h / (z * z); };
    sum += a - A; worst = Math.max(worst, Math.abs(a - A));
    worst = Math.max(worst, Math.abs(map[q] / (C.f * C.hc / (j + 0.5 - C.V0)) - 1)); n++;
L.rangeFactor = function (tanPhi, rho) { return Math.sqrt(1 + tanPhi * tanPhi + rho * rho); };

What to try. As the page opens, the program's detector on the rare case under the exam's rule: 2,404 pedestrians counted, 58.1% missed (seed 1; the tables' three-seed mean is 58.4%). Cut to 0 with the 6 px² box unticked: any pixel, 2,824 counted, 63.2% missed; cut at 50%: 1,518 counted, 43.7%. Switch to the exam's frames and repeat: 50.5% to 46.6%, because most pedestrians clear the cut. Back on the rare case, step through the nine detectors at the exam's rule: best to worst is 9.6 points, where the cut alone moves the program's detector by 19.5. Read the area with one ray per pixel, what an object-index pass delivers: 2,376 counted, 57.8% missed; after the default filter: 2,043 counted, 52.1%, six points from a reader with the detector unchanged. Choose range: of the 2,398 counted, all are closer than 12 m by depth and 2,022 by range. Tick half a pixel off on the exam's frames: violet marks where the displaced mask lands, and its overlap with the true mask is 57%.

The other column comes from nine retrained detectors. Six seeds, the differences paired by seed (the same 1,600 frames, only the label differs), because several differences are of the size of the seed wobble:

Clause: the label taughtCells changed (seed 1)Taught: exam's miss against the program's labelScored: the same detector, the clause changed in the rule
1 · whole silhouette (amodal)564 added (10.1%), all hidden+1.6 ± 0.2; rare case −2.8; false alarms within 2 px of the van 38% → 47%count the silhouette: 1,508 counted, 51.8% missed against 46.8
2 · 15% / 50% of the silhouette23 / 176+0.2 ± 0.1 / +0.1 ± 0.5any pixel 50.7, half 46.7, 35% hidden 44.7; the rare case up to 25.0 apart
3 · range for depthnone: the detector never sees a distance—3.9% of the exam's pedestrians change distance bin (12, 16, 20 m); closer than 12 m: 289 by depth, 266 by range
4 · mask half a pixel / one pixel off2,200 / 4,395 (39% / 79%)+2.3 ± 0.4 / +5.8 ± 0.5overlap with the true mask 57% / 32%, below 50% for 25% / 89% of pedestrians; the foot row moves for 51%, the distance read from it by 8%
5 · 8×8 rays / one ray / default filter908 / 1,096 / 960−0.4 ± 0.4 / −0.1 ± 0.3 / −1.8 ± 0.5the exam's decision flips for 2.5% / 5.2% / 4.1% of its pedestrians (§6)
6 · which instantnot measured

Three things stand out. Reporting effects are large and cost nothing: no detector was retrained to get the right-hand column. The clauses that move the detector are not the ones that move the number. The visibility rule, argued about most, decides 0.4 to 3.2% of the positive cells and moves the detector by less than its wobble; so do 8×8 rays and one ray, which change 16 to 20%. A mask half a pixel off moves 39% of the cells and costs 2.3 points in all six seeds, though no frame-level number of the exam depends on where a mask lies; the whole silhouette adds cells that show only van and costs 1.6. The sign of a clause is not known in advance. The softened mask of the default filter removes 18% of the positive cells and adds none: an edge pixel that is exactly half covered, which the lab's rule counts, reads just under one half once blurred, and thin parts are mostly edge, so 32% of the head's cells go, 24% of the feet's, and 75% of all that go lie on the outer columns. The detector taught on it is −1.8 points better on the exam, in all six seeds.

And a rule never reorders detectors. Across the 16 programs of Lesson 2 (seed 1, exam miss from 46.7 to 80.5%) and the 11 designs of Lesson 6, the order under each of the seven rules is the order under the exam's (rank correlation at least 0.998), and any pair that swaps places is within 0.6 points under one of the two rules. What reorders two detectors is the population. The amodal-taught detector misses −2.8 points fewer of the rare case than the program's and 1.6 more of the exam's pedestrians, in all six seeds; the default filter's does the opposite, −1.8 better on the exam and 1.2 worse on the rare case. Whose misses are averaged is part of the definition too.

Road not taken · the more complete label
An amodal mask is the more informative label: it says where the whole pedestrian is, and datasets are built to provide it. For a detector that sees only pixels it costs 1.6 points on the exam in every seed, moves false alarms near the van from 38% to 47%, and pays only on the rare case, where the hidden part is largest. A label that contradicts the pixels is the worse label for a detector that must find the pedestrian once they show. The road returns as a choice of whose misses count, which is a definition and belongs in the contract.

5 · The contract

A clause that is a habit of one author's code is invisible until two pipelines meet. The evidence for the label stage is what the benchmark wrote down, an afternoon of reading, and what a renderer returns for scenes whose answers are known in closed form: a few scenes, no labels and no street. So write each clause as a setting and a test.

ClauseThe lab's settingTestA violation looks like
1 · which pixelsmodal labels, silhouette rendered alonethe visible mask lies inside the silhouettefalse alarms on the van (47% against 38%)
2 · how much counts15% and 6 px² for the exama sliver of known size: the decision equals the closed formone detector at 44.7% or 51.8%
3 · how fardepth along the axison the ground, depth equals f·hc/(v + 0.5 − V0) at every pixela range pass is off by up to 0.165
4 · aligned howcentres at +0.5a sliver at a known sub-pixel place: its mask's centroid is thereoverlap 57%, +2.3 points
5 · counted how finely2 × 2 rays, sum of coveragea 6.0 px² sliver at sixteen sub-pixel offsets: the counted area stays within a stated errorworst error 4.0 px² with one ray, 1.5 with 2 × 2, 0.25 with 8 × 8
6 · which instantone instanta target at a known speed: the label is where the target was at the exposure's midpointan offset of f·v·Δt/z pixels

The listing under the dial holds the counting rules and two of these tests. The coverage test places a 1.2 × 5 pixel sliver, whose area is exactly 6.0 px², at sixteen sub-pixel offsets and reports the worst error of the counted area; the distance test reads a frame's depth pass on the ground and reports its largest deviation from the ground-plane law. The lab passes it to 0.0000, and a range map in its place fails it by 0.165. At two rays a 6 px² sliver is counted within 1.5 px², 25% of its area, so "at least 6 px²" is a coin toss for it: the rule is applied to a number only that accurate, and the contract says so.

6 · A renderer we did not write

All six clauses were settled in street.js by one author, who chose every convention. A production pipeline settles them with a renderer's defaults. In Blender 5.2 LTS, whose defaults were checked by running it, the Depth pass is the distance along the camera axis and the Mist pass the Euclidean distance; the Object Index pass is point-sampled at the pixel centre, a hard label with no coverage, where Cryptomatte stores coverage-weighted ids; and the default pixel filter, Blackman-Harris 1.5 px wide, lets a step edge that lies on a pixel boundary leak 12% into the next pixel, where the box filter gives exact area coverage. Each is a setting of a clause above, chosen by someone else.

The filter is the worked example. A Gaussian blur of standard deviation σ, read at a pixel centre half a pixel from the edge, leaks Φ(−0.5/σ) (Φ is the standard normal distribution function), so a leak of 12% means σ = 0.43 px. The lab's own pixel is a box of standard deviation 0.29 px; variances add, so the default filter's extra blur is 0.31 px. Read the exam's pedestrians' area through each of these, same detector (seed 1), the exam's rule:

Area is read asCountedDecisions that differ from the exam'sMiss, exam's framesMiss, rare caseTaught on that mask: exam's miss against the program's
the lab's: sum of 2 × 2 coverage1,323046.7%58.1%0
sum of 8 × 8 coverage1,32033 (2.5%)46.4%58.0%−0.4
one ray per pixel centre (an index pass)1,35869 (5.2%)47.8%57.8%−0.1
pixels at least 0.5 after the default filter1,26954 (4.1%)44.8%52.1%−1.8

None of these is wrong. Each is a different definition chosen by a default, and each changes a number a team will report, 54 pedestrians of the exam's frames and six points of the rare case for the filter alone, or a detector it will train, in a direction nobody chose.

What this lesson did not do
The labels here are exact: a human pipeline adds a noise process, disagreement between annotators, whose evidence is its cost. Which instant was written down and not measured, and the box formats were named, not run. The finding that teaching rarely depends on how much counts belongs to this learner, a logistic head on fixed filters, with cells labelled by a pixel threshold; one that regresses boxes or masks would meet clauses 1 and 4 more directly. Blender was described only as far as it was checked by running it, and nothing here ran the contract against a renderer other than the lab's: Lesson 8 runs it against Blender.

Common mistakes / failure modes

"The label stage costs nothing"
As a training label here it costs 0.3 points; as a scoring rule it moves the rare case by 25.0 (§1, §2).
"Teach under the rule you will be scored by"
The detector taught at 50% is 2.4 points worse under the half rule than the program's (§2).
"A more complete label is a better label"
The amodal mask costs 1.6 points on the exam and moves false alarms to the van (§4).
"It is only half a pixel"
39% of the positive cells change, the overlap with the true mask is 57%, the detector loses 2.3 points (§4).
"Depth and range are the same distance"
By range 17% of the close step-outs are beyond 12 m, and every weight of Lesson 6 moves with them (§3).
"A 6 px² rule is exact"
At two rays the counted area of a 6 px² sliver is off by up to 1.5 px² (§5).

Checkpoint exercise

Try it
A pedestrian stands 2.5 m right of the axis at depth 11.9 m, the camera 1.4 m above the road. Is the pedestrian in Lesson 6's case R, closer than 12 m, by depth? by range? And a sliver 1.2 px wide and 5 px tall is displaced from its true place by half a pixel in both directions: what is the overlap (intersection over union)? Answer: by depth the pedestrian is inside, 11.9 m. The range is √(2.5² + 11.9² + 1.4²) = 12.24 m, outside: defined by range the pedestrian is not in case R, and across the street's step-outs 17% of the close ones leave it. The overlap is (1.2 − 0.5)(5 − 0.5) = 3.15 px² over 6 + 6 − 3.15 = 8.85 px², 0.36: below the 0.5 that KITTI asks of a pedestrian, from the label alone.

Where this points next

The clauses of a label are written down and measured. The visibility rule moves the number a team reports by up to 25 points and the detector by about one: the same detector scores 44.7% or 51.8% on the exam depending on the rule. What moves a trained detector is a clause nobody argues about, a mask half a pixel off, the hidden part of a silhouette, a softened edge, by 2.3 and 1.6 points worse and −1.8 better. All of it was settled in a renderer we wrote ourselves, where we chose every convention, and a production renderer's defaults decide, silently, how a sliver is counted: an exam's pedestrians change status in 5.2% of cases when their area is read from an index pass and in 4.1% through the default filter. A production pipeline uses a production renderer whose conventions we did not choose. How do we make a renderer we did not write obey the same contract?

Takeaway
A label is a definition: a function of what the renderer returns, with one clause for each question a mask of a pedestrian leaves open, which pixels, how much of the silhouette counts, how far, aligned how, counted how finely, at which instant. It enters a project through two doors, the target a detector is taught and the rule that selects whom a miss rate is averaged over, and different clauses move them. Through the scoring door the visibility rule moved one detector's miss rate on the rare case by 25.0 points and never reordered two detectors; through the teaching door it moved the detector by under a point, while a mask half a pixel off cost 2.3 points, the hidden part of the silhouette 1.6, and a softened edge changed it by −1.8, an improvement. What reorders detectors is the population whose misses are averaged. Two reported numbers are comparable only if their definitions agree clause by clause, so the clauses are written as a contract with a test each, whose answers are known in closed form.

Interview prompts

Companion reads: Computer Graphics · 06 Sampling and antialiasing (coverage, rays per pixel and pixel filters), Computer Vision · 09 Object detection (overlap and the matching rule) and Computer Vision · 04 Cameras and projection geometry (pixel centres and the pinhole).