The rare case
Lesson 5 ended on a frequency. The case that leaves the shuttle least time to stop, a pedestrian who steps out from behind the van closer than 12 m, is drawn in one frame in 122, so 1,600 training frames hold 13 of them, and the detector misses 58% of them where it misses 9% of the pedestrians who stand in the open at the same distances. A program, unlike the street, can choose how often to draw a case. This lesson spends renders on the rare case on purpose and measures what the detector makes of the extra examples: they change what the loss believes, a weight restores the street at a price in effective sample size, and the share of renders and the weight of a case are two different decisions. It cannot say what a miss costs or what counts as a pedestrian.
New idea: drawing a case on purpose changes the distribution the learner sees, and a weight, target over proposal, converts it to any distribution we choose for a price in effective sample size. The share of renders sets how noisy the detector is and the weight sets which detector it is: weighted back to the street it misses 59% of the rare case, as drawn at a tenth of the renders it misses 51%.
Forces next: Drawing the rare case on purpose, 12 times as often as it occurs, gives the detector 160 examples of it instead of 13 for the same 1,600 renders, and what it makes of them depends on the weight each carries: as drawn they say the case is common, and it misses 8 points fewer of the rare case and 1 point more of everything else; weighted by p/q they restore the street, and the detector is the natural one at an effective sample size of 1,463 frames. The pictures, the cases and their frequencies are now right, and the training data and the exam agree about what is in a frame. They may still disagree about what counts: a pedestrian with two pixels showing is a pedestrian to the program and not to an exam that asks for a visible fraction, and the miss rate on the rare case moves by 19 points with the rule. What does a label mean, and where can the program's definition differ from the exam's?
1 · The case the exam barely sees
Call case R a pedestrian who steps out from behind the van at a depth under 12 m. The exact-stage program of Lesson 5, bbba (the street's scene, light and camera, the program's own label rule), draws scenes from a law it can run, so it counts its own frequencies without a label: in 106 scene draws case R came up 8,193 times, 0.822% of the frames, one in 122, and an exact integral of the scene law agrees. A set of 1,600 frames holds 13. Not all show enough of the pedestrian to count: 80% pass the exam's rule (15% of the silhouette and 6 px²), so the exam's own count is one frame in 152, and the real test set of 3,000 frames holds 19, too few to grade a slice.
So the lab grades the slice another way. It draws the street directly in case R, 3,000 frames of which 2,404 hold a pedestrian that counts, and 3,000 frames with an open pedestrian closer than 12 m. A project cannot buy these frames; this is a lab privilege. The detector is the one the exact-stage program trains on 1,600 frames (three seeds), at the exam's threshold. Lesson 5 graded the same slice on 1,200 frames of its own and read slightly different figures; the 3,000 frames here are the larger sample, and the figures of this lesson replace those.
| Slice of the street | Pedestrians that count | Miss rate |
|---|---|---|
| case R: steps out from behind the van, closer than 12 m | 2,404 | 58.4% |
| open, closer than 12 m | 3,000 | 8.9% |
| the exam: every pedestrian that counts, real test set | 1,323 | 46.8% |
A detector that finds nine in ten of the pedestrians standing at those distances misses nearly six in ten of those who step out at them, and the exam cannot see it: case R is 1.5% of the pedestrians the exam counts, so a detector that found every one would lower the exam by 0.9 points, about the 1.1 points by which three seeds of one program differ. The exam prices a case by how often the street draws it; for a case that matters for another reason it is the wrong instrument, so this lesson grades slices.
The remedy that suggests itself is more data (three seeds up to 3,200 frames, one above):
| Training frames | Frames of R among them | Miss on the rare slice | Miss, the exam |
|---|---|---|---|
| 400 | 4 | 61.7% | 49.6% |
| 1,600 | 13 | 58.4% | 46.8% |
| 11,200 | 73 | 57.1% | 46.0% |
From 1,600 to 11,200 frames, seven times the renders and 5.5 times the examples of R, the rare slice moves by 1.2 points. The case is rare in the data and more of the same data does not touch it. What a program can do that a street cannot is choose what it draws.
2 · Where the renders can go
To get more examples of R a project has three roads, which differ in how many frames must be produced per example.
| Option | Frames rendered per example of R | What it costs | What goes wrong |
|---|---|---|---|
| (a) Wait: draw naturally, keep everything | 122 | 80 examples need about 9,700 frames, 6.1 times the renders | six times the data the common cases did not need; §1 shows the slice does not move |
| (b) Draw naturally, keep every R frame and a fraction of the rest | 122 rendered, one kept | 160 examples among 1,600 kept frames: about 19,500 rendered | the same training set as (c) at 12 times the renders |
| (c) Draw R on purpose | 1, plus 17 scene draws | a conditional scene sampler | the set is no longer the street (§3) |
Option (c) is the program's privilege: a fleet collects what it meets, a program is asked for what it wants. The sampler is rejection on the scene law. The scene program draws the van first and draws everything else without asking whether a pedestrian will step out, so forcing the step-out gives exactly the street's law of the scene given a step-out, and keeping the draws with z < 12 m gives its law given R. A forced draw is kept with probability 5.9%, so a kept frame costs 17 scene draws, microseconds each, and one render. Eight thousand sampled scenes and the R scenes found among the 106 natural draws agree on the pedestrian's depth and position and the van's depth and length: the largest Kolmogorov-Smirnov distance is 0.02, where chance allows 0.031. The program can now ask for any share q of its frames to be case R. What should it ask for, and what does the answer do to the detector?
3 · What a mixture does to the statistics
Let p be the street's frequency of R and q the share the generator draws: with probability q a frame from the street's law given R, otherwise one from its law given anything else, both exactly as the street has them. The detector minimises an average of a loss ℓ over its frames, an expectation under the generator's law; the street's loss is the expectation under its own. With μR and μo the mean loss in each case, the link is one line:
Eq[ w·ℓ ] = q·(p/q)·μR + (1 − q)·((1 − p)/(1 − q))·μo = p·μR + (1 − p)·μo = Estreet[ ℓ ]
Weighting each frame by the density ratio, wR = p/q for a frame of R and wo = (1 − p)/(1 − q) for any other, makes the generator's average the street's in expectation. That is the importance-sampling identity: the plain weighted mean is unbiased, the self-normalised form that divides by the weights' sum only consistent (Owen, Monte Carlo theory, methods and examples, chapter 9). It needs the generator to draw whatever the street draws: no weight repairs a case with probability zero, which is Lesson 5's point again. At q = 25% the weights are 0.033 and 1.32.
Without the weights the learner's world is the mixture. A classifier fits P(pedestrian | picture) under the frames it sees, and by Bayes' rule its odds are the street's odds times the ratio of training prior odds to the street's. A mixture changes the prior in two ways that behave differently.
(i) A shift of the label. The share of frames that hold a pedestrian changes. Every picture's odds are multiplied by the same constant, so for a logistic model that is right the intercept moves by the log of the ratio of prior odds and the slopes stay (Saerens, Latinne and Decaestecker, 2002; King and Zeng, 2001), the ranking of frames does not move, and the exam, which fixes its threshold at 10% false alarms on pedestrian-free frames, absorbs it. This assumes frames are chosen by label alone; the detector also picks hard negatives by score, so the absorption is approximate, and §5 measures how far.
(ii) A shift inside the positive class. More of the pedestrians are R: 1.6% of the pedestrian frames of a natural set, 39% at q = 25%. Their odds of being R are multiplied by c(q) = q(1 − p) / (p(1 − q)) = 40.2. Only pictures that look like R are multiplied, and a sliver of pedestrian beside a van is not far from a van's edge, so no threshold undoes it: the loss now says R is 40.2 times as important as the street thinks, and the detector spends its capacity accordingly. The weights undo it, because they put the street's odds back.
4 · What the weights cost: effective sample size
The weights repair the expectation, not the variance. A weighted mean Σ wi fi / Σ wi of independent values of equal spread σ has variance σ² Σ wi² / (Σ wi)², and an unweighted mean of neff values has σ² / neff; they agree when
neff = (Σ wi)² / Σ wi²
This is Kish's effective sample size (Kish, 1965, used for importance sampling after Kong, 1992), a diagnostic that reads the weights alone and is the exact variance factor only when every frame's loss has the same spread, which §6 relaxes. With qN frames at weight p/q and the rest at (1 − p)/(1 − q) the sums are closed forms:
neff / N = 1 / ( p²/q + (1 − p)²/(1 − q) ) ≈ 1 − q
| Share of renders in R | R frames in 1,600 | Weight of an R frame, of any other | Effective sample size | R's odds multiplied by |
|---|---|---|---|---|
| 10% | 160 | 0.082, 1.10 | 1,463 | 13.4 |
| 25% | 400 | 0.033, 1.32 | 1,220 | 40.2 |
| 50% | 800 | 0.016, 1.98 | 813 | 121 |
Read the 25% row: the 400 frames of R weigh 13 in all, the weight of the 13 natural ones. Under the weights oversampling is a reallocation: the share q of the renders goes to R and the rest keep N(1 − q) frames to stand for the other 99.2% of the street. Over- and under-sampling are one trade: the shares fix the weights and the effective sample size, and only the number of frames rendered differs (option (b) of §2).
What to try. At the natural share the plan holds 11 frames of R; the detector misses 58.4% of the rare slice and 8.9% of open pedestrians at the same distances, and finds 1 of the eight step-outs and 7 of the eight open ones. The last panel shows why: it misses 96% of the first bin of visible fraction and 36% of the last. Slide to 10%: the rare slice falls to 50.7%, 7.7 points, while the exam goes from 46.8% to 47.7%. Tick weights: the slice is back at 59.1%, the effective sample size is 1,463, and the amber dots shrink to the weight of the thirteen they stand for. Untick and slide to 50%: the slice reaches 47.1% and no further, the exam 57.8%, the open pedestrians 17.5%, and the alarm rate at a fixed cut goes from 10.7% to 20.6%. Tick the weights: exam 48.3%, effective sample size 813. Last, set the rule to half of the silhouette at the natural share: the slice reads 44.2%, against 63.4% for any pixel.
5 · What the lab shows
Seven designs, three seeds each, 1,600 frames each, one detector, one exam (the means; the widget plots them):
| Design | Effective sample size | Rare slice | Open, same range | The exam | Alarms at a fixed cut |
|---|---|---|---|---|---|
| natural | 1,600 | 58.4% | 8.9% | 46.8% | 10.7% |
| 10%, as drawn | 1,600 | 50.7% | 9.8% | 47.7% | 11.8% |
| 10%, weighted | 1,463 | 59.1% | 9.2% | 47.4% | 11.2% |
| 25%, as drawn | 1,600 | 47.5% | 12.4% | 51.5% | 14.6% |
| 25%, weighted | 1,220 | 59.6% | 9.7% | 48.0% | 11.6% |
| 50%, as drawn | 1,600 | 47.1% | 17.5% | 57.8% | 20.6% |
| 50%, weighted | 813 | 60.3% | 10.4% | 48.3% | 12.0% |
As drawn, the examples help the slice and the loss pays for it. The rare slice falls by 4.1, 7.7, 10.9 and 11.3 points at 5, 10, 25 and 50%, and stops: about 47% is what this detector can do on R. The cost has no floor: the exam rises by 0.6, 0.9, 4.6 and 11.0 points, and the score that lets 10% of the empty street's frames alarm moves from 0.07 to 0.16, 0.32 and 0.59. This is shift (ii) of §3: the loss says R is 40 times as important, so the detector trades the common cases for it. At a tenth of the renders the exam moves by less than the seed wobble; past a quarter the trade only loses.
Weighted, the detector is the natural one at a smaller size. The rare slice stays at 59.1, 59.6 and 60.3%: no better than natural, a little worse. The exam is the natural program's at the effective sample size: the weighted designs score 47.4, 48.0 and 48.3%, and the natural program trained on exactly 1,463, 1,220 and 813 frames scores 47.9, 48.4 and 48.5%. What the weights buy is the street's exam, not the slice: 400 frames of R that weigh as much as the 13 natural ones teach the detector what those teach it. The brute-force control agrees: the natural program at 11,200 frames holds 73 frames of R, as many as the 5% design (80), and misses 57.1% on the slice where that design, as drawn, misses 54.3%. The examples were never what the detector lacked. The weight on them was.
The weight chooses the detector, the share chooses the variance. Draw 50% of the frames in R and weight them to restore a world in which R is a quarter of the frames (weights 0.5 and 1.5): the rare slice is 47.8% and the exam 51.9%. Draw 10% and weight the same way (2.5 and 0.83): 47.4% and 51.5%. As drawn at 25% the detector scores 47.5% and 51.5%. One detector three ways; the off-target designs keep an effective sample size of 1,280 where drawing the target itself keeps 1,600. A mixture as drawn is a weighting in disguise, a street in which R has the odds c(q); the weights t/q ask for the same street from any share.
Scores keep their scale only under the weights. The exam re-sets its threshold, so it cannot see that the unweighted detector's scores no longer mean what they meant: at a fixed cut the alarm rate on empty street frames goes from 10.7% to 14.6% and 20.6% at 25% and 50%, and stays near 11.6% and 12.0% when weighted; a braking rule that reads a score as a probability would inherit the wrong prior (Lesson 11). A pure shift of the label, the share of frames with a pedestrian moved from 50% to 90% or 10%, moves the threshold by +0.64 and −1.43 and the exam by +1.6 and −2.1 points, and the rare slice by at most 1.1: shift (i) is mostly absorbed by the threshold, not entirely, and cannot buy the case.
6 · Choosing the share and the weight
The lab has pulled apart two decisions that an unweighted mixture makes together. The target is the street as the job weighs it: an ordinary pedestrian counts 1 and a pedestrian in R counts c, a number this lesson cannot price (Lesson 11 does) and that the exam sets to 1; its share of R is t = cp/(1 − p + cp). The share is how the renders are spent to reach it. For a given target the weighted loss Σs ts μ̂s from ns frames in stratum s has variance Σs ts² σs² / ns, with σs the spread of a frame's loss there. Minimising it under Σs ns = N (the derivative of the variance plus a multiplier times Σ ns vanishes when ts² σs² / ns² is the same for every stratum) gives
ns ∝ ts σs
the classical optimal allocation of stratified sampling: spend renders in proportion to a stratum's weight in the target times the spread of its losses. The natural detector gives the spreads: a frame of R holds a pedestrian that counts and is missed with probability 0.47, so σR = 0.50; for any other frame it is 0.20 and σo = 0.40.
| A miss in R counts | Target share of R | Best share of renders | Exam with R counted so: natural | 10% as drawn | 25% as drawn |
|---|---|---|---|---|---|
| c = 1 (the exam) | 0.8% | 1.0% | 46.8% | 47.7% | 51.5% |
| c = 10 | 7.7% | 9.4% | 48.2% | 48.1% | 51.0% |
| c = 100 | 45.3% | 50.9% | 53.7% | 49.5% | 49.1% |
The first two columns are arithmetic (the formula, checked against a numerical minimiser); the rest is the lab: the exam's miss rate when a pedestrian in R counts c times, for designs as drawn. At the exam's c = 1 the best share is the street's own, 1.0%, and any oversampling only costs: the case of §5. At c = 10 the best share is 9.4% and the 10% design just breaks even with the natural one (the break-even is c = 8). At c = 100 the arithmetic says half the renders, but the best design in the lab is the 25% one, the 50% design reading 51.4%: the detector cannot get below 47% on R, so weight beyond a quarter of the renders is paid for and not used. What c is worth is Lesson 11's question.
7 · Which pedestrians count in the rare slice
Every miss rate here is over the pedestrians that count, and in the rare slice counting is least obvious: a pedestrian who steps out is a sliver first. The last panel of the widget cuts the detector's miss rate by how much of the silhouette shows. Up to a quarter of it the detector misses about nine in ten (96, 92, 93 and 92% in the four bins), 66% when half shows and 36% when more than four fifths does. A rule that decides who counts is a cut on this curve, and the number it reports is the average of the curve above the cut.
| Rule | Who uses it | R frames that count | Miss, natural | Miss, 10% as drawn |
|---|---|---|---|---|
| at least 1 px² visible | the program (kvis = 1) | 94.1% | 63.4% | 56.7% |
| 15% of the silhouette and 6 px² | the street's label, and the exam | 80.1% | 58.4% | 50.7% |
| half of the silhouette | a stricter rule, to show the range | 50.6% | 44.2% | 35.2% |
In 13% of the R frames in which any pixel shows, only a sliver under 6 px² does; the program labels it a pedestrian and the exam does not count it, so the same detector scores 5.0 points worse under the program's rule than under the exam's, and 14.2 points better under a rule that asks for half of the silhouette. The number of the slice that §6 told us to spend renders on moves by 19.2 points with the definition of who counts, more than any share of the renders moves it (7.7 at 10%). Training frames and exam agree about what is in a frame and how often; they do not yet agree about what counts.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
The program now draws the street's cases, draws a rare one on purpose when it must, and says with its weights what its frames stand for. The detector, the exam and the weights agree about what is in a frame and how often it happens. They do not agree about what counts: every number in the rare slice is a miss rate over the pedestrians that count, and counted by the program's rule, the exam's, or one that asks for half of the silhouette, the same detector misses 63.4%, 58.4% or 44.2% of them. A pedestrian with two pixels showing is a pedestrian to the program and not to the exam, which asks for a visible fraction. What does a label mean, and where can the program's definition differ from the exam's?
Interview prompts
- A class is 1% of production traffic and you train on a set where it is 25%. What changes in the model, and what do you do about it? (§3 — its odds are overrated by the ratio of prior odds; weight by p/q to restore the street, or re-set the threshold if only the label prior moved.)
- Define the effective sample size of a weighted sample and say what it assumes. (§4 — (Σw)²/Σw², the size of an unweighted sample with the same variance of the mean; it reads the weights alone, so it is exact only when every frame's loss has the same spread.)
- Oversampling improved your slice metric and hurt the overall metric. Is that a bug? (§5 — no: an unweighted mixture trains for a world where the case is c times as important; the trade is real and the gain has a floor.)
- What decides the trained model, the share of renders or the weights? (§5 — the weights: three shares weighted to one target gave one detector, and the share changed only the effective sample size.)
- Why can a program draw a rare case on purpose when a data campaign cannot? (§2 — it conditions its own scene law, here by rejection on a forced step-out, at one render per kept frame.)
- How would you choose the share of renders for a rare case? (§6 — fix the target weight first, then spend in proportion to target weight times spread; at the exam's cost that is the natural frequency.)
- A benchmark changes its visibility rule and your slice metric moves 15 points. Is the model worse? (§7 — no: a rule is a cut on the miss-versus-visible-fraction curve and the reported number is the average above the cut.)
Companion reads: Reinforcement Learning · 12 Importance sampling (the same density ratio for policies) and Computer Vision (the detector this track holds fixed).