all_lessons/Synthetic Vision Data/06 · The rare caselesson 6 / 12

The rare case

Lesson 5 ended on a frequency. The case that leaves the shuttle least time to stop, a pedestrian who steps out from behind the van closer than 12 m, is drawn in one frame in 122, so 1,600 training frames hold 13 of them, and the detector misses 58% of them where it misses 9% of the pedestrians who stand in the open at the same distances. A program, unlike the street, can choose how often to draw a case. This lesson spends renders on the rare case on purpose and measures what the detector makes of the extra examples: they change what the loss believes, a weight restores the street at a price in effective sample size, and the share of renders and the weight of a case are two different decisions. It cannot say what a miss costs or what counts as a pedestrian.

The thesis, here
A generator may spend its renders where it likes, but a learner takes the frequencies it is shown for the truth. Drawing a case on purpose is therefore two decisions that are easy to confuse: how many examples of it to draw, which sets the variance, and how much each example counts, which sets the detector. The weight that restores the street is the density ratio p/q, and what it costs is the effective sample size.
Linear position
Forced by: Counting cases on a small labelled sample, with each pedestrian's distance read off the foot row of its outline and a nearer van touching it as the sign of a step-out, shows that the scene program draws cells holding only 42% of the street's pedestrians: 51% of them stand beyond the 16 m it ever drew and 17% step out from behind a van it never lets hide anyone. Setting the distance range and the step-out probability from 200 outlined frames takes the miss rate from 62% to about 52%, as far as the street's own pedestrian settings do, and the clutter, once the pedestrians are right, takes it to 47%. But a cell is a box, not a frequency: the step-outs nearest the shuttle, which leave the least time to stop, occur in one frame in 122, so a training set of 1,600 frames holds about 13 of them, and the detector trained with the street's scene misses 57% of them against 10% of the pedestrians as near in the open. How should a generator spend its renders on cases that rarely happen, and what does that do to the statistics the model learns?
New idea: drawing a case on purpose changes the distribution the learner sees, and a weight, target over proposal, converts it to any distribution we choose for a price in effective sample size. The share of renders sets how noisy the detector is and the weight sets which detector it is: weighted back to the street it misses 59% of the rare case, as drawn at a tenth of the renders it misses 51%.
Forces next: Drawing the rare case on purpose, 12 times as often as it occurs, gives the detector 160 examples of it instead of 13 for the same 1,600 renders, and what it makes of them depends on the weight each carries: as drawn they say the case is common, and it misses 8 points fewer of the rare case and 1 point more of everything else; weighted by p/q they restore the street, and the detector is the natural one at an effective sample size of 1,463 frames. The pictures, the cases and their frequencies are now right, and the training data and the exam agree about what is in a frame. They may still disagree about what counts: a pedestrian with two pixels showing is a pedestrian to the program and not to an exam that asks for a visible fraction, and the miss rate on the rare case moves by 19 points with the rule. What does a label mean, and where can the program's definition differ from the exam's?
The plan
Seven moves. (1) Measure the rare case: more natural data does not repair it. (2) Compare the ways to get more examples. (3) Derive what a mixture does to what the learner believes, and the weight that undoes it. (4) Price the weight. (5) Measure it on the street. (6) Choose the share and the weight. (7) Ask which pedestrians count in the rare slice.

1 · The case the exam barely sees

Call case R a pedestrian who steps out from behind the van at a depth under 12 m. The exact-stage program of Lesson 5, bbba (the street's scene, light and camera, the program's own label rule), draws scenes from a law it can run, so it counts its own frequencies without a label: in 106 scene draws case R came up 8,193 times, 0.822% of the frames, one in 122, and an exact integral of the scene law agrees. A set of 1,600 frames holds 13. Not all show enough of the pedestrian to count: 80% pass the exam's rule (15% of the silhouette and 6 px²), so the exam's own count is one frame in 152, and the real test set of 3,000 frames holds 19, too few to grade a slice.

So the lab grades the slice another way. It draws the street directly in case R, 3,000 frames of which 2,404 hold a pedestrian that counts, and 3,000 frames with an open pedestrian closer than 12 m. A project cannot buy these frames; this is a lab privilege. The detector is the one the exact-stage program trains on 1,600 frames (three seeds), at the exam's threshold. Lesson 5 graded the same slice on 1,200 frames of its own and read slightly different figures; the 3,000 frames here are the larger sample, and the figures of this lesson replace those.

Slice of the streetPedestrians that countMiss rate
case R: steps out from behind the van, closer than 12 m2,40458.4%
open, closer than 12 m3,0008.9%
the exam: every pedestrian that counts, real test set1,32346.8%

A detector that finds nine in ten of the pedestrians standing at those distances misses nearly six in ten of those who step out at them, and the exam cannot see it: case R is 1.5% of the pedestrians the exam counts, so a detector that found every one would lower the exam by 0.9 points, about the 1.1 points by which three seeds of one program differ. The exam prices a case by how often the street draws it; for a case that matters for another reason it is the wrong instrument, so this lesson grades slices.

The remedy that suggests itself is more data (three seeds up to 3,200 frames, one above):

Training framesFrames of R among themMiss on the rare sliceMiss, the exam
400461.7%49.6%
1,6001358.4%46.8%
11,2007357.1%46.0%

From 1,600 to 11,200 frames, seven times the renders and 5.5 times the examples of R, the rare slice moves by 1.2 points. The case is rare in the data and more of the same data does not touch it. What a program can do that a street cannot is choose what it draws.

2 · Where the renders can go

To get more examples of R a project has three roads, which differ in how many frames must be produced per example.

OptionFrames rendered per example of RWhat it costsWhat goes wrong
(a) Wait: draw naturally, keep everything12280 examples need about 9,700 frames, 6.1 times the renderssix times the data the common cases did not need; §1 shows the slice does not move
(b) Draw naturally, keep every R frame and a fraction of the rest122 rendered, one kept160 examples among 1,600 kept frames: about 19,500 renderedthe same training set as (c) at 12 times the renders
(c) Draw R on purpose1, plus 17 scene drawsa conditional scene samplerthe set is no longer the street (§3)

Option (c) is the program's privilege: a fleet collects what it meets, a program is asked for what it wants. The sampler is rejection on the scene law. The scene program draws the van first and draws everything else without asking whether a pedestrian will step out, so forcing the step-out gives exactly the street's law of the scene given a step-out, and keeping the draws with z < 12 m gives its law given R. A forced draw is kept with probability 5.9%, so a kept frame costs 17 scene draws, microseconds each, and one render. Eight thousand sampled scenes and the R scenes found among the 106 natural draws agree on the pedestrian's depth and position and the van's depth and length: the largest Kolmogorov-Smirnov distance is 0.02, where chance allows 0.031. The program can now ask for any share q of its frames to be case R. What should it ask for, and what does the answer do to the detector?

3 · What a mixture does to the statistics

Let p be the street's frequency of R and q the share the generator draws: with probability q a frame from the street's law given R, otherwise one from its law given anything else, both exactly as the street has them. The detector minimises an average of a loss ℓ over its frames, an expectation under the generator's law; the street's loss is the expectation under its own. With μR and μo the mean loss in each case, the link is one line:

Eq[ w·ℓ ] = q·(p/q)·μR + (1 − q)·((1 − p)/(1 − q))·μo = p·μR + (1 − p)·μo = Estreet[ ℓ ]

Weighting each frame by the density ratio, wR = p/q for a frame of R and wo = (1 − p)/(1 − q) for any other, makes the generator's average the street's in expectation. That is the importance-sampling identity: the plain weighted mean is unbiased, the self-normalised form that divides by the weights' sum only consistent (Owen, Monte Carlo theory, methods and examples, chapter 9). It needs the generator to draw whatever the street draws: no weight repairs a case with probability zero, which is Lesson 5's point again. At q = 25% the weights are 0.033 and 1.32.

Without the weights the learner's world is the mixture. A classifier fits P(pedestrian | picture) under the frames it sees, and by Bayes' rule its odds are the street's odds times the ratio of training prior odds to the street's. A mixture changes the prior in two ways that behave differently.

(i) A shift of the label. The share of frames that hold a pedestrian changes. Every picture's odds are multiplied by the same constant, so for a logistic model that is right the intercept moves by the log of the ratio of prior odds and the slopes stay (Saerens, Latinne and Decaestecker, 2002; King and Zeng, 2001), the ranking of frames does not move, and the exam, which fixes its threshold at 10% false alarms on pedestrian-free frames, absorbs it. This assumes frames are chosen by label alone; the detector also picks hard negatives by score, so the absorption is approximate, and §5 measures how far.

(ii) A shift inside the positive class. More of the pedestrians are R: 1.6% of the pedestrian frames of a natural set, 39% at q = 25%. Their odds of being R are multiplied by c(q) = q(1 − p) / (p(1 − q)) = 40.2. Only pictures that look like R are multiplied, and a sliver of pedestrian beside a van is not far from a van's edge, so no threshold undoes it: the loss now says R is 40.2 times as important as the street thinks, and the detector spends its capacity accordingly. The weights undo it, because they put the street's odds back.

4 · What the weights cost: effective sample size

The weights repair the expectation, not the variance. A weighted mean Σ wi fi / Σ wi of independent values of equal spread σ has variance σ² Σ wi² / (Σ wi)², and an unweighted mean of neff values has σ² / neff; they agree when

neff = (Σ wi)² / Σ wi²

This is Kish's effective sample size (Kish, 1965, used for importance sampling after Kong, 1992), a diagnostic that reads the weights alone and is the exact variance factor only when every frame's loss has the same spread, which §6 relaxes. With qN frames at weight p/q and the rest at (1 − p)/(1 − q) the sums are closed forms:

neff / N = 1 / ( p²/q + (1 − p)²/(1 − q) ) ≈ 1 − q

Share of renders in RR frames in 1,600Weight of an R frame, of any otherEffective sample sizeR's odds multiplied by
10%1600.082, 1.101,46313.4
25%4000.033, 1.321,22040.2
50%8000.016, 1.98813121

Read the 25% row: the 400 frames of R weigh 13 in all, the weight of the 13 natural ones. Under the weights oversampling is a reallocation: the share q of the renders goes to R and the rest keep N(1 − q) frames to stand for the other 99.2% of the street. Over- and under-sampling are one trade: the shares fix the weights and the effective sample size, and only the number of frames rendered differs (option (b) of §2).

The mixture dial
Slide the share of renders spent on case R. Top: the 1,600 training frames of the seed-1 detector at that share, one dot each (amber = R; with weights ticked, dot area is the weight). Then eight close step-outs, one per bin of visible fraction, and eight open pedestrians at the same distances, scored by the detector of that share (green border = found, red = missed). Curves are three-seed means of detectors trained on 1,600 frames (precomputed); dots, strips and counts are computed live.
R frames in the 1,600
—
effective sample size (p/q)
—
miss, rare slice
—
miss, open, same range
—
miss, the exam
—
alarms at a fixed cut
—
found, R + open
—
Show the core JS
L.isR = function (scene) { return !!(scene.ped && scene.mode === 'emerge' && scene.ped.parts[0].z < L.ZR); };
if (stratum === 'R') { sc = SV.drawScene(cfg, rng, { ped: true, mode: 'emerge' }); if (sc.ped.parts[0].z < L.ZR) break; }
var n = Math.round(q * N), wR = t / q, wO = (1 - t) / (1 - q);
L.ess = function (w) { var s = 0, s2 = 0, i; for (i = 0; i < w.length; i++) { s += w[i]; s2 += w[i] * w[i]; } return s * s / s2; };
L.essFrac = function (p, q) { return 1 / (p * p / q + (1 - p) * (1 - p) / (1 - q)); };
L.neymanShare = function (p, c, sigR, sigO) { var a = p * c * sigR, b = (1 - p) * sigO; return a / (a + b); };
return pack(fm, cells, ci.lab, o.weights ? o.weights[si] : 1);

What to try. At the natural share the plan holds 11 frames of R; the detector misses 58.4% of the rare slice and 8.9% of open pedestrians at the same distances, and finds 1 of the eight step-outs and 7 of the eight open ones. The last panel shows why: it misses 96% of the first bin of visible fraction and 36% of the last. Slide to 10%: the rare slice falls to 50.7%, 7.7 points, while the exam goes from 46.8% to 47.7%. Tick weights: the slice is back at 59.1%, the effective sample size is 1,463, and the amber dots shrink to the weight of the thirteen they stand for. Untick and slide to 50%: the slice reaches 47.1% and no further, the exam 57.8%, the open pedestrians 17.5%, and the alarm rate at a fixed cut goes from 10.7% to 20.6%. Tick the weights: exam 48.3%, effective sample size 813. Last, set the rule to half of the silhouette at the natural share: the slice reads 44.2%, against 63.4% for any pixel.

5 · What the lab shows

Seven designs, three seeds each, 1,600 frames each, one detector, one exam (the means; the widget plots them):

DesignEffective sample sizeRare sliceOpen, same rangeThe examAlarms at a fixed cut
natural1,60058.4%8.9%46.8%10.7%
10%, as drawn1,60050.7%9.8%47.7%11.8%
10%, weighted1,46359.1%9.2%47.4%11.2%
25%, as drawn1,60047.5%12.4%51.5%14.6%
25%, weighted1,22059.6%9.7%48.0%11.6%
50%, as drawn1,60047.1%17.5%57.8%20.6%
50%, weighted81360.3%10.4%48.3%12.0%

As drawn, the examples help the slice and the loss pays for it. The rare slice falls by 4.1, 7.7, 10.9 and 11.3 points at 5, 10, 25 and 50%, and stops: about 47% is what this detector can do on R. The cost has no floor: the exam rises by 0.6, 0.9, 4.6 and 11.0 points, and the score that lets 10% of the empty street's frames alarm moves from 0.07 to 0.16, 0.32 and 0.59. This is shift (ii) of §3: the loss says R is 40 times as important, so the detector trades the common cases for it. At a tenth of the renders the exam moves by less than the seed wobble; past a quarter the trade only loses.

Weighted, the detector is the natural one at a smaller size. The rare slice stays at 59.1, 59.6 and 60.3%: no better than natural, a little worse. The exam is the natural program's at the effective sample size: the weighted designs score 47.4, 48.0 and 48.3%, and the natural program trained on exactly 1,463, 1,220 and 813 frames scores 47.9, 48.4 and 48.5%. What the weights buy is the street's exam, not the slice: 400 frames of R that weigh as much as the 13 natural ones teach the detector what those teach it. The brute-force control agrees: the natural program at 11,200 frames holds 73 frames of R, as many as the 5% design (80), and misses 57.1% on the slice where that design, as drawn, misses 54.3%. The examples were never what the detector lacked. The weight on them was.

The weight chooses the detector, the share chooses the variance. Draw 50% of the frames in R and weight them to restore a world in which R is a quarter of the frames (weights 0.5 and 1.5): the rare slice is 47.8% and the exam 51.9%. Draw 10% and weight the same way (2.5 and 0.83): 47.4% and 51.5%. As drawn at 25% the detector scores 47.5% and 51.5%. One detector three ways; the off-target designs keep an effective sample size of 1,280 where drawing the target itself keeps 1,600. A mixture as drawn is a weighting in disguise, a street in which R has the odds c(q); the weights t/q ask for the same street from any share.

Scores keep their scale only under the weights. The exam re-sets its threshold, so it cannot see that the unweighted detector's scores no longer mean what they meant: at a fixed cut the alarm rate on empty street frames goes from 10.7% to 14.6% and 20.6% at 25% and 50%, and stays near 11.6% and 12.0% when weighted; a braking rule that reads a score as a probability would inherit the wrong prior (Lesson 11). A pure shift of the label, the share of frames with a pedestrian moved from 50% to 90% or 10%, moves the threshold by +0.64 and −1.43 and the exam by +1.6 and −2.1 points, and the rare slice by at most 1.1: shift (i) is mostly absorbed by the threshold, not entirely, and cannot buy the case.

Road not taken · oversample and say nothing
It is the usual practice, and the first look rewards it: the number the team watches, the rare slice, falls by 7.7 to 10.9 points. What breaks is everything else: the exam rises 4.6 points at a quarter of the renders and 11.0 at a half ("balance the classes"), the open pedestrians at the same distances miss 8.6 points more at a half, the scores drift off the street's scale, and the gain stops at 47%. The road returns as a choice: an unweighted mixture is the target in which R is 13, 40 or 121 times as important, and §6 says how to make that choice on purpose.

6 · Choosing the share and the weight

The lab has pulled apart two decisions that an unweighted mixture makes together. The target is the street as the job weighs it: an ordinary pedestrian counts 1 and a pedestrian in R counts c, a number this lesson cannot price (Lesson 11 does) and that the exam sets to 1; its share of R is t = cp/(1 − p + cp). The share is how the renders are spent to reach it. For a given target the weighted loss Σs ts μ̂s from ns frames in stratum s has variance Σs ts² σs² / ns, with σs the spread of a frame's loss there. Minimising it under Σs ns = N (the derivative of the variance plus a multiplier times Σ ns vanishes when ts² σs² / ns² is the same for every stratum) gives

ns ∝ ts σs

the classical optimal allocation of stratified sampling: spend renders in proportion to a stratum's weight in the target times the spread of its losses. The natural detector gives the spreads: a frame of R holds a pedestrian that counts and is missed with probability 0.47, so σR = 0.50; for any other frame it is 0.20 and σo = 0.40.

A miss in R countsTarget share of RBest share of rendersExam with R counted so: natural10% as drawn25% as drawn
c = 1 (the exam)0.8%1.0%46.8%47.7%51.5%
c = 107.7%9.4%48.2%48.1%51.0%
c = 10045.3%50.9%53.7%49.5%49.1%

The first two columns are arithmetic (the formula, checked against a numerical minimiser); the rest is the lab: the exam's miss rate when a pedestrian in R counts c times, for designs as drawn. At the exam's c = 1 the best share is the street's own, 1.0%, and any oversampling only costs: the case of §5. At c = 10 the best share is 9.4% and the 10% design just breaks even with the natural one (the break-even is c = 8). At c = 100 the arithmetic says half the renders, but the best design in the lab is the 25% one, the 50% design reading 51.4%: the detector cannot get below 47% on R, so weight beyond a quarter of the renders is paid for and not used. What c is worth is Lesson 11's question.

7 · Which pedestrians count in the rare slice

Every miss rate here is over the pedestrians that count, and in the rare slice counting is least obvious: a pedestrian who steps out is a sliver first. The last panel of the widget cuts the detector's miss rate by how much of the silhouette shows. Up to a quarter of it the detector misses about nine in ten (96, 92, 93 and 92% in the four bins), 66% when half shows and 36% when more than four fifths does. A rule that decides who counts is a cut on this curve, and the number it reports is the average of the curve above the cut.

RuleWho uses itR frames that countMiss, naturalMiss, 10% as drawn
at least 1 px² visiblethe program (kvis = 1)94.1%63.4%56.7%
15% of the silhouette and 6 px²the street's label, and the exam80.1%58.4%50.7%
half of the silhouettea stricter rule, to show the range50.6%44.2%35.2%

In 13% of the R frames in which any pixel shows, only a sliver under 6 px² does; the program labels it a pedestrian and the exam does not count it, so the same detector scores 5.0 points worse under the program's rule than under the exam's, and 14.2 points better under a rule that asks for half of the silhouette. The number of the slice that §6 told us to spend renders on moves by 19.2 points with the definition of who counts, more than any share of the renders moves it (7.7 at 10%). Training frames and exam agree about what is in a frame and how often; they do not yet agree about what counts.

What this lesson did not do
It did not say what a miss in R costs, so the target of §6 stays a parameter: Lesson 11 prices a close step-out by the time left to stop. The weights used the street's frequency p, which the program can compute; a project reads it off its logs, where 10,000 frames hold about 82 such cases, so p and every weight carry an error of 11% (Lesson 5's counts). The finding that more examples did not help belongs to this learner, a logistic head on fixed filters; one with more room could turn them into skill, and the lesson does not measure it. One rare case was drawn on purpose; a real project has dozens, competing for the same renders. The label rule stayed the program's: Lesson 7.

Common mistakes / failure modes

"A rare case needs more data"
Seven times the renders moved the rare slice by 1.2 points: the detector lacked weight on the case, not frames (§1, §5).
"Oversampling and weighting do the same thing"
As drawn, 25% in R lowers the slice by 10.9 points and raises the exam by 4.6; weighted, the slice stays at 59.6%. The share sets the variance, the weights set the detector (§5).
"Weights are free, they only restore the truth"
They cost effective sample size, 1,220 of 1,600 at 25% and 813 at 50%: the natural program's score with that many frames (§4, §5).
"A threshold repairs a prior shift"
Only a shift of the label, and only when frames are chosen by label alone; a shift inside the positive class moves some pictures' odds only (§3).
"Balance the classes: draw 50%"
The slice stops improving near 47%, reached by a quarter of the renders; at 50% the exam is 11.0 points worse (§5, §6).
"The exam will show whether the rare case improved"
It is 1.5% of the exam's pedestrians: a perfect detector of it moves the exam by 0.9 points (§1).

Checkpoint exercise

Try it
A generator draws 20% of its 2,000 frames from a case that is 1% of the street's frames. What are the two weights, what is the effective sample size, how much do the case's frames weigh in all, and by what factor does a learner trained without weights overrate the case's odds? Answer: the weights are 0.01/0.2 = 0.050 for the case and 0.99/0.8 = 1.2375 for the rest, and they add to 2,000. The effective sample size is 2,000² / (400 · 0.05² + 1,600 · 1.2375²) = 1,632, which is 0.816 of N, near 1 − q = 0.8. The 400 frames of the case weigh 20, the 20 frames the street would have given. Without weights the odds that a pedestrian frame is the case are multiplied by 0.2 · 0.99 / (0.01 · 0.8) = 24.75.

Where this points next

The program now draws the street's cases, draws a rare one on purpose when it must, and says with its weights what its frames stand for. The detector, the exam and the weights agree about what is in a frame and how often it happens. They do not agree about what counts: every number in the rare slice is a miss rate over the pedestrians that count, and counted by the program's rule, the exam's, or one that asks for half of the silhouette, the same detector misses 63.4%, 58.4% or 44.2% of them. A pedestrian with two pixels showing is a pedestrian to the program and not to the exam, which asks for a visible fraction. What does a label mean, and where can the program's definition differ from the exam's?

Takeaway
A program can draw a rare case on purpose, at the price of a training set that is no longer the street. The loss then says the case is common and the detector pays for it elsewhere: at a tenth of the renders the rare slice improves by 7.7 points and the exam gets 0.9 points worse, and the gain stops near 47% while the cost does not. The weights p/q restore the street in expectation, and then the detector is the natural one trained on as many frames as the effective sample size (Σw)²/Σw², about N(1 − q). The share of the renders decides the variance and the weight decides the detector; a weight other than the street's is a statement of what a miss costs, and should be made as one. The best share for a given target is proportional to the target's weight on a case times the spread of its losses. And the number the rare slice reports depends on which pedestrians count.

Interview prompts

Companion reads: Reinforcement Learning · 12 Importance sampling (the same density ratio for policies) and Computer Vision (the detector this track holds fixed).