all_lessons/Synthetic Vision Data/10 · A little realitylesson 10 / 12

A little reality

Lesson 9 ended with a number we can trust to describe the program, and a program that is wrong in ways nothing inside it can see. Real labelled frames can see them, and they are scarce: the exam this series has used since Lesson 1 holds 4,500 of them. This lesson prices a real frame by what it is spent on, in one currency, the frames of real-only training that give the same exam miss rate. Calibrating the program on 50 frames takes the detector from 80% to 56%; fine-tuning the calibrated program on the same 50 reaches 50%, which real frames alone need 158 to match, and past a few hundred frames real frames alone do better. Grading takes the largest share: ±5 points of miss rate costs 869 frames. And a miss rate per frame is not yet a probability per approach.

The thesis, here
A real frame has no fixed worth in synthetic frames, because what it buys depends on what it is spent on. Spent on calibrating, it sets a few numbers that reshape every frame the program draws afterwards; spent on fine-tuning, it is one more training frame that corrects what the program cannot; spent on grading, it is the only thing that says how far to trust either, and that is the largest bill. The exchange rate is read off learning curves, and it falls as the budget grows.
Linear position
Forced by: A factory with a manifest that replays any frame, invariants that catch 14 of 17 injected bugs and an audit of the manifest two more, and splits by scene lineage that stop a memorizing model from scoring its own training data gives a number we can trust to describe the program: the naive program passes every check and its detector still misses 80% of the street's pedestrians. The program is still a model of the street, wrong in ways no check inside it can see, and a few real labelled frames can say how wrong. How should M real frames be spent, to grade, to calibrate the program or to train the detector, and what is a real frame worth in synthetic ones?
New idea: a real frame is worth what it is turned into, measured in the frames of real-only training that give the same exam miss rate. A frame that fine-tunes the calibrated program is worth 3 of them at 50 frames, and the program's 1,600 frames are worth 62, so one real frame is worth 26 of the program's.
Forces next: A real labelled frame is worth 26 frames of the calibrated program, and 400 of them are best spent calibrating and fine-tuning on 100 and grading with the other 300: the detector now misses 49% of the street's pedestrians where it first missed 80%, and the grade certifies at most 57%. That is a rate per frame, and the shuttle meets a pedestrian once per approach: first seen at 16 m, a pedestrian leaves 8 frames before the shuttle can no longer stop, and missing all 8 has probability 49% to the eighth power, 0.3%, if misses are independent and 49% if they repeat, which no set of single frames can say. Nor can any log of human driving say what braking would have changed, because the drivers who braked early are the ones who saw the pedestrian. Can the program say what would have happened had the shuttle braked?
The plan
Six moves. (1) The ledger: what each use of a real frame outputs, and the yardstick. (2) The price of the grade, the use nothing replaces. (3) Four ways to spend M frames on a detector, measured. (4) The exchange rate, read off the curves. (5) Spending a budget: the split that certifies the most. (6) The exam's rate per frame against a probability per approach.

1 · What a real frame can buy

The exam holds 3,000 test frames and 1,500 validation frames, 4,500 real labelled frames in all. The repair of Lessons 3 to 5 used none for the camera (a bench kit and thirty logs), none for the light (a thousand unlabelled logs) and at most 200 outlined frames for the scene, so knowing that it worked cost at least 22.5 times what making it work did. A project has an M, and a few hundred is a good one. A labelled frame can be spent three ways:

A real frame is spent on…What comes outWhat it costs
gradinga number and its interval: the detector's miss rate on the street, from frames it has never seenthe bill of §2
calibratinga program: its stages' settings, read from outlines (scene), unlabelled logs (light) and a bench kit (camera)no labelled frame for camera and light, about 50 for the scene (Lesson 5)
fine-tuninga detector: the real frames join the synthetic ones in trainingeach frame, once

They are not rivals in the same way. Calibrating reads a few numbers off the frames and leaves them where they were, so a frame that calibrated can still train; a frame that trained is no longer a fair test. The real competition is between training and grading. To compare uses we need one currency, and the exam gives only miss rates, so the currency is what a frame is worth: the frames of real training that reach the same miss rate. The yardstick is the real-only learning curve R(M), the fixed detector trained on the first M frames of the street's labelled stream and graded by the exam. Every number below averages afternoons, each one a stream of labelled frames with its own logs and kit: three for every strategy, six for the yardstick, whose trainings are small. The free repairs come first. The camera read from a kit and the brightness read from 1,000 logs, with no labelled frame, take the naive program's detector from 80.3% to 62.5%, while a detector trained on 1,600 real frames misses 46.0%. The first 18 points cost nothing; the other 17 are what a labelled frame must buy, and the question is which use buys them cheapest.

2 · Grading: the use nothing replaces, and the largest bill

A miss rate is a count: of n pedestrians, k are missed, and p̂ = k/n has standard error √(p(1 − p)/n). What counts is pedestrians, not frames: only 44.1% of the street's frames hold one the exam counts (1,323 of 3,000), so a labelled frame brings 0.441 of them and 25 frames bring about 11. At 11 pedestrians and an 80% miss rate the interval p̂ ± 1.96 √(p̂(1 − p̂)/n) covers the true rate in 90% of samples, not 95%, and it collapses at a count of zero. Wilson's interval (1927) keeps every p within 1.96 standard errors of the data, a quadratic, covers 95% there, and survives zero misses (0 of 20 gives an upper end of 16%):

p ∈ [ ( p̂ + z²/2n ± z √( p̂(1 − p̂)/n + z²/4n² ) ) / ( 1 + z²/n ) ],   z = 1.96

Inverting the normal interval gives the pedestrians a half-width h needs, n = 1.96² p(1 − p)/h², and dividing by 0.441 gives frames, for a detector that misses about half (the series' detectors) and for the naive one:

Half-width of the 95% intervalPedestrians (miss rate 47%)Frames (47%)Frames (80%)
±5 points383869558
±3 points1,0642,4131,549
±2 points2,3935,4273,486

The exam's 1,323 pedestrians read a 47% miss rate to ±2.7 points, the best case a project will see. A difference of two detectors needs more care. Graded on the same pedestrians, two detectors are right and wrong about many of the same ones, and the variance of the difference of their miss rates is (b + c)/n² − (b − c)²/n³, where b pedestrians are found by the first and missed by the second and c the reverse; independent samples would add the two binomial variances. Lesson 4's detectors with the exact camera, and with the exact camera and light, differ by 4.7 points on the exam. They disagree on 19.5% of the pedestrians, so the paired standard deviation is 1.2 points and the unpaired one 1.9: the same difference is 3.9 standard deviations from zero paired and 2.5 unpaired. A 3-point improvement needs 829 pedestrians on the same ones, and 1,952 on different ones.

The threshold is a measurement too. The exam asks for the detector that lets 10% of pedestrian-free frames alarm, a quantile estimated from M₀ such frames, so the false-alarm rate it delivers has standard deviation √(0.09/M₀). At the exact-stage detector's operating point the miss rate is 61.8, 48.3 and 39.2% at 5, 10 and 15% false alarms, 2.3 points of miss rate per point of false alarms. Redrawing the threshold 500 times:

Pedestrian-free frames M₀501505001,500
sd of the false-alarm rate, points4.42.51.30.8
sd of the miss rate, points9.35.12.81.7

The exam's own threshold comes from 798 pedestrian-free validation frames and carries ±2.4 points of miss rate, more than the ±1.4 of its pedestrians; for this detector it let 10.6% of fresh pedestrian-free frames alarm, worth 1.6 points in its favour, inside that error. What the threshold needs are pedestrian-free frames, and 98% of the unlabelled logs are: the 90th percentile of the scores on 2,000 logs lets 9.1% of pedestrian-free frames alarm, 0.9 points short of 10 because the logs' own pedestrians use up part of the 10% ((0.1 − 0.02·TP)/0.98 gives 9.2%), with no labelled frame at all. The threshold costs frames that are free; the pedestrians cost frames that are not.

So grading is the largest bill and no other use can pay it: a calibrated program cannot certify itself, because its own frames are what is under test. The exam spent 4,500 frames to read ±2.7 points; a project with M = 400 reads ±7.4 at best. Everything below was graded on the lab's big sets, and §5 prices what a project would have to spend to know it.

3 · Four ways to spend M on a detector

StrategyThe detector's training setEvidence it reads
R, real onlythe M real framesM labelled frames
F, fine-tune the naive program1,600 frames of the naive program and the M real frames, each real frame weighing 1,600/MM labelled frames
C, calibrate1,600 frames of the program rebuilt from evidencea bench kit, 1,000 logs, and M labelled frames for the scene only
CF, calibrate then fine-tune1,600 frames of the calibrated program and the M real frames, weighted as in Fall of the above

The weight is decided before the lab speaks: the real frames weigh as much in total as the synthetic ones, treating the two sources as equally informative per unit of weight; §4 measures what that costs. The calibration is Lessons 3 to 5 composed. A kit of size one (11 exposures of a grey card, one edge shot, 30 exposure-log frames) gives the camera and 1,000 logs give the brightness range; the first M labelled frames give the pedestrians' distance range and step-out probability through Lesson 5's ruler. The sun, palettes, clutter, van and label rule stay the naive program's: no estimator reaches them. What the evidence reads, three afternoons (± is the standard deviation over them), against the street's own settings, which only the lab knows:

Read from the first… frames2550200800The street
nearest pedestrian, m10.1 ± 1.48.8 ± 1.08.0 ± 0.08.0 ± 0.08
farthest pedestrian, m23.2 ± 2.524.6 ± 0.024.6 ± 0.024.6 ± 0.024
step-out probability0.30 ± 0.200.27 ± 0.100.25 ± 0.020.28 ± 0.020.28

The kit reads the full well as 4,014 ± 13 electrons (the street: 4,000) and the logs read the brightness range as 0.11 to 0.95 (the street: 0.12 to 1.00), neither with a labelled frame.

The real-frame budget
Slide the number of real labelled frames M (a log scale). Top: the exam miss rate of the four strategies, three-afternoon means (six for R) with each point's spread; every point is a trained detector, precomputed by the builder. Red dashed: the naive program; ink dashed: the program with the street's own scene, light and camera. The budget is split into frames that train and calibrate (teal) and frames held out to grade (amber); the grey band is the 95% interval the held-out frames put around the best detector. Bottom: what the first min(M, 800) labelled frames say about the street's pedestrians, computed on the spot by Lesson 5's ruler, with a bootstrap.
R, real only
—
F, naive + real
—
C, calibrated
—
CF, calibrated + real
—
best equals real-only frames
—
train · grade, frames
—
certified miss rate
—
scene read: range · step-outs
—
Show the core JS
/* Wilson's score interval for k misses among n pedestrians (Wilson, 1927): the p that a sample of n puts within z standard errors of k/n; it does not collapse at 0 or n */
L.wilson = function (k, n, z) {
  z = z || L.Z95;
  var p = k / n, z2 = z * z, d = 1 + z2 / n, c = (p + z2 / (2 * n)) / d, h = z * Math.sqrt(p * (1 - p) / n + z2 / (4 * n * n)) / d;
  return [c - h, c + h];
};

/* what G held-out labelled frames can certify about a detector whose miss rate is p: the upper end of the 95% interval, from the pedestrians that count (ppf of them per frame) */
L.certify = function (p, G, ppf) {
  var n = ppf * G;
  return n < 1 ? 1 : Math.min(1, p + L.halfWidth(p, n));
};

L.bestSplit = function (curve, M, ppf, grid) {
  var best = null, i;
  for (i = 0; i < grid.length; i++) {
    var Mt = grid[i], p = curve(Mt), b = L.certify(p, M - Mt, ppf);
    if (Mt < M && (!best || b < best.bound)) best = { train: Mt, grade: M - Mt, p: p, bound: b };
  }
  return best;
};

/* the number of real-only frames at which the curve reaches miss rate m; Infinity if m is at or below the curve's floor a */
L.realEq = function (f, m) { return m <= f.a ? Infinity : Math.pow((m - f.a) / f.b, -1 / f.c); };

L.mix = function (syn, real, wReal) {
  var w = new Array(syn.length + real.length), i;
  for (i = 0; i < syn.length; i++) w[i] = 1;
  for (i = 0; i < real.length; i++) w[syn.length + i] = wReal === undefined ? L.balance(syn.length, real.length) : wReal;
  return { frames: syn.concat(real), weights: w };
};

L.framesBefore = function (z0, v, a, tau, dt) {
  var d = z0 - L.stopDistance(v, a, tau);
  return d < 0 ? 0 : Math.floor(d / (v * dt) + 1e-9) + 1;
};

L.noDetectBeta = function (p, k, rho) {
  if (rho <= 0) return Math.pow(p, k);
  if (rho >= 1) return p;
  var a = p * (1 - rho) / rho, b = (1 - p) * (1 - rho) / rho, v = 1, i;
  for (i = 0; i < k; i++) v *= (a + i) / (a + b + i);
  return v;
};

What to try. Slide to 25 frames: the same frames feed all four boxes, and R misses 68% against CF's 55%. Slide to 100 and compare none with the best split (50 · 50): the certificate box fills, at most 71% for a detector that misses 50%. At 400 the best split is 100 · 300 and certifies at most 57.3% (±8.5); half certifies 58.3% and a quarter 62.7%, because the interval widens faster than the miss rate falls. Tick the street's settings at 25 and then at 200 frames and watch the teal range settle on the red ticks. At 1,600, R (46%) has dropped below CF (49%).

4 · The exchange rate

Mean exam miss rate, %, over the afternoons, every entry a trained detector:

Frames M025501002004008001,600
R, real only—68.159.151.448.847.445.946.0
F, naive + real—60.660.455.654.354.154.054.1
C, calibrated62.559.755.655.955.556.056.056.1
CF, calibrated + real—55.049.948.847.948.849.248.6

C is flat from 50 frames on: three numbers are all the labelled frames set, and 50 frames set them. It stops 3.7 points short of the program with the street's own camera, light and pedestrian settings (Lesson 5: 52.2%), because the light's sun and palettes and the clutter come from no estimator. R catches C at 68 frames and CF at 264. CF misses 14.8, 9.9 and 2.5 points less than R at 25, 50 and 100 frames (in all three afternoons), 1.0 less at 200, and 1.5, 3.0 and 3.0 points more at 400, 800 and 1,600: half of its loss is the program's frames, which hold it near 49%, while real frames alone keep improving. F beats R only at 25 frames and is 8 points behind at 1,600. Inside the mixture the calibration is worth 6, 10, 7 and 6 points at 25, 50, 100 and 200 frames. The afternoons differ (CF at 25 frames ranges over 15 points), so gaps under about 3 points are inside the spread.

For the conversion, fit the real-only curve R(M) = a + b M−c: a floor a = 44.8%, c = 0.83, residual 3.0 points, the afternoon spread. The real-only size that matches a miss rate m is E(m) = ((m − a)/b)−1/c; a synthetic frame is worth E/N real ones and a real frame spent on a use E/M:

What the frames areReal frames spentExam missReal-only frames that match itPer frame spent
the calibrated program, scene from 50 frames (C)5055.6%641.3
calibrate, then fine-tune on the same 50 (CF)5049.9%1583.2
… on 10010048.8%2172.2
… on 20020047.9%2891.4
… on 40040048.8%2150.5

The calibrated program's 1,600 frames are worth 62 real frames, 0.039 each, so one real frame is worth 26 of them; the naive program's are worth 0.009 each (15 real frames in all) and the free evidence alone 35. The ceiling is the street's own scene, light and camera, which no project can build: in Lesson 2's table its 1,600 frames equal 1,600 real ones (0.3 points apart), worth 1 each. A frame that fine-tunes the calibrated program returns 3 real frames up to 100 frames, 1.4 at 200 and under one at 400, where CF falls behind R. The rate falls with the budget because the program's contribution is a head start, worth at most about 289 real frames, and a head start matters less the farther the real frames have run.

The weight is the exchange rate in disguise: a real frame of weight w counts as w synthetic ones in the loss. At 200 frames (one afternoon) F misses 62.8, 57.4, 54.9, 52.1 and 48.8% at w = 1, 4, 8 (the rule), 32 and 128, a spread of 14 points; CF misses 49.5, 50.2, 49.7, 49.8 and 48.7%, a spread of 1.5. A naive frame is worth 0.009 real ones, about a hundred to one, and F is best at the swept weight nearest that, 128. The calibrated program is right at any weight and the naive one only at the weight its worth implies: what calibration buys inside the mixture is a weight nobody had to tune.

Road not taken · learn a refiner
A project with logs and no physical model has a tempting road: train a network to turn the program's pictures into pictures that look like the logs, using no street labels, then train the detector on the refined frames (Shrivastava et al., 2017, for simulated eyes and hands; Hoffman et al., 2018, which adds cycle consistency and a task loss). It asks for no camera model, no light model and no outlines. What breaks is what the program's frames were for. Matching the logs is a statement about whole pictures, and a pedestrian is in 2% of the logs and 50% of the program's frames, so a refiner that matched them exactly would remove the pedestrian from 96% of the program's pedestrian frames. What stops that is a term that keeps the refined picture near the program's: a prior that encourages the label to survive, not a check that it did, and no unlabelled frame measures how many pedestrians survived. Grading a refiner takes labelled frames, the bill of §2, and one that fits the logs' pixel statistics is the calibration of Lessons 3 and 4 with its parameters unnamed and so unchecked. This lesson did not train one.

5 · Spending a budget

A frame can calibrate and train at once and cannot train and grade, so a budget M splits into Mt frames that calibrate and fine-tune and G = M − Mt held out. The detector trained on Mt misses p(Mt), the best of the four curves, and the held-out frames hold 0.441G pedestrians, so the strictest claim they support, the upper end of the 95% interval (planned with the normal interval), is

bound(Mt) = p(Mt) + 1.96 √( p(1 − p) / (0.441 (M − Mt)) )

Training more lowers the first term and widens the second; the best split is where a frame moved from grading to training lowers the miss rate exactly as much as it widens the interval. Over the measured training budgets:

Budget MTrain (calibrate and fine-tune)GradeDetector missesCertified: at mostPrice of certainty
2005015049.9%62.0%12.0
40010030048.8%57.3%8.5
80020060047.9%53.9%6.0
1,60080080045.9%51.1%5.2
4,500 (the exam)8003,70045.9%48.3%2.4

The rule has three parts, each with its break-even. Calibrate first: it costs no frame it cannot give back, and below 68 frames it beats training outright. Fine-tune on what the curve still falls for: with the calibrated program the curve is flat from about 100 frames (48.8, 47.9, 48.8%), and past 264 frames real frames alone beat the mixture. Grade with the rest: from 400 frames on, more than half the budget goes to the grade, because a trained frame now buys under a point and a held-out one narrows the interval. The price of certainty, how far the certified bound sits above the miss rate, is 12 points at 200 frames and 2.4 at the exam's 4,500: a project cannot show what it has built. Turned around, certifying a miss rate of at most 55% takes 200 frames of training and 434 of grade, 634 in all, and at most 52% takes 1,378.

6 · A rate per frame is not a probability per approach

Every price above is in the exam's currency, a miss rate per frame, and the detector the best split trains at 400 frames misses 49% where the naive one missed 80%. The shuttle does not live per frame. Take the next lesson's shuttle, the lab's assumptions and not facts about any vehicle: it drives at 7 m/s (25 km/h), decelerates at 3 m/s² once it has decided, and takes 0.4 s from decision to deceleration, so it stops after dstop = vτ + v²/2a = 2.8 + 8.2 = 11.0 m. A pedestrian first seen at depth z₀ is safe only if the detector fires while the shuttle is still farther away than dstop. The camera records ten frames a second, the shuttle moves 0.7 m between frames, and the frames that can still help number ⌊(z₀ − dstop)/0.7⌋ + 1. If the detector misses a frame with probability p = 49%:

Pedestrian first seen at12 m14 m16 m20 m
frames before the last useful moment, k25813
no detection, misses independent: pk23.8%2.8%0.3%0.01%
no detection, pedestrians differ (correlation 0.5 between a pedestrian's frames)36%24%18%15%
no detection, misses repeat in every frame: p49%49%49%49%

The middle row gives every pedestrian their own miss probability (some are easy, some hard, the same in every frame of one approach) with mean p and correlation 0.5 between two frames of the same pedestrian: the chance that all k frames miss is E[qk] for q beta-distributed, a product of k terms. At 16 m the rows differ by a factor of 153, and each describes some detector with a 49% miss rate per frame. The exam cannot say which describes this one: its test set holds single frames, and the correlation is a property of pairs of frames of one approach. This detector misses partly pedestrians who are hard to see and stay hard as the shuttle nears (misses that repeat) and partly noise (misses that do not), and nothing in a set of single frames weighs the two. The number the shuttle needs exists only in the program's closed loop.

What this lesson did not do
It measured one detector, a logistic head on 110 fixed features that learns the street from about a hundred frames; with a learner of millions of weights the real-only curve is slower, the same program is worth more, and the exchange rate moves. The weight of a real frame was one rule and a sweep at three budgets in one afternoon, where a project would choose it by cross-validation on its own frames. Three afternoons leave gaps under about 3 points unresolved. It did not price a miss (Lesson 11), say what counts as a pedestrian in real labels (Lesson 7) or train a refiner. Costs are in frames, since a frame's price in money depends on a labelling specification, and the grading arithmetic treats frames as independent, which frames of one drive are not (Lesson 9's lineage splits are the repair).

Common mistakes / failure modes

"A real frame is worth one real frame"
Fine-tuning the calibrated program it is worth 3 at 50 frames and under 1 at 400; the naive program's frames are worth 0.009 each (§4).
"Mix real and synthetic frames at equal weight"
With the naive program that is 54.9% at 200 frames against 48.8% at a weight of 128; the calibrated program is right at any weight (§4).
"More labelled frames always improve the calibration"
C is flat from 50 frames: 55.6% there, 56.1% at 1,600 (§4).
"A 3-point gain is a gain"
Graded on different frames it needs 1,952 pedestrians; on the same ones 829 (§2).
"10% false alarms means 10%"
A threshold from 50 pedestrian-free frames moves the miss rate by ±9 points (§2).
"The simulator pays at every budget"
CF misses 9.9 points less than R at 50 frames and 3.0 more at 1,600 (§4).
"47% per frame is 47% per approach"
With 8 frames to spare it is 0.3% if misses are independent and 49% if they repeat (§6).

Checkpoint exercise

Try it
A project holds 600 labelled frames. (a) Planning for a miss rate of 50%, the widest case, how many pedestrians must the held-out frames hold to read it to ±8 points, and how many frames is that? (b) How many frames are left to calibrate and fine-tune on, and what does the certificate then read if the detector misses what the best curve says? Answer: (a) 1.96² × 0.25/0.08² = 150.1, so 151 pedestrians, and 151/0.441 = 342.4, so 343 frames. (b) 257 frames are left; the best curve misses about 48% there, so the certificate reads at most 56%.

Where this points next

A real frame now has a price in every currency the project uses: 26 frames of the calibrated program, about 3 frames of real training when it fine-tunes that program, and nothing else when it grades, which at 400 frames certifies 57% for a detector that misses 49%. But that is a rate per frame, and the shuttle meets a pedestrian once per approach: first seen at 16 m, a pedestrian leaves 8 frames, and missing all of them has probability 0.3% if misses are independent and 49% if they repeat, which a set of single frames cannot say. And braking is an action whose consequences no log of human driving contains, because the drivers who braked early are the ones who saw the pedestrian. Can the program say what would have happened had the shuttle braked?

Takeaway
A real frame is worth what it is turned into, measured in the frames of real-only training that give the same exam miss rate. The camera and the light cost no labelled frame and take the detector from 80% to 62%; 50 frames set the scene and take it to 56%, and the program's 1,600 frames are then worth 62 real ones, so a real frame is worth 26 of them. Fine-tuning the calibrated program on the same 50 reaches 50%, 3 frames of real training per frame, and the return falls with the budget until real frames alone do better, near 264 frames. A naive program's frames are worth 0.009 each and need a tuned weight; a calibrated program's do not. Grading is the largest bill: ±5 points costs 869 frames, and a 3-point gain needs 829 pedestrians even paired. The best split trains on a few hundred frames and grades with the rest, and what 400 frames certify sits 8.5 points above what they build. All of it is a rate per frame, and the shuttle needs a probability per approach.

Interview prompts

Companion reads: Training a Robot Model · 16 Exchange rates (the same unit for hours of robot data), Training a Robot Model · 15 Evaluation (what grading costs there) and Computer Vision · 19 Evaluation, deployment, and CV system design.