all_lessons/Synthetic Vision Data/04 · Same scene, other pixelslesson 4 / 12

Same scene, other pixels

Lesson 3 gave the program the street's camera and took the miss rate from 80.3% to 65.8%, but it still lights every frame with one sun at one brightness, while the street shows the same scenes at dusk, in glare and under cloud, and the label does not care. Which of the light's ten parameters must the detector see vary, and how far? Here one does, the brightness, because the camera turns light into noise; its range can be read from the sky of unlabelled frames. Widening the whole look past the street's range costs the program its own labels.

The thesis, here
A nuisance is a factor the label does not depend on, and a learner becomes invariant to it only by being shown it vary. Which nuisances to show is measured on the program's own free labels; how far, by the range the world has, read from pixels of known content, and no wider: past a point the pedestrian leaves the pixels and the empty frames raise the alarm threshold for everyone.
Linear position
Forced by: Flat frames of a grey card gave the camera's full well and read noise, an edge gave its blur, and logs gave its exposure control; added to the program, they take the miss rate from 80% to 66%, within a point of what the street's own camera numbers give, and the program's pictures now carry the grain and softness of the real camera. A calibrated camera still sees only the light the program gives it: one sun behind the camera, white vans, bright shirts, a blue sky. The program's frames leave its exposure controller at a gain near 2, the street's run from 1.9 to the cap of 8 with a third of them at the cap, and the controller that is worth 1 point in the program's light is worth 7 in the street's. The real street shows the same scenes at dusk, in glare and under cloud. What should be allowed to vary between two pictures of the same scene, and what must not?
New idea: a nuisance is a factor the label ignores, and invariance to it is bought by showing the learner the range the world has, no wider. Only the brightness moves this exam; its range, read from the sky, takes the miss rate from 65.8% to 61.2%.
Forces next: Randomising the light over the range read from the sky of unlabelled frames takes the miss rate from 66% to 61%, within a point of the 62% the street's exact light gives, and randomising wider than that buys no accuracy while the program's own error climbs from 7% to 17%. With the pixels right, what the program draws starts to matter in a way it did not: a scene repair that was worth a little over 5 points while the pictures were wrong is worth 15 now. And the scene is where the program is silent about whole kinds of case: every pedestrian it drew stood in the open and within 16 m. Which real cases does the scene never draw, and how do we count them?
The plan
Six moves: (1) one scene under many lights; (2) which lights the detector tells apart; (3) where a nuisance stops being one; (4) the street's range from unlabelled frames; (5) how wide; (6) the price against the exact program, and what is left.

1 · One scene, many lights

The program of Lesson 3 has the street's camera and the light of SIM0: a sun 55° high behind the camera, one brightness, a white van, bright shirts, a blue sky. The gap shows in the camera's exposure control: over 1,000 street logs its gain runs from 1.9 to its cap of 8 and sits at the cap in 34% of them; over the program's frames (the 1,000 that Lesson 3 sent through this camera) it runs from 1.8 to 2.3, so the detector has never met a frame taken at the cap.

Group (parameters)The programThe street
sun (elevation, bearing)55°, behind the camera12°–65°, within ±150°
brightness (level, ambient)1, 0.450.12 to 1 (8.3-fold), 0.2–0.5
colours (clothes, van, sky, road)bright, white, blue, 0.3any, any, 3 skies, 0.15–0.45
textures (windows, stripes)none, 0.20.3, 0.05–0.35

None of the ten enters the label, a function of the scene and the pinhole: which pixels the pedestrian covers, how much of the silhouette shows, whether the pedestrian counts. Draw 8 scenes under 8 looks each: the masks of one scene agree to the last bit (largest difference 0), and so do labels and areas, while the pictures differ. All ten are nuisances, factors the label does not depend on, the setting of augmentation theory, which averages over transformations that keep the data distribution, and with it the label relation, approximately invariant (Chen, Dobriban and Lee, 2020).

A learner cannot be told which factors are nuisances. It fits whatever separates its training pictures, and if the light never changes, how this pedestrian looks in this light separates them as well as what a pedestrian is. Only a scene shown in other light says that the factor moves and the label does not, and every factor shown costs frames: which of the ten does the detector need?

2 · Which lights the detector can tell apart

Take four looks from the street's range, each a single point: noon, the program's own (sun 55° high behind the camera, ambient 0.45, brightness 1); low (15° high, 90° to the side, ambient 0.30, brightness 0.7); against (30° high, 150° round, nearly facing the camera, ambient 0.30, brightness 0.8); dim (noon's geometry at brightness 0.15). Train the fixed detector on 800 frames of one look and grade it on the program's frames of each, at the threshold that lets 10% of that look's empty frames alarm. Cells are miss rates in per cent; no street, no real label.

trained on \ graded onnoonlowagainstdim
noon0.00.80.425.4
low0.00.30.022.1
against0.10.10.114.7
dim2.74.80.63.9
200 of each0.00.10.05.9

Read it by columns. Among the three bright looks no off-diagonal cell exceeds 0.8%, though the sun moved from behind the camera to the side and into the lens: to this detector they are one look. The dim look is another world: detectors trained on the bright looks miss 14.7 to 25.4% of its pedestrians, the dim-trained one at most 4.8% anywhere else, so transfer runs from hard light to easy and not back. Training on 200 frames of each look mends the dim column to 5.9% but costs 2.0 points against the specialist: the first sight of the price of invariance.

Why brightness and not the sun? The detector reads bars and steps in both polarities, so a pedestrian against a road is an edge whichever side is lit; a sun in another direction only moves light and shade about. Brightness changes what the filters cannot read around, the noise. A pixel of radiance L collects ev·L electrons and its relative noise is √(ev·L + r2) / (ev·L), with ev = 600 electrons per unit radiance and read noise r = 3 electrons (Lesson 3): 7.6% at L = 0.30 and 22.2% at 0.045, the same road at brightness 0.15, 2.9 times as much, and auto-exposure lifts the picture back to the same average level with the noise in it. Light reaches the picture by two doors, pattern, which this detector ignores, and noise, which it does not.

3 · Where a nuisance stops being one

Why not show every nuisance as widely as it goes? Because at some brightness the pedestrian is no longer in the picture. The pedestrian's signal is the picture with the pedestrian minus the same picture without, as the camera records it, blur included: si at each pixel and channel. The camera adds independent noise of variance σi2 = (ev·Li + r2) / ev2 there, Li the radiance behind the pixel. The signal's length in units of noise is the contrast-to-noise ratio:

CNR = √( Σi si2 / σi2 )

It is the detectability d′ of an observer who knows the signal exactly, and under this noise model no detector that sees only these pixels does better: a matched filter run on 4 scenes of the lab's own noisy frames, at brightnesses from 0.02 to 0.5, measures a d′ within about 5% of it. A frame has K = 576 cells that can alarm and one empty frame in ten may, so each cell may alarm with probability 1 − 0.91/K: a threshold of z = 3.56 noise sd, confirmed by simulating the largest of 576 normals. The ideal observer finds a pedestrian of CNR c with probability Φ(c − z): half at 3.56, nearly all at 8.

The detector is not ideal. Grade the detector of the street's own look on the program's frames, drawn in the street's look (open pedestrians 8 to 16 m away), and bin its pedestrians by CNR. This is a different population and a different detector from Lesson 3's line of about 14, which was the street-trained detector on the street's own test frames, where some pedestrians step out from behind the van; the two lines need not agree:

CNR of the pedestrianbelow 88–1212–2020–35above 35
pedestrians41984226381
found75%68%79%93%99%

A logistic fit of hits against log CNR crosses one half at c0 = 8.5 (90% of 200 bootstrap fits lie between 6.5 and 10.4; only 23 pedestrians are under 12, so the line is a soft gauge). Widen every numeric range of the street's look 1.5 times about its centre, brightness down to 0.01: at the street's own width 0.8% of the program's pedestrians fall below the line, at 1.5 times 6.2%, positives with little evidence behind them that free labels cannot reveal. The empty dim frames cost as much: 17% of the frames are dim (brightness under 0.2), and at the threshold pooled over all frames an empty one alarms 25% of the time against 7% for a bright one, so the threshold sits too high for the bright frames and their miss rate rises from 7.7% to 13.6%. So a width has a limit the program can measure with no street: the share of pedestrians below the line, and its own miss rate, the exam run on frames it draws.

4 · The street's range, from pixels whose content is known

The street's range of brightness is in no label, but it is in pixels whose content is known without one. The top row is sky (the far wall ends below it): the brightness times the sky's colour. The bottom rows are road: the brightness times the ground albedo times the road's light, ambient + (1 − ambient) · sin(elevation). Read each band through the calibrated camera, DNγ / (κ · gain) with the logged gain, which turns the picture back into scene radiance, and take the band's median, so that a tree, a van or a pole over part of it moves nothing.

The sky tells two things. Its colour, the channel ratios with the level divided out, names which of three skies the frame shows (blue, grey, warm; the program is the calibration standard of its own probes and supplies their colours): right in 99.9% of the 96% of frames it classifies. Its level, divided by the level the program's probe reports for that sky at brightness 1, is the frame's brightness: against the hidden truth (a lab privilege) it is 0.95 of it, middle 80% from 0.90 to 0.99, a little low because the blur leaks the darker wall into the top row. If the brightness is uniform on [lo, hi], its p-quantile is lo + p(hi − lo), so the 5% and 95% quantiles give the whole range and a few outliers move neither:

frames of evidence n101001,000
estimated range0.24–0.840.13–0.940.12–0.94
sd of its upper end0.080.03—

The street's own range, which no project could read, is 0.12 to 1.00. The sd shrinks like 1/√n; the offset of the upper end does not, for it is the probe's low reading. A hundred logs are enough.

The road tells less. Dividing its level by the brightness leaves G = ground albedo × road light, a product of three things it cannot separate: a road of albedo 0.30 under a 55° sun and ambient light 0.45 and one of albedo 0.45 under a 30° sun and ambient light 0.20 have the same G, 0.27 and 0.27. The probe reads them alike (ratio 1.00) and the wall behind barely tells them apart (1.17). The probes constrain these parameters jointly and identify only the brightness and the sky; the sun's bearing and the clothes have no probe. The principle is to match the distribution of what you can observe, only through a cause you can attribute: fitting a simulator's parameter distribution to real observations is the idea of Chebotar et al. (2019, who matched real-robot trajectories), here with pixel statistics as the observation.

The width of the brightness range
Top: eight frames, one scene re-lit by looks drawn at the chosen width of the estimated range (green: CNR above the line, amber: below; "missed": that width's detector scores the frame under its own threshold), or the first eight street logs (red where it alarms). Bottom: the sweep, and the sky brightness of the first n logs with the estimated (blue) and drawn (teal) ranges.
exam miss (brightness)
—
own miss (brightness)
—
alarm rate, 1,000 logs
—
whole look: exam miss
—
whole look: own miss
—
top row
—
range from n logs
—
Show the core JS
L.uniformRange = function (v, a, b) {
  var qa = quantile(v, a), qb = quantile(v, b), span = (qb - qa) / (b - a);
  return [qa - a * span, qb + (1 - b) * span];
};
L.read = function (frame, tab) {
  var p = L.probes(frame), k = L.classify(p, tab), l;
  if (k < 0) return null;
  l = p.level / tab[k].level;
  return { sky: k, lum: l, G: p.ground / l };
};

  var db = L.blur(d, cam.blur), rb = L.blur(rest, cam.blur);
  for (i = 0; i < d.length; i++) { v = (cam.ev * Math.max(rb[i], 0) + cam.read * cam.read) / (cam.ev * cam.ev); s2 += db[i] * db[i] / v; }
  return Math.sqrt(s2);
L.zAlarm = function (K, fa) { return L.PhiInv(Math.pow(1 - (fa === undefined ? 0.1 : fa), 1 / K)); };

What to try. Open as set: width 1, 100 logs, the 3rd scene (a pedestrian 11 m away, CNR 42). That width's detector has an exam miss rate of 61.5% and an own miss rate of 0.4%; 0 of the 8 frames fall below the line, 0 missed. Slide the width to 0: the eight frames share one brightness, 0.53, yet the exam is 62.0% against 61.5%: the level mattered, not the width. Go to 1.5 and 3: the range reaches 0.01 and 1.8, the exam is 60.7% and 60.7% and the own miss rate at or below 2.5%, while the dashed curves, the whole look widened, climb to an own miss rate of 16.7% at 1.5. Switch the top row to the logs: width 0 alarms on 5 of the first eight frames, width 1.5 on 2, and over 1,000 logs the alarm rate falls from 74.2% to 30.6% while the exam does not move. The farthest scene (15 m) has CNR 20 against 42 at 11 m: distance costs CNR as dim light does (§6).

5 · How wide? The sweeps

Two sweeps scale the width s of a range about its centre: one for the brightness alone (the street's 0.12 to 1, kept above 0.01), the other for every numeric range of the street's look together (sun elevation and bearing, ambient fraction, brightness, ground albedo, stripes) with the street's clothes, vans and skies, a lab privilege. "Line" is the share of pedestrians below c0. Cells are per cent, means of three seeds at 1,600 frames; the exam wobbles by about 1.5 points between seeds.

width sbrightness alonethe street's whole look
examownexamownline
062.00.066.54.90.0
0.562.00.260.95.90.0
161.50.461.66.70.8
1.560.72.563.416.76.2
260.91.763.617.34.7
360.70.662.912.82.3

Brightness alone: level before width. The exam is flat in the width (the six values lie within 1.3 points) and the own miss rate stays at or below 2.5%: at 1,600 frames the width of the brightness range buys nothing. What helps is the level: the program's one brightness, 1, sat at the top of the street's range, and moving it to the centre, with no variation, takes the exam from 65.8% to 62.0% (3.8 points), the whole range then adding 0.5 more. The matrix says why: trained at 1 the detector misses 25.4% at 0.15, and the street's frames lie anywhere in the range, nearly all dimmer than 1; a brightness at the centre is at most 4.7 times from any of them, one at the top 8.3 times from the dimmest.

The whole look: a shallow U and a cliff. Widened together the look gives 60.9% at width 0.5, 61.6% at the street's own width and 63.4 to 63.6% beyond it, while the program's own miss rate, which needs no street, stays at or below 6.7% to width 1 and jumps to 16.7% at 1.5, as the share below the line goes from 0.8% to 6.2% and the dim frames lift the threshold (§3). That is a stopping rule that needs no real label: widen until the program's own error jumps. The alarm rate on 1,000 unlabelled frames, a meter, does not give it: it falls from 74.2% to 30.6% as the brightness alone widens to 1.5 and from 25.5% to 17.8% as the whole look widens from 1 to 2, while the exam stops improving. It says how alien the frames are, not whether more variety helps the labels.

The price of invariance. Whole look, width 3 against width 1: exam 69.2 against 66.2% at 200 frames, 64.6 against 64.1% at 400, 62.6 against 62.7% at 800 and 62.9 against 61.6% at 1,600 (one seed below 1,600). The exam gap is a few points at 200 frames and within the seed wobble beyond; the wide program's own miss rate is higher at every size, 18.2 against 12.3% at 200 and 12.8 against 6.7% at 1,600.

Road not taken · randomise as widely as it goes
Tobin et al. (2017) took this road. They randomised lights, textures, distractors, the camera and the noise widely and trained on rendered images alone (no real image of the task; most runs started from an ImageNet-initialised network); objects in real webcam images were localised to 1.5 cm on average, poorly with fewer than a thousand distinct textures at ten thousand images and better up to about fifty thousand. A wide range covers a world one has not measured and pays when frames are plentiful. Here, at 1,600 frames, the whole look three times wider buys nothing: the exam is 1.3 points above the street's width, inside the seed wobble, and the own miss rate 12.8%; once the probes have measured the range, width beyond it is data spent on frames no street produces.

6 · The price, and what the repaired pixels leave

The program built from the evidence draws the brightness from [0.12, 0.94], read from the sky of 956 usable frames of 1,000 unlabelled logs; nothing else changes. Trained on 1,600 frames and graded with the Lesson 1 exam (three seeds), it misses 61.2% of the street's pedestrians, against 65.8% for the program before the repair (aaba) and 61.6% for the exact-stage look (abba): 0.4 points from the exact one (a lab comparison: a project cannot build abba), inside the 1.5-point wobble between seeds.

Road not taken · match every probe
The probes also see the road, and a program that matches the logs on every probe sounds better than one that matches the sky. Carry the road's light in the ground albedo alone and add the sky's three colours: the road's statistic matches by construction and the exam is 70.6% against 61.2% for the brightness alone (9.4 points worse, two seeds), though the meter prefers it (44.0% alarms against 47.2%). The road cannot say which of its three causes varied (§4), and the wrong one draws roads, under high suns and in dark ambient light, that the street never has. Match what the evidence attributes.

What the repaired pixels leave is the scene. With the light right, repairing it is worth more: the scene stage was worth 5.5 points while the pictures were wrong (baaa against aaaa) and is worth 14.8 now (bbba against abba). And the scene is where the program is silent. It draws every pedestrian 8 to 16 m from the camera, in the open, none stepping out from behind the van: of 600 program pedestrians that count, the nearest is at 8 m, the farthest at 16 m, 0% step out. The street's are weaker: the median contrast-to-noise ratio of its exam pedestrians is 22 against 38 for the program's, 1.7 times as strong, and the 5% quantile 7 against 17, below the line 8.5 for the street and above it for the program. A median hides which cases are missing and how many, and a project has no labels of the street's pedestrians to find out.

What this lesson did not do
It repaired the light for one learner and one exam: this detector reads around the sun's direction and not around the noise, and a network with other filters would give another matrix (the test carries over, not the verdict). The range came from a sky and a road; the sun's bearing, the clothes and the stripes have no probe in unlabelled frames, and the estimate assumes a uniform spread. The three sky colours the classifier tells apart came from the program's own look code, a catalogue a real project would first have to learn by clustering the skies of its logs. It left the scene alone (Lessons 5 and 6), the camera (Lesson 3) and what counts as a pedestrian (Lesson 7), and it never used a real label (Lesson 10).

Common mistakes / failure modes

"If the label does not depend on a factor, randomise it"
Only what the learner is not already invariant to: the three bright looks are one look to this detector (no transfer misses more than 0.8%), and the whole look's other nuisances add frames, not accuracy (§2, §5).
"Wider is safer"
Not once nuisances widen together: the whole look at width 2 has an exam of 63.6% against 60.9% at 0.5, and its own miss rate leaves 6.7% for 16.7% at 1.5 (§5).
"A pedestrian in the label is a pedestrian in the pixels"
In dim light the noise swallows the pedestrian: at width 1.5, 6.2% of the pedestrians fall below the line and bright ones are missed 13.6% of the time, not 7.7% (§3).
"Match every probe and the program matches the street"
The road's statistic can be matched through the wrong cause: the exam is 70.6% against 61.2%, though the meter prefers it (§4, §6).

Checkpoint exercise

Try it
A pedestrian covers 12 px² and is 0.03 brighter than the road in each channel; the road's radiance is 0.08, ev = 600, r = 3 (blur ignored). What is the pedestrian's CNR, which side of the line c0 are they on, and where are they when the light falls to a quarter? Answer: a road pixel's noise sd is √(600 · 0.08 + 9) / 600 = 0.0126 and the signal is 36 samples (12 px × 3 channels) of 0.03, so the CNR is √36 · 0.03 / 0.0126 = 14.3, above the line (8.5). At a quarter of the light the signal is 0.0075 on a road of 0.02, the noise sd 0.0076 and the CNR 5.9: below the line, though the ideal observer would still find the pedestrian with probability 99%. A fourfold fall of light cost a factor 2.4, not four: the noise fell too (§3).

Where this points next

The light is repaired as far as this exam can feel it: a program shown the street's range of brightness, read from the sky of unlabelled frames, misses 61.2% against 61.6% for the exact-stage look and 65.8% before. With the pixels right, a scene repair worth 5.5 points while the pictures were wrong is worth 14.8 now, and the scene is where the program is silent about whole kinds of case: every pedestrian it draws stands in the open, 8 to 16 m from the camera, 1.7 times as strong as the street's median pedestrian. Which real cases does the scene never draw, and how do we count them?

Takeaway
A nuisance is a factor the label does not depend on, and all ten parameters of the light are. A learner becomes invariant to one only by being shown it vary; which are worth showing is measured on free labels, by training under one look and grading under another, and for this detector the brightness matters and the sun's direction does not, because the camera turns brightness into noise. A nuisance may be shown only as widely as the pedestrian stays in the pixels, and its range comes from pixels of known content: the sky gives the brightness, the road only a product of three causes. The evidence-built program misses 61.2%, within 0.4 points of the exact-stage look, and its gain is mostly level, not width.

Interview prompts

Companion reads: Computer Graphics, from first principles (the terms of a look) and Computer Vision · 07 Training CV models (augmentation).