all_lessons/Synthetic Vision Data/03 · The camera is a measurementlesson 3 / 12

The camera is a measurement

Lesson 2 priced the camera at sixteen points of miss rate, the dearest stage, and named the evidence that measures it without a label: a grey card, an edge, the exposure log. This lesson takes the measurements. A pixel is not the light that fell on it but a count of electrons, blurred by a lens, scaled by an exposure controller, clipped, bent by a gamma curve and rounded to a code; each step has a few numbers, and each number has evidence that varies it alone. Seven numbers read off eleven exposures of a card, one printed edge and a log become a sensor stage, and the program drawn through it misses 66% of the street's pedestrians instead of 80%, within a point of the program built with the street's own camera. It cannot say what light the camera should be shown.

The thesis, here
A camera is a channel from light to numbers, and a channel is measured by what it does to inputs whose relations we know. A card whose brightness doubles from frame to frame shows the response curve and the noise law; a printed edge shows the blur; the camera's own log shows its controller. The count of photons supplies the absolute scale, because the variance of a count is its mean, and what comes out is a program: a sensor stage that any radiance image can be sent through, whose worth is the exam's verdict.
Linear position
Forced by: Swapping the stages one at a time prices them: of the 33 points of miss rate between a detector trained on the program (80%) and one trained on the street itself (47%), the camera accounts for 16, the light and the scene for about 8 each, and the label for nothing, and the repairs compound, so the camera is worth 14.5 points while the other stages are wrong and 21 once they are right. The camera is therefore the place to start, because it is both the most expensive gap and the cheapest to measure: a few frames of a grey card pin down its noise without a single label. How does a camera turn light into numbers, and how do we measure it well enough to simulate it?
New idea: a pixel is a measurement: a signal plus noise whose law has a few numbers, and each number can be read off evidence that varies one thing while the rest holds still. Seven numbers from flat frames, an edge and a log rebuild the street's camera to within the error bars of §3 and §4, and the program drawn through them lands within 0.1 points of the program built with the street's true numbers.
Forces next: Flat frames of a grey card gave the camera's full well and read noise, an edge gave its blur, and logs gave its exposure control; added to the program, they take the miss rate from 80% to 66%, within a point of what the street's own camera numbers give, and the program's pictures now carry the grain and softness of the real camera. A calibrated camera still sees only the light the program gives it: one sun behind the camera, white vans, bright shirts, a blue sky. The program's frames leave its exposure controller at a gain near 2, the street's run from 1.9 to the cap of 8 with a third of them at the cap, and the controller that is worth 1 point in the program's light is worth 7 in the street's. The real street shows the same scenes at dusk, in glare and under cloud. What should be allowed to vary between two pictures of the same scene, and what must not?
The plan
Seven moves. (1) Say what a pixel is and which numbers a measurement of the camera must return. (2) Read the gamma curve and the linear scale off a ladder of flat frames. (3) Find the noise law by elimination, and the full well and read noise on its line. (4) Read the blur off an edge and the controller off the log. (5) Assemble the sensor stage, price it, and price its parts. (6) Ask which pedestrians the noise erases. (7) Show what the calibrated camera still gets wrong.

1 · The ideal sensor, and the channel it leaves out

Lesson 2 found the camera the dearest stage. The detector trained on the program misses 80.3% of the street's pedestrians; trained on the program with the street's camera and nothing else changed (aaba) it misses 65.8%, while the alarm rate on unlabelled street frames moves only from 100% to 93.7%. The program's sensor is ideal: it multiplies the renderer's radiance by a fixed exposure, bends it with the gamma curve and rounds it to eight bits, so a pixel is a function of its radiance. The street's camera is not: send one radiance through it twice and the two pictures differ. The widget of §5 does exactly this; at noon the street's picture is already grainy and soft, and a few stops dimmer it is mostly grain.

What the program needs from a camera is therefore not a number per pixel but a distribution: given a radiance image, the probability of each code image. The street's camera builds it in six steps, and each adds a few numbers to find.

StepWhat happens to the light of a pixelNumber to find
Lenseach point is spread over its neighbours: a Gaussian blur of width σ pixelsσ
Countradiance L becomes e = ev·L electrons on average; the number actually counted is randomev
Readthe amplifier adds noise of r electrons rms, whatever the signalr
Gainthe exposure controller multiplies by a gain g between 1 and maxGain, chosen from the frame's own brightness to bring its mean to a targettarget, maxGain
Clipthe converter saturates at fw electrons (at gain 1): q = g·e/fw is cut to the range 0 to 1fw
Codethe gamma curve DN = q1/γ, then rounding to eight bitsγ

The count and the clip meet in κ = ev/fw, the fraction of full scale that one unit of radiance fills at gain 1, while fw alone sets the size of the noise (§3). That makes seven numbers: γ, κ, fw, r, σ, target and maxGain. None can be read off a street frame, where the scene is unknown; each shows in an input whose structure we control. A flat card, lit at levels that double from frame to frame, has no structure for a lens to change, so every difference between its pixels is noise, while the way its mean changes from level to level is the response curve; exposure time doubles the photons exactly, so the levels are known as ratios without knowing the light in absolute terms. A printed edge is the one structure whose blur is visible. The exposure log, the gain the camera chose for each frame it took in service, is the only place the controller shows itself, because in the flats and the edge it is off. One thing is declared, not measured: the unit, a card of radiance 1 under the program's unit light.

2 · A ladder of flat frames: the gamma curve and the linear scale

Shoot the card at eleven exposures, stops −7 to +3, each double the last, so that at stop s its radiance is 2s. The camera must be in manual mode, gain 1 and no auto-exposure. With the controller on, every level dimmer than its target is brightened until its mean reaches the target, and the ladder collapses: at stops −3, −2 and −1 the mean linear value comes out 0.120, 0.120 and 0.120 instead of 0.019, 0.038 and 0.075.

If DN = q1/γ with q = κ·2s, then log₂ DN = (s + log₂ κ)/γ: the mean code against the stop is a straight line of slope 1/γ. Three levels leave it: at stops −7 and −6 noise pushes 10.3% and 1.5% of the pixels below zero, where the camera clips them to code 0, and at stop 3 every pixel sits at code 255. A level is used when under 0.5% of its pixels sit on an end code, which leaves eight; the dimmest of those are left out of the γ fit too, because where noise is large against the level it skews the mean of a power of a noisy number: the levels with mean code above 0.2 give γ = 2.198, all eight would give 2.191. With γ in hand the frames are linearised, q = DNγ, and a least-squares line through the origin of the level means against 2s gives κ = 0.1501: a unit card at gain 1 fills 15.0% of the scale.

3 · The noise law, by elimination

The same frames carry the noise. A flat card has no structure, so the spread of q across the pixels of one frame is all noise (a real sensor has a fixed pattern, which subtracting two frames of the same card cancels; the lab camera has none, and the two agree to 1.0%). Each of the eight usable levels gives a pair, the mean μ and the variance σ² of q; over the ladder μ grows 128-fold. What does σ² do? Five laws are candidates, each fitted to the eight pairs and scored by the rms of log10(measured / fitted), which treats a factor of two alike at every level.

Lawσ² =rms residualVerdict
none0—the smallest σ² measured, 1.7·10−6, is 235 standard errors above zero
constantb0.64σ² rises 86-fold across the ladder
proportionalc·μ²0.74the noise fraction falls from 28% to 2%
counting onlya·μ0.056the fit is 1.3 times too low at the darkest level, 11% too high at the brightest
counting + constanta·μ + b0.002within the sampling error (0.002)

Why the last one is not a lucky fit but the physics. The electrons a pixel collects from a steady light are random with a Poisson law, whose variance equals its mean: of e electrons on average the spread is √e. The amplifier adds a constant amount of its own, r electrons rms, whatever the signal, so in electrons the variance is e + r². At gain 1 the code is q = e/fw, and

σ² = μ/fw + (r/fw)²

a straight line in (μ, σ²) with slope 1/fw and intercept (r/fw)². The slope is the reciprocal of the electrons that fill the scale, the full well, and the intercept gives the read noise, r = fw·√intercept, in electrons: absolute units from a card whose absolute brightness we never knew, because the count's own variance is the electron counter. This is the photon-transfer method of Janesick et al. (1985); the EMVA 1288 standard (Release 4.0 Linear, 2021) writes the same line with the slope K in digital numbers per electron (here 1/fw, with full scale as 1). In this camera "full well" means the electrons that reach the top code, which the standard calls the saturation capacity; it keeps that apart from the pixel's physical full well, and in a real camera the first is normally the lower.

One correction. The converter rounds to eight bits and so adds a uniform error of variance step²/12, where a code step of 1/255 is γ·DNγ−1/255 in q: 1.0% of the variance at the darkest usable level and 2.3% at the brightest. Left in, it raises the slope by 2.1% and the full well reads 3,906 instead of 3,988; it is subtracted. Fitted by weighted least squares (weights 1/σ⁴: every level's variance has the same relative error), eleven exposures of 16 frames each give fw = 3,988 ± 6 electrons and r = 3.02 ± 0.02, the spread over 40 repeated afternoons, and with κ the electrons per unit radiance are ev = κ·fw = 599. One frame per exposure, an hour's work, gives ± 28 and ± 0.09: eleven frames of a card read the full well to 0.7%.

4 · The other two parts: blur from an edge, the controller from the log

Blur. A flat frame cannot show blur: the lens kernel is normalised, so averaging a constant returns the constant. An edge can. Print a card, black on one side and white on the other, shoot it in manual mode with the edge somewhere between two pixel columns (where, the camera is not told), and average the 24 rows and 3 colours of each column. The tempting reading takes the distance between 10% and 90% of the rise and divides by 2.563, the ratio for a Gaussian (2·1.2816 σ = 2.563 σ). On 32 frames it gives σ = 0.85 ± 0.02 px, 21% too wide and different from frame to frame: the pixel is a square that averages the light over its own area, and where the edge falls inside it decides how the first sample reads.

So calibrate the difference. The program's renderer already draws an edge as a pixel grid sees it, each pixel averaging two rays across its width; what the program lacks is the lens. Render the step as the program would, blur it with a trial σ using the program's own blur code, slide the unknown sub-pixel position of the edge, and keep the pair that lays the render on the camera's row (least squares over the 13 columns around the edge). That σ is the one blur the program has to add: σ = 0.700 ± 0.006 px from one frame, ± 0.001 from 32.

Controller. It is off in the flats and invisible in an edge, so its trace is the camera's own log: each logged frame comes with the gain g it was taken at, and from the frame we recover the level the controller saw, level = mean(q)/g. The rule is g = clamp(target/level, 1, maxGain), and it has three signatures (the widget shows them). A plateau: the largest gain in 1,000 logged street frames is 8.0, the cap, and 33.8% of the frames sit on it. A hyperbola: below the cap the mean after the gain, g × level, is the target, 0.1201 from 30 logged frames and 0.1200 from 960. And a floor at 1 that the log never shows (the smallest gain among the 1,000 frames is 1.86): declared, not measured. A frame at the cap is brightened by 8 but not made more informative, since the gain multiplies the count and its noise together: at the street's median level a pixel counts 81 electrons and its signal-to-noise ratio is 8.5, in the dimmest twentieth 4.2.

Road not taken · time
A camera integrates over an exposure time t, and what moves during it smears. A point at depth Z crossing the view at speed v moves f·v/Z pixels per second, so the smear is f·v·t/Z pixels. A pedestrian stepping out at 1.5 m/s at 8 m, exposed for 10 ms, smears 0.16 px for this camera (f = 83.1 px) and 3.1 px for one 1,920 px wide with the same field of view (f = 1,663 px). At 96 px time is below a pixel and the program leaves it out; at 1,920 px it would not be.

5 · Assemble the sensor, and price it

The seven estimates are a sensor stage: the lab's own sensor code with a switch for each part (with all three on and the street's own numbers it reproduces the street's camera bit for bit on 60 frames), fed the estimates and drawing the program's scenes and lights. The yardstick is the exact-stage program aaba, the same program with the street's true numbers; both train on the same scenes and lights (the table's seeds), so only the numbers differ. A kit of size k is k frames at each of the 11 exposures, k edge frames and 30k logged frames.

ProgramEvidenceMiss rate on the streetAlarm rate, unlabelled frames
aaaa, ideal sensornone80.3%100%
calibrated, k = 111 + 1 frames, 30 logs66.1%93.7%
calibrated, k = 16176 + 16 frames, 480 logs65.9%93.3%
aaba, the street's own numbers(the lab's privilege)65.8%93.7%

The estimated stage lands within 0.3 points of the exact stage at every size of kit (a cell moves by 0.8 points from seed to seed, and the two share their seeds): the evidence is not the bottleneck, which is the point of reading seven numbers rather than learning a picture. Which of the three parts earned the 14.5 points? Switch them on one at a time and in pairs, with the street's own numbers (a laboratory privilege), and train on each (three seeds):

Parts switched onMiss rate
noise70.0%
blur81.0%
exposure80.0%
noise + blur66.8%
noise + exposure69.7%
blur + exposure82.1%
all three (aaba)65.8%

Averaged over the six orders in which the parts could be switched on, as Lesson 2 averaged its stages, the noise is worth 12.9 points, the blur 1.3 and the exposure controller 0.3; they add to 14.5. The blur does nothing alone and 3.2 points once the noise is there; the controller does nothing in the program's light, where every frame is exposed alike. The label-free meter sees a different part from the exam: it reads 100% for noise alone, 100% for noise and blur and 92% once the controller is on, because the street's pictures are bright after the gain and the program's were dim. The controller couples the camera to the light, and §7 returns to it.

The camera as a measurement
Top: one scene of the program, dimmed by the light slider, through the ideal sensor, the sensor calibrated from the kit the main slider picks (its parts switched by the boxes) and the street's own camera, each with the chosen detector's score map (box: the pedestrian; square: the largest score) and, below, that score against the detector's threshold. Middle: the kit's evidence, drawn from fresh frames. Bottom: miss rate of the programs, from the tables. Estimates carry the spread of 40 repeated afternoons.
full well, electrons
—
read noise, electrons
—
gamma
—
blur σ, pixels
—
target · cap
—
miss, calibrated · exact
—
gain at this light
—
pedestrian CNR
—
Show the core JS
var step = gamma * Math.pow(md, gamma - 1) / 255;
o.mean = mq / nf; o.varRaw = vq / nf; o.varQ = step * step / 12; o.var = Math.max(o.varRaw - o.varQ, 1e-12);
a = (sw * sxy - sx * sy) / (sw * sxx - sx * sx); b = (sy - a * sx) / sw;
if (b < 0) { b = 0; a = sxy / sxx; }
for (i = 0; i < n; i++) { var f = a * pts[i].mean + b; w[i] = 1 / (f * f); }
return { a: a, b: b, fw: 1 / a, read: Math.sqrt(b) / a, kappa: kap, n: n, c: { const: cb, prop: cp, shot: cs },
if (sw.noise) e = e + Math.sqrt(Math.max(e, 0)) * SV.randn(rng) + cam.read * SV.randn(rng);
var q = Math.min(1, Math.max(0, gain * e / cam.fw));
out[c * HW + i] = Math.round(Math.pow(q, ig) * levels) / levels;

What to try. (1) At noon, with the flat frames shown, drag k from 1 to 32: the full well reads 4,038, then 3,998, the read noise 3.20, then 3.03, and the error bars shrink 5-fold; the street's own values, which the estimators were never told, are 4,000 and 3. (2) Show the edge, then the log: the blur reads 0.700 ± 0.006 from one frame (the street's is 0.7), and the target 0.1201 (the street's, 0.12) with the cap at 8. (3) Lower the light two ticks at a time, a stop each: the gain readout doubles, 1.94, 3.88, 7.77, then stays at the cap, and the pedestrian's CNR falls from 51.0 to 22.2 and, at 5 stops, to 3.2, where the ideal observer finds the pedestrian 37% of the time. (4) Choose the ideal-sensor program's detector: it finds the pedestrian in all three pictures at noon and, at 3 stops, raises an alarm elsewhere on the calibrated and the street's. Choose the calibrated program's: at 3 stops it finds the pedestrian in all three.

6 · Which pedestrians does noise erase?

A pedestrian of A pixels whose level differs from that of the surroundings by Δ stands out of the noise by a contrast-to-noise ratio

CNR = Δ·√A / σ

because the mean of A pixels has noise σ/√A. The calibrated law gives σ at the pedestrian's level and the frame's logged gain (σ² = (g/fw)·q + (g·r/fw)², over the three colours, with the noise's own contribution removed from the squared contrast); the labelled mask gives the pedestrian's pixels, and a ring two pixels wide around it the surroundings. Each stop of light lost multiplies the CNR by between 1/√2 (photons dominate) and 1/2 (read noise does): 0.67 for the widget's first stop, 0.55 for its fourth. A detector that knew the pedestrian's pixels exactly would find the pedestrian with probability Φ(CNR − z), Φ the normal distribution function and z set by the false-alarm budget: the exam lets one pedestrian-free frame in ten alarm, a frame has 576 cells, so each may alarm on noise alone at most 1.8 times in 10,000, which is z = 3.56 standard deviations. Under this model of the noise no detector that sees only these pixels does better than this bound.

CNR of the pedestrianIdeal observer missesExact-stage detector missesStreet-trained detector misses
under 476%80%82%
4 to 85%81%74%
8 to 120%81%63%
12 to 200%70%45%
20 to 400%52%22%
40 and over0%30%2%
all 1,3238.1%66.7%47.5%

Read it from the left. Noise erases the pedestrians below CNR 4 for any detector, 9% of the exam's, and that is all it costs the ideal observer, 8.1% of the pedestrians. The street-trained detector has seen the street's grain and still misses 48%; it needs a CNR of about 14 to find half of them, 4.0 times the bound's. The exact-stage detector draws the street's grain on the program's scenes and light, and misses 70% of the pedestrians whose CNR is between 12 and 20. So noise sets a floor and most misses sit far above it, in the other stages. For a program the floor is a rule: a pedestrian drawn at CNR below about 4 is a label that no pixel supports, and Lesson 4 returns to it when it decides how dark to draw.

7 · What the calibrated camera still gets wrong

The camera is calibrated, and the program's pictures now carry the street's grain and softness. Every frame carries the gain it was taken at, and the gain is a light meter, since level = target/gain. The program's frames, sent through the street's camera, come out at gains between 1.80 and 2.29, median 1.98: one noon sun, one white van, one blue sky. The street's 1,000 logged frames run from 1.86 to 8, median 5.9, with 34% at the cap, and 99.7% of them are darker than the program's darkest 1%. In electrons the program's median frame counts 242 per pixel (signal-to-noise 15.3), the street's 81 (8.5).

The label-free meter says the same without the arithmetic: 93.7% of unlabelled street frames still alarm for aaba and 93.3% for the calibrated program, so the camera repair was real and the gap is not closed. And the controller, worth 1.1 points under the program's one light, is worth more where the light varies: under the street's light, noise and blur without it miss 68.4% of the pedestrians, and with it (abba) 61.6%. Which pictures of one scene a program should draw is a question about light, not about the camera.

What this lesson did not do
The lab camera is exactly in the family the estimators assume: Poisson counting, one constant read noise, a Gaussian blur, a two-number controller. A real camera adds fixed-pattern noise, dark current, a black-level offset, hot pixels, shading and blur that vary across the field, a colour filter array with demosaicing, a tone curve that is not a pure power (the sRGB curve has a linear toe and a 2.4 exponent with an offset, close to 2.2 only in the mid and high tones), a rolling shutter and the processing that follows; the EMVA 1288 standard (Release 4.0 Linear, 2021) asks for at least 50 exposures with two images at each, plus dark frames, where eleven doublings suffice here. Misspecification, not sampling error, is the risk a real calibration carries, and the exam on a real street is its only test. The light is Lesson 4's, what exists in the scene Lessons 5 and 6, and what the label means Lesson 7.

Common mistakes / failure modes

"Read the noise off a dark frame, lens capped"
This camera has no black-level offset, so 50% of a dark frame's pixels clip to code 0 and the rest read r = 1.8 instead of 3.0; take it from the intercept of the line (§3).
"Noise is a fixed amount of grain: add Gaussian noise to the codes"
A constant variance misses the measured pairs by 0.64 in log10, where counting plus read noise misses by 0.002: the variance grows 86-fold over the ladder (§3).
"Read the blur off the 10–90% width of the edge"
It gives 0.85 px; the pixel's own area and the edge's phase are in the width. Fitting against the program's own renderer gives 0.70 (§4).
"More flat frames make a better camera"
They shrink the error bar of the full well from ±28 to ±6 electrons, and the miss rate stays at 66.1% for k = 1 and 65.9% for k = 16: the estimators are matched to the lab camera, so the evidence is not the bottleneck (§5).
"Noise is why the real street is hard"
An ideal observer would miss 8.1% of the exam's pedestrians for noise alone; the street-trained detector misses 48% (§6).

Checkpoint exercise

Try it
A camera in manual mode, gain 1, shows a flat card at two levels. The linearised value has mean and variance (μ, σ²) = (0.02, 1.4·10−5) at the first and (0.20, 1.04·10−4) at the second, and neither is clipped. What are its full well and its read noise, and what is the signal-to-noise ratio of a pixel at the first level? Answer: the line through the two points has slope (1.04·10−4 − 1.4·10−5)/(0.20 − 0.02) = 5.0·10−4, so fw = 1/slope = 2,000 electrons. Its intercept is 1.4·10−5 − slope·0.02 = 4.0·10−6, which is (r/fw)², so r = fw·√intercept = 4.0 electrons. A pixel at μ = 0.02 counts 0.02·fw = 40 electrons, and its signal-to-noise ratio is e/√(e + r²) = 5.3.

Where this points next

The measurement worked: the program drawn through the seven estimates misses 65.9% of the street's pedestrians against 65.8% for the exact-stage program, 13 of the 14.5 points being the noise. What a camera cannot give is the light: the program's frames leave the calibrated camera at gains between 1.8 and 2.3, while the street's run from 1.9 to 8 with 34% at the cap, and the label-free meter still reads 93%. The real street shows the same scenes at dusk, in glare and under cloud. What should be allowed to vary between two pictures of the same scene, and what must not?

Takeaway
A pixel is a measurement: a count of electrons whose variance equals its mean, plus a constant read noise, blurred by a lens, scaled by an exposure controller, clipped, and rounded after a gamma curve. Seven numbers describe that channel, and each can be read off evidence that varies one thing: flat frames at doubling exposures give the gamma curve, the linear scale and, through the line σ² = μ/fw + (r/fw)², the full well and the read noise in absolute electrons; a printed edge, fitted against the program's own renderer, gives the blur; the gain log gives the controller's target and cap. Eleven frames of a card, one of an edge and thirty log frames already give the program a sensor that misses 66.1% against 65.8% with the street's true numbers, and the noise is 13 of the 14.5 points. Noise also sets a floor: a pedestrian below a contrast-to-noise ratio of about 4 cannot be found by anyone. What the calibrated camera cannot do is choose the light, and the street's light is not the program's.

Interview prompts

Companion reads: Computer Vision · 04 Cameras and projection geometry (the pinhole half of the camera) and Computer Graphics, from first principles (the renderer whose radiance the sensor measures).