The camera is a measurement
Lesson 2 priced the camera at sixteen points of miss rate, the dearest stage, and named the evidence that measures it without a label: a grey card, an edge, the exposure log. This lesson takes the measurements. A pixel is not the light that fell on it but a count of electrons, blurred by a lens, scaled by an exposure controller, clipped, bent by a gamma curve and rounded to a code; each step has a few numbers, and each number has evidence that varies it alone. Seven numbers read off eleven exposures of a card, one printed edge and a log become a sensor stage, and the program drawn through it misses 66% of the street's pedestrians instead of 80%, within a point of the program built with the street's own camera. It cannot say what light the camera should be shown.
New idea: a pixel is a measurement: a signal plus noise whose law has a few numbers, and each number can be read off evidence that varies one thing while the rest holds still. Seven numbers from flat frames, an edge and a log rebuild the street's camera to within the error bars of §3 and §4, and the program drawn through them lands within 0.1 points of the program built with the street's true numbers.
Forces next: Flat frames of a grey card gave the camera's full well and read noise, an edge gave its blur, and logs gave its exposure control; added to the program, they take the miss rate from 80% to 66%, within a point of what the street's own camera numbers give, and the program's pictures now carry the grain and softness of the real camera. A calibrated camera still sees only the light the program gives it: one sun behind the camera, white vans, bright shirts, a blue sky. The program's frames leave its exposure controller at a gain near 2, the street's run from 1.9 to the cap of 8 with a third of them at the cap, and the controller that is worth 1 point in the program's light is worth 7 in the street's. The real street shows the same scenes at dusk, in glare and under cloud. What should be allowed to vary between two pictures of the same scene, and what must not?
1 · The ideal sensor, and the channel it leaves out
Lesson 2 found the camera the dearest stage. The detector trained on the program misses 80.3% of the street's pedestrians; trained on the program with the street's camera and nothing else changed (aaba) it misses 65.8%, while the alarm rate on unlabelled street frames moves only from 100% to 93.7%. The program's sensor is ideal: it multiplies the renderer's radiance by a fixed exposure, bends it with the gamma curve and rounds it to eight bits, so a pixel is a function of its radiance. The street's camera is not: send one radiance through it twice and the two pictures differ. The widget of §5 does exactly this; at noon the street's picture is already grainy and soft, and a few stops dimmer it is mostly grain.
What the program needs from a camera is therefore not a number per pixel but a distribution: given a radiance image, the probability of each code image. The street's camera builds it in six steps, and each adds a few numbers to find.
| Step | What happens to the light of a pixel | Number to find |
|---|---|---|
| Lens | each point is spread over its neighbours: a Gaussian blur of width σ pixels | σ |
| Count | radiance L becomes e = ev·L electrons on average; the number actually counted is random | ev |
| Read | the amplifier adds noise of r electrons rms, whatever the signal | r |
| Gain | the exposure controller multiplies by a gain g between 1 and maxGain, chosen from the frame's own brightness to bring its mean to a target | target, maxGain |
| Clip | the converter saturates at fw electrons (at gain 1): q = g·e/fw is cut to the range 0 to 1 | fw |
| Code | the gamma curve DN = q1/γ, then rounding to eight bits | γ |
The count and the clip meet in κ = ev/fw, the fraction of full scale that one unit of radiance fills at gain 1, while fw alone sets the size of the noise (§3). That makes seven numbers: γ, κ, fw, r, σ, target and maxGain. None can be read off a street frame, where the scene is unknown; each shows in an input whose structure we control. A flat card, lit at levels that double from frame to frame, has no structure for a lens to change, so every difference between its pixels is noise, while the way its mean changes from level to level is the response curve; exposure time doubles the photons exactly, so the levels are known as ratios without knowing the light in absolute terms. A printed edge is the one structure whose blur is visible. The exposure log, the gain the camera chose for each frame it took in service, is the only place the controller shows itself, because in the flats and the edge it is off. One thing is declared, not measured: the unit, a card of radiance 1 under the program's unit light.
2 · A ladder of flat frames: the gamma curve and the linear scale
Shoot the card at eleven exposures, stops −7 to +3, each double the last, so that at stop s its radiance is 2s. The camera must be in manual mode, gain 1 and no auto-exposure. With the controller on, every level dimmer than its target is brightened until its mean reaches the target, and the ladder collapses: at stops −3, −2 and −1 the mean linear value comes out 0.120, 0.120 and 0.120 instead of 0.019, 0.038 and 0.075.
If DN = q1/γ with q = κ·2s, then log₂ DN = (s + log₂ κ)/γ: the mean code against the stop is a straight line of slope 1/γ. Three levels leave it: at stops −7 and −6 noise pushes 10.3% and 1.5% of the pixels below zero, where the camera clips them to code 0, and at stop 3 every pixel sits at code 255. A level is used when under 0.5% of its pixels sit on an end code, which leaves eight; the dimmest of those are left out of the γ fit too, because where noise is large against the level it skews the mean of a power of a noisy number: the levels with mean code above 0.2 give γ = 2.198, all eight would give 2.191. With γ in hand the frames are linearised, q = DNγ, and a least-squares line through the origin of the level means against 2s gives κ = 0.1501: a unit card at gain 1 fills 15.0% of the scale.
3 · The noise law, by elimination
The same frames carry the noise. A flat card has no structure, so the spread of q across the pixels of one frame is all noise (a real sensor has a fixed pattern, which subtracting two frames of the same card cancels; the lab camera has none, and the two agree to 1.0%). Each of the eight usable levels gives a pair, the mean μ and the variance σ² of q; over the ladder μ grows 128-fold. What does σ² do? Five laws are candidates, each fitted to the eight pairs and scored by the rms of log10(measured / fitted), which treats a factor of two alike at every level.
| Law | σ² = | rms residual | Verdict |
|---|---|---|---|
| none | 0 | — | the smallest σ² measured, 1.7·10−6, is 235 standard errors above zero |
| constant | b | 0.64 | σ² rises 86-fold across the ladder |
| proportional | c·μ² | 0.74 | the noise fraction falls from 28% to 2% |
| counting only | a·μ | 0.056 | the fit is 1.3 times too low at the darkest level, 11% too high at the brightest |
| counting + constant | a·μ + b | 0.002 | within the sampling error (0.002) |
Why the last one is not a lucky fit but the physics. The electrons a pixel collects from a steady light are random with a Poisson law, whose variance equals its mean: of e electrons on average the spread is √e. The amplifier adds a constant amount of its own, r electrons rms, whatever the signal, so in electrons the variance is e + r². At gain 1 the code is q = e/fw, and
σ² = μ/fw + (r/fw)²
a straight line in (μ, σ²) with slope 1/fw and intercept (r/fw)². The slope is the reciprocal of the electrons that fill the scale, the full well, and the intercept gives the read noise, r = fw·√intercept, in electrons: absolute units from a card whose absolute brightness we never knew, because the count's own variance is the electron counter. This is the photon-transfer method of Janesick et al. (1985); the EMVA 1288 standard (Release 4.0 Linear, 2021) writes the same line with the slope K in digital numbers per electron (here 1/fw, with full scale as 1). In this camera "full well" means the electrons that reach the top code, which the standard calls the saturation capacity; it keeps that apart from the pixel's physical full well, and in a real camera the first is normally the lower.
One correction. The converter rounds to eight bits and so adds a uniform error of variance step²/12, where a code step of 1/255 is γ·DNγ−1/255 in q: 1.0% of the variance at the darkest usable level and 2.3% at the brightest. Left in, it raises the slope by 2.1% and the full well reads 3,906 instead of 3,988; it is subtracted. Fitted by weighted least squares (weights 1/σ⁴: every level's variance has the same relative error), eleven exposures of 16 frames each give fw = 3,988 ± 6 electrons and r = 3.02 ± 0.02, the spread over 40 repeated afternoons, and with κ the electrons per unit radiance are ev = κ·fw = 599. One frame per exposure, an hour's work, gives ± 28 and ± 0.09: eleven frames of a card read the full well to 0.7%.
4 · The other two parts: blur from an edge, the controller from the log
Blur. A flat frame cannot show blur: the lens kernel is normalised, so averaging a constant returns the constant. An edge can. Print a card, black on one side and white on the other, shoot it in manual mode with the edge somewhere between two pixel columns (where, the camera is not told), and average the 24 rows and 3 colours of each column. The tempting reading takes the distance between 10% and 90% of the rise and divides by 2.563, the ratio for a Gaussian (2·1.2816 σ = 2.563 σ). On 32 frames it gives σ = 0.85 ± 0.02 px, 21% too wide and different from frame to frame: the pixel is a square that averages the light over its own area, and where the edge falls inside it decides how the first sample reads.
So calibrate the difference. The program's renderer already draws an edge as a pixel grid sees it, each pixel averaging two rays across its width; what the program lacks is the lens. Render the step as the program would, blur it with a trial σ using the program's own blur code, slide the unknown sub-pixel position of the edge, and keep the pair that lays the render on the camera's row (least squares over the 13 columns around the edge). That σ is the one blur the program has to add: σ = 0.700 ± 0.006 px from one frame, ± 0.001 from 32.
Controller. It is off in the flats and invisible in an edge, so its trace is the camera's own log: each logged frame comes with the gain g it was taken at, and from the frame we recover the level the controller saw, level = mean(q)/g. The rule is g = clamp(target/level, 1, maxGain), and it has three signatures (the widget shows them). A plateau: the largest gain in 1,000 logged street frames is 8.0, the cap, and 33.8% of the frames sit on it. A hyperbola: below the cap the mean after the gain, g × level, is the target, 0.1201 from 30 logged frames and 0.1200 from 960. And a floor at 1 that the log never shows (the smallest gain among the 1,000 frames is 1.86): declared, not measured. A frame at the cap is brightened by 8 but not made more informative, since the gain multiplies the count and its noise together: at the street's median level a pixel counts 81 electrons and its signal-to-noise ratio is 8.5, in the dimmest twentieth 4.2.
5 · Assemble the sensor, and price it
The seven estimates are a sensor stage: the lab's own sensor code with a switch for each part (with all three on and the street's own numbers it reproduces the street's camera bit for bit on 60 frames), fed the estimates and drawing the program's scenes and lights. The yardstick is the exact-stage program aaba, the same program with the street's true numbers; both train on the same scenes and lights (the table's seeds), so only the numbers differ. A kit of size k is k frames at each of the 11 exposures, k edge frames and 30k logged frames.
| Program | Evidence | Miss rate on the street | Alarm rate, unlabelled frames |
|---|---|---|---|
aaaa, ideal sensor | none | 80.3% | 100% |
| calibrated, k = 1 | 11 + 1 frames, 30 logs | 66.1% | 93.7% |
| calibrated, k = 16 | 176 + 16 frames, 480 logs | 65.9% | 93.3% |
aaba, the street's own numbers | (the lab's privilege) | 65.8% | 93.7% |
The estimated stage lands within 0.3 points of the exact stage at every size of kit (a cell moves by 0.8 points from seed to seed, and the two share their seeds): the evidence is not the bottleneck, which is the point of reading seven numbers rather than learning a picture. Which of the three parts earned the 14.5 points? Switch them on one at a time and in pairs, with the street's own numbers (a laboratory privilege), and train on each (three seeds):
| Parts switched on | Miss rate |
|---|---|
| noise | 70.0% |
| blur | 81.0% |
| exposure | 80.0% |
| noise + blur | 66.8% |
| noise + exposure | 69.7% |
| blur + exposure | 82.1% |
all three (aaba) | 65.8% |
Averaged over the six orders in which the parts could be switched on, as Lesson 2 averaged its stages, the noise is worth 12.9 points, the blur 1.3 and the exposure controller 0.3; they add to 14.5. The blur does nothing alone and 3.2 points once the noise is there; the controller does nothing in the program's light, where every frame is exposed alike. The label-free meter sees a different part from the exam: it reads 100% for noise alone, 100% for noise and blur and 92% once the controller is on, because the street's pictures are bright after the gain and the program's were dim. The controller couples the camera to the light, and §7 returns to it.
What to try. (1) At noon, with the flat frames shown, drag k from 1 to 32: the full well reads 4,038, then 3,998, the read noise 3.20, then 3.03, and the error bars shrink 5-fold; the street's own values, which the estimators were never told, are 4,000 and 3. (2) Show the edge, then the log: the blur reads 0.700 ± 0.006 from one frame (the street's is 0.7), and the target 0.1201 (the street's, 0.12) with the cap at 8. (3) Lower the light two ticks at a time, a stop each: the gain readout doubles, 1.94, 3.88, 7.77, then stays at the cap, and the pedestrian's CNR falls from 51.0 to 22.2 and, at 5 stops, to 3.2, where the ideal observer finds the pedestrian 37% of the time. (4) Choose the ideal-sensor program's detector: it finds the pedestrian in all three pictures at noon and, at 3 stops, raises an alarm elsewhere on the calibrated and the street's. Choose the calibrated program's: at 3 stops it finds the pedestrian in all three.
6 · Which pedestrians does noise erase?
A pedestrian of A pixels whose level differs from that of the surroundings by Δ stands out of the noise by a contrast-to-noise ratio
CNR = Δ·√A / σ
because the mean of A pixels has noise σ/√A. The calibrated law gives σ at the pedestrian's level and the frame's logged gain (σ² = (g/fw)·q + (g·r/fw)², over the three colours, with the noise's own contribution removed from the squared contrast); the labelled mask gives the pedestrian's pixels, and a ring two pixels wide around it the surroundings. Each stop of light lost multiplies the CNR by between 1/√2 (photons dominate) and 1/2 (read noise does): 0.67 for the widget's first stop, 0.55 for its fourth. A detector that knew the pedestrian's pixels exactly would find the pedestrian with probability Φ(CNR − z), Φ the normal distribution function and z set by the false-alarm budget: the exam lets one pedestrian-free frame in ten alarm, a frame has 576 cells, so each may alarm on noise alone at most 1.8 times in 10,000, which is z = 3.56 standard deviations. Under this model of the noise no detector that sees only these pixels does better than this bound.
| CNR of the pedestrian | Ideal observer misses | Exact-stage detector misses | Street-trained detector misses |
|---|---|---|---|
| under 4 | 76% | 80% | 82% |
| 4 to 8 | 5% | 81% | 74% |
| 8 to 12 | 0% | 81% | 63% |
| 12 to 20 | 0% | 70% | 45% |
| 20 to 40 | 0% | 52% | 22% |
| 40 and over | 0% | 30% | 2% |
| all 1,323 | 8.1% | 66.7% | 47.5% |
Read it from the left. Noise erases the pedestrians below CNR 4 for any detector, 9% of the exam's, and that is all it costs the ideal observer, 8.1% of the pedestrians. The street-trained detector has seen the street's grain and still misses 48%; it needs a CNR of about 14 to find half of them, 4.0 times the bound's. The exact-stage detector draws the street's grain on the program's scenes and light, and misses 70% of the pedestrians whose CNR is between 12 and 20. So noise sets a floor and most misses sit far above it, in the other stages. For a program the floor is a rule: a pedestrian drawn at CNR below about 4 is a label that no pixel supports, and Lesson 4 returns to it when it decides how dark to draw.
7 · What the calibrated camera still gets wrong
The camera is calibrated, and the program's pictures now carry the street's grain and softness. Every frame carries the gain it was taken at, and the gain is a light meter, since level = target/gain. The program's frames, sent through the street's camera, come out at gains between 1.80 and 2.29, median 1.98: one noon sun, one white van, one blue sky. The street's 1,000 logged frames run from 1.86 to 8, median 5.9, with 34% at the cap, and 99.7% of them are darker than the program's darkest 1%. In electrons the program's median frame counts 242 per pixel (signal-to-noise 15.3), the street's 81 (8.5).
The label-free meter says the same without the arithmetic: 93.7% of unlabelled street frames still alarm for aaba and 93.3% for the calibrated program, so the camera repair was real and the gap is not closed. And the controller, worth 1.1 points under the program's one light, is worth more where the light varies: under the street's light, noise and blur without it miss 68.4% of the pedestrians, and with it (abba) 61.6%. Which pictures of one scene a program should draw is a question about light, not about the camera.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
The measurement worked: the program drawn through the seven estimates misses 65.9% of the street's pedestrians against 65.8% for the exact-stage program, 13 of the 14.5 points being the noise. What a camera cannot give is the light: the program's frames leave the calibrated camera at gains between 1.8 and 2.3, while the street's run from 1.9 to 8 with 34% at the cap, and the label-free meter still reads 93%. The real street shows the same scenes at dusk, in glare and under cloud. What should be allowed to vary between two pictures of the same scene, and what must not?
Interview prompts
- What does a pixel value measure, and what is random about it? (§1, §3 — a count of electrons whose Poisson variance equals its mean, plus a read noise independent of the signal; then blur, gain, clip and gamma; the program has to sample from that channel.)
- How do flat frames give a camera's full well and read noise when the card's absolute brightness is unknown? (§3 — variance against mean is a line, σ² = μ/fw + (r/fw)²; the count's Poisson variance supplies the scale: the slope is 1/fw and the intercept gives r.)
- Why is the 10–90% width of an edge a poor estimate of blur? (§4 — it contains the pixel's own area and the edge's sub-pixel phase; fit the difference between the program's own render of the edge and the camera's.)
- Which part of a camera model is worth simulating, and how would you know? (§5 — switch the parts on one by one and in pairs and average over the orders: here the noise, with the blur worth something only beside it and the controller nothing in one light.)
- A synthetic pedestrian has a contrast-to-noise ratio of 2. What do you do with the label? (§6 — nobody can find the pedestrian: the ideal observer's z for a ten per cent false-alarm budget over 576 cells is about 3.6, so the frame teaches the detector to alarm on noise; drop it or draw it brighter.)
- The program's camera is calibrated. Why does the label-free meter still alarm on most street frames? (§7 — the camera is only as good as the light it is shown; the program's frames sit at one gain, the street's range over several stops.)
Companion reads: Computer Vision · 04 Cameras and projection geometry (the pinhole half of the camera) and Computer Graphics, from first principles (the renderer whose radiance the sensor measures).