What a pixel measures: the forward model and the exam
Lesson 6 ended with a surface that answers "where does this ray first hit?" and has no colour. Give the exam that surface, exact and painted one colour, and it scores 19.2 dB on cameras held out of the data, because a photograph is made of colours. This lesson derives what a pixel measures (for a diffuse surface, its albedo times a cosine, the same from every camera), writes the forward model that renders it, and defines the exam that scores it. Since the colour belongs to the point, photographs can grade a geometry without ground truth. Then it fits the geometry by gradient descent on the rendering error and meets the wall: with first-hit visibility the loss is a staircase and its derivative is exactly zero.
New idea: the forward model: a pixel is the shaded colour of the first surface its ray meets, R(θ), and the exam scores R(θ) against photographs from cameras it has not seen. Because a diffuse point has one colour for every camera, photographs can grade geometry; because the first hit is a discrete choice, the grade cannot yet be descended.
Forces next: We can render a model and score it against a photograph, but we cannot yet improve it by gradient: where one surface hides another, a pixel is a cliff in the geometry, so the loss is flat almost everywhere and says nothing about which way to move. What would a renderer look like whose pixels change smoothly with the scene, and still reduce to the hard one when the scene is solid?
1 · A perfect surface, a failed exam
Lesson 6 left a surface. Grant it the best case, the exact outlines of the statue, the crate and the ball, and set the exam. Twelve cameras on a ring of radius 6 m photograph the scene (focal length 56 px, 64 pixels, a field of view of 59.5°); these photographs are the data. Twenty-four more cameras sit on the same ring between them, at angles the data never used. A model is anything that returns a photograph when given a camera. Its exam score is the average, over those 24 held-out cameras, of the PSNR between its photograph and the real one:
PSNR = −10·log10(MSE), MSE = mean over pixels and colour channels of (prediction − photograph)², colours in [0, 1]
Each 10 dB is a factor of ten in MSE: 30 dB is an rms error of 3% of the colour range, 20 dB is 10%, 10 dB is 32%. The cameras must be held out. A model that stores the 12 photographs scores an infinite PSNR on the cameras it was built from; made to answer a new camera with the nearest stored photograph it scores 14.0 dB, and answering every pixel with the background colour scores 4.6 dB. Remembering the data is free. Predicting is the exam.
Now paint the perfect surface with one colour, the average colour of the object pixels in the data, and render the 24 held-out cameras: 19.2 dB. The geometry is exact and the score is poor, because the held-out photographs are made of colours. What colour does a surface point have, and does it depend on who is looking?
2 · What a pixel measures
A pixel counts light. What it counts is radiance: power per unit of area across the ray, per unit of solid angle (in Flatland, per unit of length and of angle). Radiance does not change along a ray through empty space, so a pixel reports the radiance leaving the first surface point its ray meets, in the direction back toward the camera. What is that radiance? Take the simplest surface, a diffuse (Lambertian) one, and ask how much light arrives and how much leaves.
How much arrives. Parallel light from direction l (a unit vector pointing at the lamp) carries power E0 per metre of cross-section across the beam. A piece of surface of length s, whose unit normal n makes angle θ with l, presents only a cross-section s·cos θ to the beam, so it catches E0·s·cos θ: per metre of surface, E0(n·l). The cosine is foreshortening and nothing more. A surface facing away (n·l < 0) catches nothing.
How much leaves. A fraction a, the albedo (one number per colour channel), is sent back, and "diffuse" means that it leaves with the same radiance in every direction. (Less power goes toward grazing directions, but the surface also looks narrower from there, by the same cosine.) The light that arrives from everywhere else, sky and neighbouring surfaces, is lumped into one constant, amb. The colour of a surface point is therefore (in units where a white surface facing the lamp reads 1)
I = a · ( amb + (1 − amb) · max(0, n·l) ), amb = 0.3, l = (−0.57, 0.82)
(l points up and to the left in the top-down views.) Two things stand out. First, no direction of observation appears: every camera that sees the point reads the same colour. Second, the colour is a product of the paint a, carried by the surface, and the shading I(n), set by how the surface tilts toward the lamp. The shading is not a detail: its factor runs from 1.00 where the surface faces the lamp (θ = 0°), through 0.91 at 30° and 0.65 at 60°, to 0.30 from 90° on, a factor of 3.3. Shadows, interreflection, gloss and transparency are ignored here (Computer Graphics 07 and 08 do them properly).
3 · The forward model, and the whole exam
Now the photograph that a model predicts can be written down. A scene is a set of parameters θ: a shape (a signed distance function), an albedo for every surface point, the lamp l and amb. For a camera and one of its pixels, with ray ρ,
R(θ; camera)pixel = a(x*) · I( n(x*) ), x* = the first point where ρ meets the surface (the background colour if it never does)
Two steps: visibility, which point is nearest along the ray, then shading, the formula above. Visibility is found by casting the ray (sphere tracing on the signed distance field, as in lesson 6) or by rasterising: project every primitive and keep the nearest one per pixel in a depth buffer. Either way each pixel takes the colour of exactly one point, which makes R a hard renderer. It is the forward model, from scene to photographs, and the exam measures how far R(θ) is from photographs it was not fitted to. Run it on a ladder of models; each row adds an ingredient of R:
| model (all given the true surface except the first two) | held-out PSNR |
|---|---|
| every pixel the background colour | 4.6 dB |
| the nearest stored photograph, no geometry at all | 14.0 dB |
| one colour for every surface | 19.2 dB |
| true albedo, no light | 15.6 dB |
| true light, one albedo for everything | 21.9 dB |
| true light, each object its own base colour (no stripes) | 23.0 dB |
| colour copied from the nearest data photograph that sees the point | 34.6 dB |
| true albedo times true light, which is R(θ) | infinite: it is the photograph |
Read the ladder. The light matters as much as the paint: with the true shading and one albedo the score is 21.9 dB, but the true albedo without shading scores only 15.6 dB, worse than one flat colour. Per-object base colours add about a decibel, and the stripes are the rest. Colours copied from the data reach 34.6 dB, because of §2: a diffuse point has one colour for every camera, so a pixel of a data photograph that sees the point is a valid sample of its colour for any other camera that sees it. (Lesson 6's copy did not ask who sees the point. This row moves between 27.9 and 34.6 dB as the number of data cameras goes from 6 to 24, depending on how the pixel grids happen to line up.) Once the surface is known, colour is cheap.
But every row was handed the true surface. A phone has no range sensor; the surface has to come from the photographs themselves.
4 · Turn Lambert around: photo-consistency
The copy row works because a point has one colour for every camera. Read that backwards. Suppose we only hypothesise where the surface is. Take a pixel of camera A and a candidate depth z along its ray, which gives a point P(z). Project P into other cameras and read the colours there. If P is on the surface and the cameras can see it, the colours agree. If not, each camera reads some other surface point, and they disagree. So
C(z) = variance of the colours that the cameras read at P(z), camera A included
scores a geometry from photographs alone: no ground truth, no colour model. This is photo-consistency, and sweeping z along one ray is a plane sweep for one pixel. Test it on the 228 pixels of the data cameras that see the statue, with z from 2.5 to 10 m in 1 cm steps, and call a pixel right when the minimum of C is within 10 cm of the truth:
| cameras that vote with camera A | how many | pixels right | median error |
|---|---|---|---|
| its two neighbours, 30° either side | 2 | 65% | 3.1 cm |
| every camera within 90° either side | 6 | 38% | 18 cm |
| only the cameras that see the true point | 3.4 on average | 94% | 1.9 cm |
Three readings. (a) It works: with the right cameras voting, nearly every minimum is within centimetres. (b) More cameras should help, and they hurt. A camera on the far side of the statue cannot see P; it reads some other surface and spoils the vote. Which cameras see P depends on where the surface is, which is what we are looking for, and the last row cheated by asking the true surface. Visibility is inside the score. (c) It is blind where the photograph is flat. In the quarter of the pixels where colour changes least from one pixel to the next, the valley C ≤ 0.001 is 0.69 m wide; in the steepest quarter, 0.15 m.
Photo-consistency says whether a hypothesis is good, one pixel at a time, and its vote needs the visibility that only a model of the whole scene can supply. Lesson 5 showed the way: gather the unknowns in one model, predict the data, compare, update. The model is now R, and it brings shading and visibility with it.
5 · Fit the model: why a gradient
Gather the unknowns into θ and score them on the data cameras:
E(θ) = (1/N) · Σcameras, pixels | R(θ; camera)pixel − photographpixel |², N = 12 × 64 = 768 pixels, mean over the 3 channels
This is the exam's squared error, taken on the data cameras instead of the held-out ones. Why update by gradient and not by search? Count. A search tries n values of each of d parameters, nd renders: 106 for the three numbers of one circle at n = 100, 102000 for a thousand numbers. A gradient costs d + 1 renders by finite differences, and a small multiple of one render (a forward and a backward pass) by reverse-mode differentiation, whatever d is. Only the gradient scales. It is worth having only if it says which way to move, so test the smallest unknown there is.
6 · The wall: a pixel is a cliff in the geometry
Let the model be one circle of radius r, centred at (δ, 0), painted one flat colour so that nothing but visibility can change: a pixel shows that colour if its ray meets the circle, the background otherwise. The cleanest case is a disc of radius 1 m photographed by six cameras on a ring of radius 5 m (48 pixels each, f = 40 px), with a model that is the same disc, its centre off by δ.
Take one pixel. Its ray, with unit direction d, crosses the line z = 0 at a point x0, at an angle whose sine is |dz|. The distance from the circle's centre (δ, 0) to the ray's line is therefore |δ − x0|·|dz|, and the ray meets the circle exactly when that is at most r (the circle is always in front of the camera here):
|δ − x0| ≤ w, w = r / |dz|
So pixel k shows the circle's colour for δ in an interval of half-width wk around x0,k, and the background colour outside it. Its squared error against the photograph is e¹k in the first case and e⁰k in the second (two fixed numbers), and the loss is a sum of boxes, 1(·) being 1 when its condition holds and 0 otherwise:
E(δ) = (1/N) Σk [ e⁰k + (e¹k − e⁰k) · 1( |δ − x0,k| ≤ wk ) ]
Read it. Between two interval ends nothing in any picture changes, so E is constant: a plateau, with derivative exactly 0. At an interval end a ray grazes the circle, which starts or stops hiding what lies behind it; one pixel changes colour, and E jumps by (e¹ − e⁰)/N: a cliff. Cliffs sit where a ray is tangent to the model, wherever the photographs are; the photographs only set their heights. So the derivative of the loss has two parts. One is the change in the colours of pixels whose visibility stays the same: differentiating the program, by hand or by autodiff, finds it, and for a flat colour it is zero. The other is the impulse at each cliff, the boundary term, and differentiating never sees it, because the switch is a comparison and the derivative of a comparison is zero. Count the cliffs over a sweep of δ from −1.6 m to 1.6 m in steps of one millimetre:
| unit disc, 6 cameras (the clean case) | the statue's photographs, 12 cameras, r = 1.15 m | |
|---|---|---|
| cliffs inside the sweep | 228 | 500 |
| distinct values of E over the sweep | 59 | 472 |
| 1 mm steps on which E does not change | 96.4% | 85.3% |
(Opposite cameras of the disc see the same rays, so its 228 cliffs sit at only 114 places, and E, which only counts how many pixels are wrong, takes few values. The statue's model is a circle of its mean radius, painted the average colour of the object pixels.) Zoom out and the staircase becomes a bowl, because hundreds of tiny steps add up to a large-scale slope. That slope is real, but it exists only across many plateaus, and measuring it costs two renders per parameter. The derivative a machine computes in one backward pass for millions of parameters is the one at scale zero, and there it is zero.
Anti-aliasing. A pixel is a footprint, not a ray. With S sub-rays per pixel, each cliff becomes S cliffs 1/S as tall: a finer staircase. With 16 sub-rays, 61% of the positions on the sweep still have an exactly zero derivative at step 0.1 mm, against 97% with one ray. Only the limit S → ∞ removes the steps: the pixel shows the circle with weight c, the fraction of its footprint the circle covers, which is the overlap of the pixel [i, i+1] with the circle's image, the interval between its two tangent rays. As δ changes, each cliff becomes a ramp one footprint wide (0.107 m at 6 m for a camera looking across the motion). That limit is the orange curve in the widget.
What to try. The page opens with the model 1.5 m to the left of the statue (δ = −1.5 m), the whole sweep in view, and held-out camera 17, whose image runs left to right like the map. At this scale the two losses are one bowl: 0.2012 for the first-hit renderer, 0.1981 anti-aliased, and the model's exam score is 7.5 dB. The first-hit derivative reads exactly 0; the anti-aliased one reads −0.0675 per metre, and differencing the first-hit loss across ±0.1 m, a window holding 32 cliffs, gives −0.0691: the bowl's slope is real, and exists only at that scale. Press descend (first hit): nothing moves, and nothing would from 97% of the starting positions on the slider. Press descend (anti-aliased): 100 steps of Adam walk the model to −0.133 m. Now set magnify to 4, a 25 mm window. The plateau under δ is 8.08 mm wide, between cliffs at −1.5058 m and −1.4977 m; step δ a millimetre at a time across the right-hand one and E drops by 0.00078, one pixel in 768 changing colour. The orange curve is a smooth ramp, and "pixels that feel δ" says why: 0 for the first-hit renderer, 24 of 768 for the anti-aliased one, the pixels that contain the circle's edge. The violet exam curve has cliffs of its own and disagrees with the training loss about the best offset: the exam peaks at +0.085 m with 9.1 dB (one flat circle is a poor model of this scene), the anti-aliased training loss is lowest at −0.067 m. Finally, in the full window, start from δ = −2.0: the anti-aliased descent stops at −0.461 m, in a shallow valley 0.012 above that minimum. A gradient exists there, and it is local.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
Rendering is now a function we can write down and score: first hit, then Lambert, compared with photographs from cameras the model never saw. With the true surface and colours copied from the data it earns 34.6 dB, and photographs can grade a geometry with no ground truth. It is also a function we cannot yet descend: on the statue's photographs the loss of the simplest model is constant between 500 cliffs inside a 3.2 m sweep, and the first-hit derivative is exactly zero from 97% of the positions on the slider (the clean disc of §6 is where lesson 8 begins). Exact anti-aliasing gives the model's own edge a slope that rests on 24 of 768 pixels, and a surface that no ray grazes, hidden by another surface or not there yet, gets nothing. The cause is the first-hit rule itself: which surface a ray meets first is a discrete choice. What would a renderer look like whose pixels change smoothly with the scene, and still reduce to the hard one when the scene is solid?
Interview prompts
- Derive the Lambertian shading factor, and say why the colour of a point does not depend on the camera. (§2 — foreshortening gives E0 cos θ per metre of surface; diffuse reflection has equal radiance in every direction, so no viewing direction appears.)
- What is a forward model, and how is it scored? Why are the cameras held out? (§1, §3 — R(θ) maps scene to photographs; the score is PSNR on cameras not used, because a model that stores its photographs is perfect on them and predicts nothing.)
- What is photo-consistency, and why can adding cameras make it worse? (§4 — the variance of the colours the cameras read at a hypothesised point; a camera that cannot see the point reads another surface, so visibility is inside the score.)
- Why update a scene by gradient rather than by search? (§5 — search costs nd renders, a gradient a small multiple of one, whatever the number of parameters.)
- Why is the derivative of a first-hit rendering loss zero almost everywhere, and where are the cliffs of a circle? (§6 — a pixel changes only when its ray grazes the circle, at |δ − x0| = r/|dz|; between those offsets the picture, and so the loss, is constant.)
- What does anti-aliasing change, and what does it leave? (§6, Road — each cliff becomes a ramp a pixel wide, so the model's own edge has a slope; hidden or missing surfaces have no edge pixels and get none.)
Companion reads: Computer Graphics · 07 Light and the rendering equation (radiance), Computer Graphics · 08 Local shading and PBR (Lambert and what replaces it), Computer Graphics · 06 Sampling and anti-aliasing (the pixel as a footprint), and Computer Vision · 18 Image restoration (what PSNR does and does not measure).