all_lessons/3D Vision/07 · What a pixel measures: the forward model and the examlesson 7 / 14

What a pixel measures: the forward model and the exam

Lesson 6 ended with a surface that answers "where does this ray first hit?" and has no colour. Give the exam that surface, exact and painted one colour, and it scores 19.2 dB on cameras held out of the data, because a photograph is made of colours. This lesson derives what a pixel measures (for a diffuse surface, its albedo times a cosine, the same from every camera), writes the forward model that renders it, and defines the exam that scores it. Since the colour belongs to the point, photographs can grade a geometry without ground truth. Then it fits the geometry by gradient descent on the rendering error and meets the wall: with first-hit visibility the loss is a staircase and its derivative is exactly zero.

The thesis, here
A pixel records the light leaving the first surface its ray meets. For a diffuse surface that light is albedo times a cosine and does not depend on the camera, so a model of a scene is a function from scene to photographs (the forward model), the exam compares it with photographs from cameras it never saw, and photographs can grade a geometry on their own, because a correct surface must look the same to every camera that sees it. The same first-surface rule makes the rendering error a staircase in the geometry, and a staircase has no gradient to follow.
Linear position
Forced by: Fusing noisy depth into a signed distance field gives a surface that answers "where is the first hit along this ray?", with a normal and an inside and an outside. But the held-out photograph is made of colors, not positions, and so far nothing has a color. What does a pixel actually measure, and how do we write down, and score, the image a model predicts?
New idea: the forward model: a pixel is the shaded colour of the first surface its ray meets, R(θ), and the exam scores R(θ) against photographs from cameras it has not seen. Because a diffuse point has one colour for every camera, photographs can grade geometry; because the first hit is a discrete choice, the grade cannot yet be descended.
Forces next: We can render a model and score it against a photograph, but we cannot yet improve it by gradient: where one surface hides another, a pixel is a cliff in the geometry, so the loss is flat almost everywhere and says nothing about which way to move. What would a renderer look like whose pixels change smoothly with the scene, and still reduce to the hard one when the scene is solid?
The plan
Six moves. (1) Give the exam a perfect surface painted one colour, and watch it fail. (2) Derive what a pixel measures: the Lambert colour of the first hit. (3) Write the forward model and run the whole exam on a ladder of models. (4) Turn Lambert around: photo-consistency grades geometry from photographs alone, if we know who can see what. (5) Fit the model by descent on the rendering error, and say why descent. (6) Find the wall: for first-hit rendering the loss is a staircase, and anti-aliasing mends only the edge.

1 · A perfect surface, a failed exam

Lesson 6 left a surface. Grant it the best case, the exact outlines of the statue, the crate and the ball, and set the exam. Twelve cameras on a ring of radius 6 m photograph the scene (focal length 56 px, 64 pixels, a field of view of 59.5°); these photographs are the data. Twenty-four more cameras sit on the same ring between them, at angles the data never used. A model is anything that returns a photograph when given a camera. Its exam score is the average, over those 24 held-out cameras, of the PSNR between its photograph and the real one:

PSNR = −10·log10(MSE),   MSE = mean over pixels and colour channels of (prediction − photograph)²,   colours in [0, 1]

Each 10 dB is a factor of ten in MSE: 30 dB is an rms error of 3% of the colour range, 20 dB is 10%, 10 dB is 32%. The cameras must be held out. A model that stores the 12 photographs scores an infinite PSNR on the cameras it was built from; made to answer a new camera with the nearest stored photograph it scores 14.0 dB, and answering every pixel with the background colour scores 4.6 dB. Remembering the data is free. Predicting is the exam.

Now paint the perfect surface with one colour, the average colour of the object pixels in the data, and render the 24 held-out cameras: 19.2 dB. The geometry is exact and the score is poor, because the held-out photographs are made of colours. What colour does a surface point have, and does it depend on who is looking?

2 · What a pixel measures

A pixel counts light. What it counts is radiance: power per unit of area across the ray, per unit of solid angle (in Flatland, per unit of length and of angle). Radiance does not change along a ray through empty space, so a pixel reports the radiance leaving the first surface point its ray meets, in the direction back toward the camera. What is that radiance? Take the simplest surface, a diffuse (Lambertian) one, and ask how much light arrives and how much leaves.

How much arrives. Parallel light from direction l (a unit vector pointing at the lamp) carries power E0 per metre of cross-section across the beam. A piece of surface of length s, whose unit normal n makes angle θ with l, presents only a cross-section s·cos θ to the beam, so it catches E0·s·cos θ: per metre of surface, E0(n·l). The cosine is foreshortening and nothing more. A surface facing away (n·l < 0) catches nothing.

How much leaves. A fraction a, the albedo (one number per colour channel), is sent back, and "diffuse" means that it leaves with the same radiance in every direction. (Less power goes toward grazing directions, but the surface also looks narrower from there, by the same cosine.) The light that arrives from everywhere else, sky and neighbouring surfaces, is lumped into one constant, amb. The colour of a surface point is therefore (in units where a white surface facing the lamp reads 1)

I = a · ( amb + (1 − amb) · max(0, n·l) ),   amb = 0.3,   l = (−0.57, 0.82)

(l points up and to the left in the top-down views.) Two things stand out. First, no direction of observation appears: every camera that sees the point reads the same colour. Second, the colour is a product of the paint a, carried by the surface, and the shading I(n), set by how the surface tilts toward the lamp. The shading is not a detail: its factor runs from 1.00 where the surface faces the lamp (θ = 0°), through 0.91 at 30° and 0.65 at 60°, to 0.30 from 90° on, a factor of 3.3. Shadows, interreflection, gloss and transparency are ignored here (Computer Graphics 07 and 08 do them properly).

3 · The forward model, and the whole exam

Now the photograph that a model predicts can be written down. A scene is a set of parameters θ: a shape (a signed distance function), an albedo for every surface point, the lamp l and amb. For a camera and one of its pixels, with ray ρ,

R(θ; camera)pixel = a(x*) · I( n(x*) ),   x* = the first point where ρ meets the surface (the background colour if it never does)

Two steps: visibility, which point is nearest along the ray, then shading, the formula above. Visibility is found by casting the ray (sphere tracing on the signed distance field, as in lesson 6) or by rasterising: project every primitive and keep the nearest one per pixel in a depth buffer. Either way each pixel takes the colour of exactly one point, which makes R a hard renderer. It is the forward model, from scene to photographs, and the exam measures how far R(θ) is from photographs it was not fitted to. Run it on a ladder of models; each row adds an ingredient of R:

model (all given the true surface except the first two)held-out PSNR
every pixel the background colour4.6 dB
the nearest stored photograph, no geometry at all14.0 dB
one colour for every surface19.2 dB
true albedo, no light15.6 dB
true light, one albedo for everything21.9 dB
true light, each object its own base colour (no stripes)23.0 dB
colour copied from the nearest data photograph that sees the point34.6 dB
true albedo times true light, which is R(θ)infinite: it is the photograph

Read the ladder. The light matters as much as the paint: with the true shading and one albedo the score is 21.9 dB, but the true albedo without shading scores only 15.6 dB, worse than one flat colour. Per-object base colours add about a decibel, and the stripes are the rest. Colours copied from the data reach 34.6 dB, because of §2: a diffuse point has one colour for every camera, so a pixel of a data photograph that sees the point is a valid sample of its colour for any other camera that sees it. (Lesson 6's copy did not ask who sees the point. This row moves between 27.9 and 34.6 dB as the number of data cameras goes from 6 to 24, depending on how the pixel grids happen to line up.) Once the surface is known, colour is cheap.

But every row was handed the true surface. A phone has no range sensor; the surface has to come from the photographs themselves.

4 · Turn Lambert around: photo-consistency

The copy row works because a point has one colour for every camera. Read that backwards. Suppose we only hypothesise where the surface is. Take a pixel of camera A and a candidate depth z along its ray, which gives a point P(z). Project P into other cameras and read the colours there. If P is on the surface and the cameras can see it, the colours agree. If not, each camera reads some other surface point, and they disagree. So

C(z) = variance of the colours that the cameras read at P(z), camera A included

scores a geometry from photographs alone: no ground truth, no colour model. This is photo-consistency, and sweeping z along one ray is a plane sweep for one pixel. Test it on the 228 pixels of the data cameras that see the statue, with z from 2.5 to 10 m in 1 cm steps, and call a pixel right when the minimum of C is within 10 cm of the truth:

cameras that vote with camera Ahow manypixels rightmedian error
its two neighbours, 30° either side265%3.1 cm
every camera within 90° either side638%18 cm
only the cameras that see the true point3.4 on average94%1.9 cm

Three readings. (a) It works: with the right cameras voting, nearly every minimum is within centimetres. (b) More cameras should help, and they hurt. A camera on the far side of the statue cannot see P; it reads some other surface and spoils the vote. Which cameras see P depends on where the surface is, which is what we are looking for, and the last row cheated by asking the true surface. Visibility is inside the score. (c) It is blind where the photograph is flat. In the quarter of the pixels where colour changes least from one pixel to the next, the valley C ≤ 0.001 is 0.69 m wide; in the steepest quarter, 0.15 m.

Photo-consistency says whether a hypothesis is good, one pixel at a time, and its vote needs the visibility that only a model of the whole scene can supply. Lesson 5 showed the way: gather the unknowns in one model, predict the data, compare, update. The model is now R, and it brings shading and visibility with it.

5 · Fit the model: why a gradient

Gather the unknowns into θ and score them on the data cameras:

E(θ) = (1/N) · Σcameras, pixels | R(θ; camera)pixel − photographpixel |²,   N = 12 × 64 = 768 pixels, mean over the 3 channels

This is the exam's squared error, taken on the data cameras instead of the held-out ones. Why update by gradient and not by search? Count. A search tries n values of each of d parameters, nd renders: 106 for the three numbers of one circle at n = 100, 102000 for a thousand numbers. A gradient costs d + 1 renders by finite differences, and a small multiple of one render (a forward and a backward pass) by reverse-mode differentiation, whatever d is. Only the gradient scales. It is worth having only if it says which way to move, so test the smallest unknown there is.

6 · The wall: a pixel is a cliff in the geometry

Let the model be one circle of radius r, centred at (δ, 0), painted one flat colour so that nothing but visibility can change: a pixel shows that colour if its ray meets the circle, the background otherwise. The cleanest case is a disc of radius 1 m photographed by six cameras on a ring of radius 5 m (48 pixels each, f = 40 px), with a model that is the same disc, its centre off by δ.

Take one pixel. Its ray, with unit direction d, crosses the line z = 0 at a point x0, at an angle whose sine is |dz|. The distance from the circle's centre (δ, 0) to the ray's line is therefore |δ − x0|·|dz|, and the ray meets the circle exactly when that is at most r (the circle is always in front of the camera here):

|δ − x0| ≤ w,   w = r / |dz|

So pixel k shows the circle's colour for δ in an interval of half-width wk around x0,k, and the background colour outside it. Its squared error against the photograph is e¹k in the first case and e⁰k in the second (two fixed numbers), and the loss is a sum of boxes, 1(·) being 1 when its condition holds and 0 otherwise:

E(δ) = (1/N) Σk [ e⁰k + (e¹k − e⁰k) · 1( |δ − x0,k| ≤ wk ) ]

Read it. Between two interval ends nothing in any picture changes, so E is constant: a plateau, with derivative exactly 0. At an interval end a ray grazes the circle, which starts or stops hiding what lies behind it; one pixel changes colour, and E jumps by (e¹ − e⁰)/N: a cliff. Cliffs sit where a ray is tangent to the model, wherever the photographs are; the photographs only set their heights. So the derivative of the loss has two parts. One is the change in the colours of pixels whose visibility stays the same: differentiating the program, by hand or by autodiff, finds it, and for a flat colour it is zero. The other is the impulse at each cliff, the boundary term, and differentiating never sees it, because the switch is a comparison and the derivative of a comparison is zero. Count the cliffs over a sweep of δ from −1.6 m to 1.6 m in steps of one millimetre:

unit disc, 6 cameras (the clean case)the statue's photographs, 12 cameras, r = 1.15 m
cliffs inside the sweep228500
distinct values of E over the sweep59472
1 mm steps on which E does not change96.4%85.3%

(Opposite cameras of the disc see the same rays, so its 228 cliffs sit at only 114 places, and E, which only counts how many pixels are wrong, takes few values. The statue's model is a circle of its mean radius, painted the average colour of the object pixels.) Zoom out and the staircase becomes a bowl, because hundreds of tiny steps add up to a large-scale slope. That slope is real, but it exists only across many plateaus, and measuring it costs two renders per parameter. The derivative a machine computes in one backward pass for millions of parameters is the one at scale zero, and there it is zero.

Anti-aliasing. A pixel is a footprint, not a ray. With S sub-rays per pixel, each cliff becomes S cliffs 1/S as tall: a finer staircase. With 16 sub-rays, 61% of the positions on the sweep still have an exactly zero derivative at step 0.1 mm, against 97% with one ray. Only the limit S → ∞ removes the steps: the pixel shows the circle with weight c, the fraction of its footprint the circle covers, which is the overlap of the pixel [i, i+1] with the circle's image, the interval between its two tangent rays. As δ changes, each cliff becomes a ramp one footprint wide (0.107 m at 6 m for a camera looking across the motion). That limit is the orange curve in the widget.

One circle against the statue's photographs
Left: the scene from above, the model circle (amber), the 12 data cameras (grey) and one held-out camera (blue). Right: its photograph, the circle as the first-hit rule renders it, where they disagree, and the exam score of the model (mean PSNR over all 24 held-out cameras) against the offset. Bottom: the training loss E against the offset, first hit (grey) and exactly anti-aliased (orange), in a window you can shrink; the band is the plateau under δ. The buttons run 100 steps of Adam on δ, using a 0.1 mm finite difference of each loss.
E, first hit
—
dE/dδ, first hit
—
E, anti-aliased
—
dE/dδ, anti-aliased
—
exam: held-out PSNR
—
plateau under δ
—
pixels that feel δ
—
last descent ends at
—
Show the core JS
o.x0.push(r.ox - r.oz * r.dx / r.dz); o.w.push(R / Math.abs(r.dz)); o.e1.push(sq(MC, ph[v], i)); o.e0.push(sq(BG, ph[v], i));
...
function Ehard(d) { var e = 0; for (var k = 0; k < N; k++) e += Math.abs(d - A.x0[k]) <= A.w[k] ? A.e1[k] : A.e0[k]; return e / N; }
...
var t = FL.toCam(cam, d, 0), r = Math.hypot(t.xc, t.zc), phi = Math.atan2(t.xc, t.zc), al = Math.asin(R / r);
return [cam.W / 2 + cam.f * Math.tan(phi - al), cam.W / 2 + cam.f * Math.tan(phi + al)];
...
var k = v * W + i, c = Math.max(0, Math.min(i + 1, s[1]) - Math.max(i, s[0]));
e += A.e0[k] + (A.e1[k] - A.e0[k] - Q) * c + Q * c * c; if (c > 0 && c < 1) edges++;
...
function slope(f, d) { return (f(d + H1) - f(d - H1)) / (2 * H1); }

What to try. The page opens with the model 1.5 m to the left of the statue (δ = −1.5 m), the whole sweep in view, and held-out camera 17, whose image runs left to right like the map. At this scale the two losses are one bowl: 0.2012 for the first-hit renderer, 0.1981 anti-aliased, and the model's exam score is 7.5 dB. The first-hit derivative reads exactly 0; the anti-aliased one reads −0.0675 per metre, and differencing the first-hit loss across ±0.1 m, a window holding 32 cliffs, gives −0.0691: the bowl's slope is real, and exists only at that scale. Press descend (first hit): nothing moves, and nothing would from 97% of the starting positions on the slider. Press descend (anti-aliased): 100 steps of Adam walk the model to −0.133 m. Now set magnify to 4, a 25 mm window. The plateau under δ is 8.08 mm wide, between cliffs at −1.5058 m and −1.4977 m; step δ a millimetre at a time across the right-hand one and E drops by 0.00078, one pixel in 768 changing colour. The orange curve is a smooth ramp, and "pixels that feel δ" says why: 0 for the first-hit renderer, 24 of 768 for the anti-aliased one, the pixels that contain the circle's edge. The violet exam curve has cliffs of its own and disagrees with the training loss about the best offset: the exam peaks at +0.085 m with 9.1 dB (one flat circle is a poor model of this scene), the anti-aliased training loss is lowest at −0.067 m. Finally, in the full window, start from δ = −2.0: the anti-aliased descent stops at −0.461 m, in a shallow valley 0.012 above that minimum. A gradient exists there, and it is local.

Road not taken · smooth the silhouette
Edge-sampling renderers, and rasterisers with approximate gradients at silhouettes, keep the first hit and add a correction at each shape's boundary. The orange curve is the exact version of that idea for a circle, and the model's own edge gets a slope, and a mesh of known topology can be fitted with it. But count what carries the slope: only the pixels that contain an edge, 24 of 768 at the opening state. A surface that no ray grazes gets none. In the statue's scene the ball is hidden behind the statue from 2 of the 12 data cameras: slide it by 0.3 m and not one pixel of those two photographs changes, while another camera's changes in up to 43 pixels. Nor can a boundary correction create an edge where the model has no surface. Estimating the gradient of the discrete choice by sampling is unbiased but very noisy; matching features instead of pixels is lesson 5's world, with no renderer to descend.
What this lesson did not do
It did not say how to represent a scene that is not one circle (lesson 9), nor how to make the rendering error smooth (lesson 8). The ladder was handed the true surface, and the best row of photo-consistency asked the true surface who sees what; nothing here estimates visibility together with geometry. Lambert is the whole reflectance model: shadows, interreflection, gloss and transparency make a point look different to different cameras and break photo-consistency. PSNR is a statement about squared colour error, not about what looks right, and the cameras are perfect: no noise, exposure or lens.

Common mistakes / failure modes

"a pixel has the colour of the object"
A pixel is albedo times shading at the first hit. The true albedo without light scores 15.6 dB, worse than one flat colour at 19.2 dB (§2, §3).
"a model that reproduces its photographs is a good model"
Storing them scores infinity on the cameras they came from and 14.0 dB on held-out ones. The exam is about cameras not used (§1).
"more cameras make photo-consistency better"
Six cameras voting blindly get 38% of the pixels right, two neighbours 65%, and 94% only when visibility is known (§4).
"the loss has no gradient because the code is not differentiable"
The function itself is piecewise constant: between cliffs its derivative is zero in the mathematical sense, however it is computed (§6).
"anti-aliasing fixes the staircase"
Exact anti-aliasing turns the cliffs of the model's own edge into ramps. Sixteen sub-rays leave 61% of positions with a zero derivative, and even the exact limit rests on 24 of 768 pixels (§6).

Checkpoint exercise

Try it
A camera at (0, 6) looks straight down (heading −90°), with f = 56 px and 64 pixels. For which offsets δ does the centre ray of pixel 37 (u = 37.5) meet a disc of radius 1 m centred at (δ, 0)? How far apart are the cliffs of pixels 37 and 38? Answer: the ray's direction is proportional to (−5.5/56, −1), so it crosses z = 0 at x0 = −6·5.5/56 = −0.589 m, and |dz| = 1/√(1 + (5.5/56)²) = 0.9952, so w = 1/0.9952 = 1.005 m. The pixel sees the disc for −1.594 ≤ δ ≤ 0.416 m; those are its two cliffs. The next pixel's ray crosses z = 0 another 6/56 = 0.107 m along, so the cliffs of neighbouring pixels are about that far apart: one footprint of a pixel at 6 m.

Where this points next

Rendering is now a function we can write down and score: first hit, then Lambert, compared with photographs from cameras the model never saw. With the true surface and colours copied from the data it earns 34.6 dB, and photographs can grade a geometry with no ground truth. It is also a function we cannot yet descend: on the statue's photographs the loss of the simplest model is constant between 500 cliffs inside a 3.2 m sweep, and the first-hit derivative is exactly zero from 97% of the positions on the slider (the clean disc of §6 is where lesson 8 begins). Exact anti-aliasing gives the model's own edge a slope that rests on 24 of 768 pixels, and a surface that no ray grazes, hidden by another surface or not there yet, gets nothing. The cause is the first-hit rule itself: which surface a ray meets first is a discrete choice. What would a renderer look like whose pixels change smoothly with the scene, and still reduce to the hard one when the scene is solid?

Takeaway
A pixel measures the radiance leaving the first surface its ray meets; for a diffuse surface that is a·(amb + (1 − amb)·max(0, n·l)), the same from every camera. A model is therefore a forward model R(θ), first hit and then shading, and the exam is the PSNR between R(θ) and photographs from held-out cameras: given the true surface, one colour scores 19.2 dB, shading 21.9 dB, and colours copied from the data 34.6 dB, so colour is cheap once the surface is known. Because a diffuse point has one colour, photographs grade geometry with no ground truth (photo-consistency), but only through the cameras that can see the point: visibility is inside the score. Fitting θ by descent on the rendering error is the way to scale, and for first-hit rendering it meets a wall: the loss is a sum of boxes, exactly constant between cliffs where rays graze the model, with derivative exactly 0. Exact anti-aliasing turns the cliffs of a model's own edge into one-pixel ramps and says nothing about surfaces that are hidden or missing.

Interview prompts

Companion reads: Computer Graphics · 07 Light and the rendering equation (radiance), Computer Graphics · 08 Local shading and PBR (Lambert and what replaces it), Computer Graphics · 06 Sampling and anti-aliasing (the pixel as a footprint), and Computer Vision · 18 Image restoration (what PSNR does and does not measure).