all_lessons/3D Vision/01 · A pixel is a raylesson 1 / 14

A pixel is a ray

The Computer Vision track taught the camera as a function: a 3D point goes in, a pixel comes out. Vision has to run it backwards, because pixels are what we have and the scene is what we want. This lesson asks what one pixel keeps and what it throws away, and finds a single missing number, the distance. That number is also exactly what a different viewpoint needs: slide the camera sideways and every pixel moves by an amount that depends on it. So the exam this whole track is graded by, predict the photograph from a camera nobody used, and the measurement it demands, how far away is each thing?, are one problem, and two photographs are enough to start on it.

The thesis, here
Every lesson in this track is graded by one exam: given photographs of a scene from some cameras, predict the photograph from a camera that was not among them. Run the exam on the simplest model there is, the first photograph itself, and it fails in one specific way: it does not know how far away anything is. Each later representation (points, surfaces, fog, fields, splats, learned priors) is a better answer to a failure the previous one left standing.
Linear position
Forced by: The Computer Vision track taught the camera as a function: a 3D point goes in, a pixel comes out. Vision has to run it backwards — the pixels are what we have and the scene is what we want. What does one pixel tell us about the point that made it, and what does it leave out?
New idea: a pixel is a ray: it records the direction in which a point lies and nothing about how far away it is. That missing number is the same one a new viewpoint needs, because sliding the camera sideways by b moves a pixel by f·b/Z. The exam "predict the view from a new camera" and the measurement "recover the depth" are one problem.
Forces next: A pixel pins down a direction and nothing else: every point along its ray lands on the same pixel, so depth is the one number per pixel that the image destroys. It is also exactly the number a new viewpoint needs, because moving the camera sideways by b shifts a pixel by f·b/Z. So the shift between two views measures depth. How precisely can two rays pin down a point, and what limits that precision?
The plan
Five moves. (1) Write the camera as a function, count what goes in and what comes out, and find that a pixel is a ray, not a point. (2) Ask what a different camera needs from this photograph, find that it needs the number the pixel lost, and compute how much: f·b/Z for a slide, nothing at all for a turn. (3) Run the exam with a guessed distance and with the true one. (4) See what is still missing even with perfect depths. (5) Pin down the small world that most of the next thirteen lessons share (two leave it for true 3D), and say what it leaves out of the real one.

1 · A camera in Flatland, and what a pixel keeps

A real camera turns a 3D world into a 2D image. This track works one dimension down, in Flatland: the world is a plane (x to the right, z up), a camera is a point with a heading, and its photograph is a single row of pixels. Everything in this lesson is the same in three dimensions with one more coordinate; §5 says what the missing dimension costs. A camera has a focal length f, measured in pixels, and W pixels, with the optical axis through the middle.

Describe a point in the camera's own frame: X metres to the right of the axis and Z metres straight ahead. Z is the point's depth, measured along the axis, not along the ray to the point. The point lands on pixel

u = W/2 + f·X / Z

Count what goes in and what comes out. Two numbers, (X, Z), go in; one number, u, comes out. A map from two numbers to one cannot be run backwards, and the reason is easy to see by asking which points land on the same pixel u0. They satisfy X/Z = (u0 − W/2)/f, a constant k, so X = k·Z: a straight line through the camera centre. That line is the pixel ray. Every point on it paints the same pixel. The pixel therefore records one number about where the point is, the ray's direction, and discards the other, the position along it. In three dimensions the count is three numbers in and two out, and the answer is again a ray.

The discarded number is not a technicality. A crate 1 m wide at 2.5 m, a billboard 2 m wide at 5 m and a wall 4 m wide at 10 m, all centred on the same ray, cover exactly the same pixels, because the width on the sensor is f·w/Z and all three have the same w/Z:

ObjectWidth wDepth ZWidth on the sensor, f·w/Z (f = 44 px)
crate1 m2.5 m17.6 px
billboard2 m5 m17.6 px
wall4 m10 m17.6 px

Nothing in the photograph can tell the three apart, so no amount of staring at one photograph recovers Z: the information is not in it. What the photograph does hold, for every pixel, is the direction and the colour of the first thing along that direction.

2 · What a different camera needs from this photograph

Now state the exam. We hold photograph A, taken by a camera whose position we know. We are asked for photograph B, taken by a different camera whose position we also know. Each pixel of A already has its colour. What it lacks is where that colour belongs in B, and that is a question about the point, which lies somewhere on A's pixel ray.

Work the simplest move, a slide: B is A moved b metres along its own right-hand direction, heading unchanged. Moving sideways changes a point's X and leaves its Z alone, so a point at (X, Z) in A's frame is at (X − b, Z) in B's, and

uB = W/2 + f·(X − b)/Z = uA − f·b / Z

The pixel moves left by d = f·b/Z, the parallax (stereo vision calls it disparity). It grows with the baseline b and the focal length f and shrinks with the depth, so near things move a lot and far things hardly move. For f = 44 px and b = 0.5 m:

Depth Z2.5 m5 m10 m40 mvery far
Shift d = f·b/Z8.8 px4.4 px2.2 px0.55 px0

Read this equation in both directions, because the whole track turns on it. Forwards: if we know Z for a pixel, we know where it goes in B, so we can predict B. Backwards: if we can see how far a feature moved between A and B, we know Z = f·b/d. The exam question and the depth measurement are the same equation read in opposite directions.

The other simple move is a turn: B is A rotated by θ about the camera centre. A pixel whose ray makes angle φ with the axis, tan φ = (uA − W/2)/f, now makes angle φ − θ, so

uB = W/2 + f·tan(φ − θ)

and Z has disappeared. A turn relabels the pixels without asking how far away anything is. That gives three facts about the exam at once: a turn can be predicted with no depth, a turn tells us no depth, and so distance becomes visible only when the camera moves. A tripod panorama stitches without any depth; a hand-held pair measures it.

Camera B is A…Does the shift depend on depth?Does B show things A could not see?Can B's image measure depth?
slid sideways by byes, f·b/Zyes, behind near things (§4)yes, Z = f·b/d
turned by θnoonly beyond the frame edgeno

3 · The exam, run

To run the exam we need a way to build the prediction. Take each of A's 120 pixels as a unit-wide interval, shift it by its own f·b/Z using whatever depth we choose, and paint it into B; where two pixels land on the same place the nearer wins, which is what a depth buffer does. Two numbers score the result. Accuracy is the PSNR between the painted pixels and the real B, counting only pixels that were painted. Unseen is the share of B's pixels that no pixel of A reached. The widget lets you choose the depth: one guessed value for every pixel, which is all a single photograph can honestly offer, or each pixel's true depth, which is what a measurement would supply.

Predict the other eye
Photograph A (top strip) is what we have. The strips below are camera B's real photograph, the prediction built from A, and where they disagree. Slide the baseline, switch between sliding and turning, and choose the depth the prediction may use. The world view shows both cameras and the three things in the street. Each tick on A's strip is joined to where that same point lands in B, and the number is how far it moved.
crate moved
—
ball moved
—
tower moved
—
accuracy (painted pixels)
—
unseen by A
—
Show the core JS
function predict(imgA, camB, depth) {
  var n = camB.W, m = n * SS, zbuf = new Float64Array(m).fill(Infinity), R = new Float64Array(m), G = new Float64Array(m), Bl = new Float64Array(m), hit = new Uint8Array(m), i, k;
  for (i = 0; i < A.W; i++) {
    var Z = depth[i], P0 = FL.backproject(A, i, Z), P1 = FL.backproject(A, i + 1, Z);
    var q0 = FL.project(camB, P0.x, P0.z), q1 = FL.project(camB, P1.x, P1.z);
    var lo = Math.min(q0.u, q1.u), hi = Math.max(q0.u, q1.u), zb = (q0.zc + q1.zc) / 2;
    for (k = Math.max(0, Math.ceil(lo * SS - 0.5)); k <= Math.min(m - 1, Math.floor(hi * SS - 0.5)); k++)
      if (zb < zbuf[k]) { zbuf[k] = zb; R[k] = imgA.r[i]; G[k] = imgA.g[i]; Bl[k] = imgA.b[i]; hit[k] = 1; }
  }
  ...
}

What to try. Start as the page opens: slide, b = 0.5 m, one guessed depth Z0 = 6 m. The readouts give the three shifts, 9.0 px for the crate, 4.3 px for the ball and 2.2 px for the tower, which are f·b/Z at their true depths (about 2.5, 5 and 10 m). A single guess moves everything by the same amount, so at most one of them can be right, and the prediction scores 13.3 dB with 3.1% unseen. Drag Z0: the best single guess is about 4.8 m, near the ball, and even then the score is only 13.5 dB, since no single number is right for all three things. Now choose each pixel's true Z. The three things land where B really has them and the accuracy jumps to 33.5 dB, but the unseen share grows from 3.1% to 10.6%. Drag the baseline with true depths on: the accuracy wobbles between about 26 and 34 dB (an edge cannot be moved by a fraction of a pixel, so how well it lands depends on where it falls on the pixel grid), while the unseen share rises steadily. Finally switch to turned on the spot and set the baseline to 0.5, which now means a 5° turn. The shifts are no longer ordered by depth (the crate moves 7.0 px, the ball 3.9 and the tower 4.2): a turn moves a pixel by an amount that depends on where it sits in the image, not on how far away it is. The accuracy is 30 dB, and it is the same for every Z0 and for the true depths. A turn does not need the depth, and does not reveal it.

Road not taken · guess the distance from the picture
A neural network can look at A and output a plausible Z for every pixel, from the sizes of familiar things and the way floors recede. That estimate can be good enough to try the exam, but it is a guess made from experience, not a measurement the pixels contain: the table above shows a crate and a wall painting identical pixels, and a learned guess will be wrong exactly when the scene is not like the experience. Lesson 11 takes this road on purpose and says what it can and cannot promise. This track first builds the measurement.

4 · What no depth map can supply

With true depths the three things landed in the right place, and yet 10.6% of B's pixels were never painted. Look at where they are in the world view: just to the right of each near object. A's pixels there showed the object, so they said nothing about what sits behind it, and B, looking from a little further right, can see behind it.

The size of the gap follows from the same equation. A near object at depth Zn in front of a far background at Zf moves by f·b/Zn, the background by f·b/Zf, and a strip opens between them, in pixels, of width

f·b·(1/Zn − 1/Zf)

For the crate, whose front face is at 2.45 m, with the background at 40 m and b = 0.5, that is 8.4 px of B that A has no information about, and it grows in proportion to the baseline. So the exam has two kinds of failure, and they are fixed by different means. Unmeasured: the geometry is in A and we have not recovered it; the next lessons measure it, from lesson 2 on. Unseen: A does not contain it at all; only more viewpoints (lessons 6–10) or knowledge from outside the photographs (lessons 11–13) can supply it. A better depth estimate never closes the second kind of gap.

5 · The world these lessons share

Every widget in the track runs in Flatland, except those of lessons 3 and 5, which need true 3D and say so. Flatland is one small shared engine (flatland.js) that does the ray tracing, cameras and drawing, and a short list of scenes. This lesson used the street: a crate, a ball and a tower at about 2.5, 5 and 10 m, all textured with stripes so that features can be matched later. The engine is deterministic, so a number quoted in a lesson is the number you will see.

IdeaSurvives in Flatland?Why
rays, parallax, triangulation, error lawsyesthe algebra is the same with one coordinate fewer
fusion, volume rendering, splatting, priorsyesnone of them depends on the dimension
epipolar constraint (lesson 5)noin a plane any two rays meet, so a match carries no constraint; that lesson works in true 3D
the curvature of rotations (lesson 3)norotations in a plane commute and form a circle; that lesson works in true 3D
a vertical axis (lesson 12)noa plane has no "up" to collapse, so that lesson treats the bird's-eye view as arithmetic, not as a widget
What this lesson did not do
It did not measure anything: the true depths in the widget are read off the renderer, which cheats. It did not say how precisely a shift can be turned into a depth, or what sets that limit. It did not address cameras whose positions are unknown. Lesson 2 measures depth from a shift and finds its error law; lesson 3 describes a camera's pose so that a computer can estimate it, and lessons 4 and 5 estimate it, from scans and from images.

Common mistakes / failure modes

"a pixel has a depth"
A pixel has a ray. Depth belongs to the point in the world; a "depth map" is a measurement or an estimate attached to pixels afterwards (§1).
"depth is the distance to the camera"
Here depth is Z, measured along the optical axis. The distance along the ray is longer by a factor 1/cos φ. The formulas u = W/2 + fX/Z and d = fb/Z use Z (§1).
"turning the camera reveals depth too"
It gives a new photograph and no new information about distance: the shift is the same at every depth. Only a change of position makes depth visible (§2).
"parallax measures how far away things look"
The shift is inversely proportional to depth: near things move most, far things barely, so equal steps in depth give unequal shifts (§2 table).
"if I knew every depth the new view would be perfect"
Pixels behind near objects were never seen. A perfect depth map still leaves a strip of width f·b·(1/Zn − 1/Zf) at every edge (§4).
"two cameras are needed"
Two viewpoints are needed. One camera that moves will do, provided the scene holds still (lesson 14 relaxes that).

Checkpoint exercise

Try it
A camera with f = 44 px looks at a post 4 m away and at a distant mountain. You slide it 0.2 m to the right. By how many pixels does each move, and in which direction? What would a post at 2 m do? Then turn the camera by 5° instead: how far does the post move, and how far the mountain? Answer: with d = f·b/Z the post moves 2.2 px to the left and the mountain by about zero; a post at 2 m moves 4.4 px, twice as far for half the distance. A 5° turn moves points on the optical axis by f·tan 5° = 3.8 px for both the post and the mountain, because a turn ignores depth.

Where this points next

Two photographs from cameras we know turn the lost number into a measurement: a feature that moved by d pixels lies at depth Z = f·b/d. But d is read off a pixel grid, to a fraction of a pixel at best, and it is the denominator. Suppose it is off by half a pixel. For the crate at 2.5 m, d = 8.8 px and the estimate lands between 2.37 m and 2.65 m. For the tower at 10 m, d = 2.2 px and the same half pixel puts it anywhere between 8.1 m and 12.9 m. The error is not the same size for near and far: it grows with distance. How precisely can two rays pin down a point, and what limits that precision?

Takeaway
A pixel is a ray: the camera map u = W/2 + f·X/Z sends two numbers to one, so every point on a line through the camera centre paints the same pixel, and the depth Z is the one number the photograph destroys. A new viewpoint needs that number: a slide of b shifts a pixel by f·b/Z, while a turn shifts it by an amount independent of Z, so distance is visible only when the camera moves. The exam "predict the photograph from a new camera" therefore contains the measurement "recover the depth", and it has two failure modes: unmeasured (fixable by measuring) and unseen (not in the photograph at all). Everything that follows is a better answer to one of them.

Interview prompts

Companion reads: Computer Vision · 04 Cameras and projection geometry (the pinhole model this lesson starts from), Computer Vision · 05 Multi-view, depth and SLAM (the standard presentation of disparity and epipolar geometry), and Computer Graphics · 02 Transforms and spaces (the same maps run forwards).