A pixel is a ray
The Computer Vision track taught the camera as a function: a 3D point goes in, a pixel comes out. Vision has to run it backwards, because pixels are what we have and the scene is what we want. This lesson asks what one pixel keeps and what it throws away, and finds a single missing number, the distance. That number is also exactly what a different viewpoint needs: slide the camera sideways and every pixel moves by an amount that depends on it. So the exam this whole track is graded by, predict the photograph from a camera nobody used, and the measurement it demands, how far away is each thing?, are one problem, and two photographs are enough to start on it.
New idea: a pixel is a ray: it records the direction in which a point lies and nothing about how far away it is. That missing number is the same one a new viewpoint needs, because sliding the camera sideways by b moves a pixel by f·b/Z. The exam "predict the view from a new camera" and the measurement "recover the depth" are one problem.
Forces next: A pixel pins down a direction and nothing else: every point along its ray lands on the same pixel, so depth is the one number per pixel that the image destroys. It is also exactly the number a new viewpoint needs, because moving the camera sideways by b shifts a pixel by f·b/Z. So the shift between two views measures depth. How precisely can two rays pin down a point, and what limits that precision?
1 · A camera in Flatland, and what a pixel keeps
A real camera turns a 3D world into a 2D image. This track works one dimension down, in Flatland: the world is a plane (x to the right, z up), a camera is a point with a heading, and its photograph is a single row of pixels. Everything in this lesson is the same in three dimensions with one more coordinate; §5 says what the missing dimension costs. A camera has a focal length f, measured in pixels, and W pixels, with the optical axis through the middle.
Describe a point in the camera's own frame: X metres to the right of the axis and Z metres straight ahead. Z is the point's depth, measured along the axis, not along the ray to the point. The point lands on pixel
u = W/2 + f·X / Z
Count what goes in and what comes out. Two numbers, (X, Z), go in; one number, u, comes out. A map from two numbers to one cannot be run backwards, and the reason is easy to see by asking which points land on the same pixel u0. They satisfy X/Z = (u0 − W/2)/f, a constant k, so X = k·Z: a straight line through the camera centre. That line is the pixel ray. Every point on it paints the same pixel. The pixel therefore records one number about where the point is, the ray's direction, and discards the other, the position along it. In three dimensions the count is three numbers in and two out, and the answer is again a ray.
The discarded number is not a technicality. A crate 1 m wide at 2.5 m, a billboard 2 m wide at 5 m and a wall 4 m wide at 10 m, all centred on the same ray, cover exactly the same pixels, because the width on the sensor is f·w/Z and all three have the same w/Z:
| Object | Width w | Depth Z | Width on the sensor, f·w/Z (f = 44 px) |
|---|---|---|---|
| crate | 1 m | 2.5 m | 17.6 px |
| billboard | 2 m | 5 m | 17.6 px |
| wall | 4 m | 10 m | 17.6 px |
Nothing in the photograph can tell the three apart, so no amount of staring at one photograph recovers Z: the information is not in it. What the photograph does hold, for every pixel, is the direction and the colour of the first thing along that direction.
2 · What a different camera needs from this photograph
Now state the exam. We hold photograph A, taken by a camera whose position we know. We are asked for photograph B, taken by a different camera whose position we also know. Each pixel of A already has its colour. What it lacks is where that colour belongs in B, and that is a question about the point, which lies somewhere on A's pixel ray.
Work the simplest move, a slide: B is A moved b metres along its own right-hand direction, heading unchanged. Moving sideways changes a point's X and leaves its Z alone, so a point at (X, Z) in A's frame is at (X − b, Z) in B's, and
uB = W/2 + f·(X − b)/Z = uA − f·b / Z
The pixel moves left by d = f·b/Z, the parallax (stereo vision calls it disparity). It grows with the baseline b and the focal length f and shrinks with the depth, so near things move a lot and far things hardly move. For f = 44 px and b = 0.5 m:
| Depth Z | 2.5 m | 5 m | 10 m | 40 m | very far |
|---|---|---|---|---|---|
| Shift d = f·b/Z | 8.8 px | 4.4 px | 2.2 px | 0.55 px | 0 |
Read this equation in both directions, because the whole track turns on it. Forwards: if we know Z for a pixel, we know where it goes in B, so we can predict B. Backwards: if we can see how far a feature moved between A and B, we know Z = f·b/d. The exam question and the depth measurement are the same equation read in opposite directions.
The other simple move is a turn: B is A rotated by θ about the camera centre. A pixel whose ray makes angle φ with the axis, tan φ = (uA − W/2)/f, now makes angle φ − θ, so
uB = W/2 + f·tan(φ − θ)
and Z has disappeared. A turn relabels the pixels without asking how far away anything is. That gives three facts about the exam at once: a turn can be predicted with no depth, a turn tells us no depth, and so distance becomes visible only when the camera moves. A tripod panorama stitches without any depth; a hand-held pair measures it.
| Camera B is A… | Does the shift depend on depth? | Does B show things A could not see? | Can B's image measure depth? |
|---|---|---|---|
| slid sideways by b | yes, f·b/Z | yes, behind near things (§4) | yes, Z = f·b/d |
| turned by θ | no | only beyond the frame edge | no |
3 · The exam, run
To run the exam we need a way to build the prediction. Take each of A's 120 pixels as a unit-wide interval, shift it by its own f·b/Z using whatever depth we choose, and paint it into B; where two pixels land on the same place the nearer wins, which is what a depth buffer does. Two numbers score the result. Accuracy is the PSNR between the painted pixels and the real B, counting only pixels that were painted. Unseen is the share of B's pixels that no pixel of A reached. The widget lets you choose the depth: one guessed value for every pixel, which is all a single photograph can honestly offer, or each pixel's true depth, which is what a measurement would supply.
What to try. Start as the page opens: slide, b = 0.5 m, one guessed depth Z0 = 6 m. The readouts give the three shifts, 9.0 px for the crate, 4.3 px for the ball and 2.2 px for the tower, which are f·b/Z at their true depths (about 2.5, 5 and 10 m). A single guess moves everything by the same amount, so at most one of them can be right, and the prediction scores 13.3 dB with 3.1% unseen. Drag Z0: the best single guess is about 4.8 m, near the ball, and even then the score is only 13.5 dB, since no single number is right for all three things. Now choose each pixel's true Z. The three things land where B really has them and the accuracy jumps to 33.5 dB, but the unseen share grows from 3.1% to 10.6%. Drag the baseline with true depths on: the accuracy wobbles between about 26 and 34 dB (an edge cannot be moved by a fraction of a pixel, so how well it lands depends on where it falls on the pixel grid), while the unseen share rises steadily. Finally switch to turned on the spot and set the baseline to 0.5, which now means a 5° turn. The shifts are no longer ordered by depth (the crate moves 7.0 px, the ball 3.9 and the tower 4.2): a turn moves a pixel by an amount that depends on where it sits in the image, not on how far away it is. The accuracy is 30 dB, and it is the same for every Z0 and for the true depths. A turn does not need the depth, and does not reveal it.
4 · What no depth map can supply
With true depths the three things landed in the right place, and yet 10.6% of B's pixels were never painted. Look at where they are in the world view: just to the right of each near object. A's pixels there showed the object, so they said nothing about what sits behind it, and B, looking from a little further right, can see behind it.
The size of the gap follows from the same equation. A near object at depth Zn in front of a far background at Zf moves by f·b/Zn, the background by f·b/Zf, and a strip opens between them, in pixels, of width
f·b·(1/Zn − 1/Zf)
For the crate, whose front face is at 2.45 m, with the background at 40 m and b = 0.5, that is 8.4 px of B that A has no information about, and it grows in proportion to the baseline. So the exam has two kinds of failure, and they are fixed by different means. Unmeasured: the geometry is in A and we have not recovered it; the next lessons measure it, from lesson 2 on. Unseen: A does not contain it at all; only more viewpoints (lessons 6–10) or knowledge from outside the photographs (lessons 11–13) can supply it. A better depth estimate never closes the second kind of gap.
5 · The world these lessons share
Every widget in the track runs in Flatland, except those of lessons 3 and 5, which need true 3D and say so. Flatland is one small shared engine (flatland.js) that does the ray tracing, cameras and drawing, and a short list of scenes. This lesson used the street: a crate, a ball and a tower at about 2.5, 5 and 10 m, all textured with stripes so that features can be matched later. The engine is deterministic, so a number quoted in a lesson is the number you will see.
| Idea | Survives in Flatland? | Why |
|---|---|---|
| rays, parallax, triangulation, error laws | yes | the algebra is the same with one coordinate fewer |
| fusion, volume rendering, splatting, priors | yes | none of them depends on the dimension |
| epipolar constraint (lesson 5) | no | in a plane any two rays meet, so a match carries no constraint; that lesson works in true 3D |
| the curvature of rotations (lesson 3) | no | rotations in a plane commute and form a circle; that lesson works in true 3D |
| a vertical axis (lesson 12) | no | a plane has no "up" to collapse, so that lesson treats the bird's-eye view as arithmetic, not as a widget |
Common mistakes / failure modes
Checkpoint exercise
Where this points next
Two photographs from cameras we know turn the lost number into a measurement: a feature that moved by d pixels lies at depth Z = f·b/d. But d is read off a pixel grid, to a fraction of a pixel at best, and it is the denominator. Suppose it is off by half a pixel. For the crate at 2.5 m, d = 8.8 px and the estimate lands between 2.37 m and 2.65 m. For the tower at 10 m, d = 2.2 px and the same half pixel puts it anywhere between 8.1 m and 12.9 m. The error is not the same size for near and far: it grows with distance. How precisely can two rays pin down a point, and what limits that precision?
Interview prompts
- Why can a single image not tell you how far away something is? (§1 — the camera map sends two numbers to one; all points on a line through the centre land on the same pixel.)
- Derive the parallax of a point when the camera slides sideways by b. (§2 — only X changes, to X − b, so uB = uA − f·b/Z.)
- Why does rotating a camera in place give no depth information, while translating it does? (§2 — a turn changes each ray's angle by the same amount whatever its distance; a slide changes it by an amount that depends on distance.)
- Two objects at 2 m and 20 m; the camera slides 10 cm. Compare the shifts. (§2 — shift is proportional to 1/Z, so the near one moves ten times as far.)
- Why is "depth" measured along the optical axis rather than along the ray? (§1 — then X/Z gives the pixel directly and the parallax is exactly f·b/Z.)
- If you had a perfect depth map for photograph A, could you render the view from a camera 1 m to the right perfectly? (§4 — no: pixels behind near objects were never seen, and the missing strip grows with the baseline.)
- What does Flatland preserve of the real problem, and what does it lose? (§5 — rays, parallax, triangulation, fusion and rendering carry over; the epipolar constraint and the curvature of rotations do not, so those lessons use true 3D.)
Companion reads: Computer Vision · 04 Cameras and projection geometry (the pinhole model this lesson starts from), Computer Vision · 05 Multi-view, depth and SLAM (the standard presentation of disparity and epipolar geometry), and Computer Graphics · 02 Transforms and spaces (the same maps run forwards).