A second ray pins the point
Lesson 1 ended on Z = f·b/d and a warning: d is read off a pixel grid and sits in the denominator, so half a pixel moves the tower by metres and the crate by centimetres. This lesson makes the warning exact. Two pixel rays are a 2×2 linear system whose conditioning is the angle between them, about b/Z, so a pixel error costs a depth error that grows as Z² and a sideways error that grows only as Z: a triangulated point is a needle along the rays. Depth maps from either source back-project into clouds that do not fit together until we know where the sensors were.
New idea: two pixel rays are a 2×2 linear system whose conditioning is the triangulation angle θ ≈ b/Z, so pixel noise becomes an error needle: σZ = √2·Z²σu/(f·b) in depth, σX = Z·σu/f sideways. Precision has to be bought with baseline, with a smaller σu, or with a sensor that does not triangulate, and each purchase has a price.
Forces next: Two rays pin a point only because we were told where the second camera is. In practice that is the unknown: a robot, a phone or a scanner moves between measurements, and every depth map lives in its own sensor frame. To compare or merge frames we must write down a rotation and a translation, and a rotation cannot be added, averaged or stepped along like a vector. How do we describe a rigid motion so that a computer can estimate it?
1 · Two rays, two unknowns
Lesson 1 closed on two numbers. A half-pixel error in d leaves the crate anywhere in an interval 0.3 m long and the tower in one 4.8 m long: four times the distance, 17 times the interval, close to 4² = 16. To see where the square comes from, write down exactly what two rays determine.
Camera A sits at the origin looking along +z, and camera B is A slid right by b, as in lesson 1 (image width W = 120 px, focal length f in pixels). A point P = (X, Z) is described in A's frame. Pixel uA fixes a ray from A with slope kA = (uA − W/2)/f, so X = kAZ; pixel uB fixes a ray from B, X − b = kBZ. Two equations, two unknowns:
X − kAZ = 0, X − kBZ = b ⟹ Z = b/(kA − kB) = f·b/(uA − uB) = f·b/d
Lesson 1's formula is the solution of this system, and its denominator is the system's determinant, kA − kB = d/f. The system is singular exactly when d = 0: the rays are parallel and meet only at infinity. "The half pixel sits in the denominator" therefore has a name: a small determinant, an ill-conditioned system.
For two arbitrary cameras the system is [eA −eB]·(tA, tB)ᵀ = oB − oA, with unit ray directions e, camera centres o and distances t along the rays; the engine's intersection routine solves exactly this. Its determinant is sin θ, with θ the angle between the rays, and its singular values are √2·cos(θ/2) and √2·sin(θ/2), so its condition number is
κ = cot(θ/2) ≈ 2/θ
the factor by which a relative error in the ray directions can be amplified in the solution. For a point at depth Z on A's axis, θ = arctan(b/Z) ≈ b/Z and κ ≈ 2Z/b. With f = 44 px and b = 0.5 m:
| Depth Z | 2.5 m | 5 m | 10 m | 40 m |
|---|---|---|---|---|
| disparity d = f·b/Z | 8.8 px | 4.4 px | 2.2 px | 0.55 px |
| angle θ between the rays | 11.31° | 5.71° | 2.86° | 0.72° |
| condition number κ = cot(θ/2) | 10.1 | 20.0 | 40.0 | 160 |
Doubling the depth halves the angle and doubles the conditioning. Nothing here involves noise: κ belongs to the geometry alone.
2 · What a pixel error does to the pin
A matcher never reports a pixel position exactly. Say each pixel is wrong by an independent error of standard deviation σu. To first order the error of a result is its gradient times the errors of the inputs. With d = uA − uB and a point on A's axis (uA = W/2):
Z = f·b/d: ∂Z/∂uA = −Z²/(f·b), ∂Z/∂uB = +Z²/(f·b) | X = (uA − W/2)Z/f: ∂X/∂uA = Z/f, ∂X/∂uB = 0
Independent errors add in variance, so
σZ = √2·Z²σu/(f·b), σX = Z·σu/f, Cov(X, Z) = −Z³σu²/(f²b)
The √2 is the two pixels: the disparity has error σd = √2·σu, and σZ = Z²σd/(f·b) is the form that stereo papers print (Gallup et al., 2008) and depth-camera guides use, where σd is called the sub-pixel error. Depth error grows as Z² and falls as 1/(f·b); sideways error grows only as Z; the ratio of the two standard deviations is √2·Z/b. The covariance is negative: a wrong uA slides the point along B's ray, which leans by θ from the axis, so the errors in X and Z are tied together and the error region is a tilted needle.
Its shape is the conditioning of §1 made visible. Turning ray A by a small angle ε slides the intersection along ray B by rε/sin θ, with r the distance to P, and turning ray B slides it along ray A by the same. With unit directions eA, eB and equal distances the covariance is (rε/sin θ)²·(eAeAᵀ + eBeBᵀ) = (rε/sin θ)²·M·Mᵀ for M = [eA eB], whose eigenvalues are 1 ± cos θ. The error ellipse has axes in the ratio √((1 + cos θ)/(1 − cos θ)) = cot(θ/2) = κ ≈ 2Z/b: the condition number drawn in the plane. It is √2 times the ratio √2·Z/b of the two standard deviations, because the ellipse is tilted and its axes are not along X and Z. (Equivalently, two pixel-wide wedges cross in a lozenge whose length-to-width ratio tends to κ as the pixels shrink.)
This is first-order theory, so test it. Put P on A's axis, give both pixels independent errors of σu = 0.1 px, triangulate with the general intersection, repeat 20 000 times and measure the cloud (b = 0.5 m):
| Depth Z | σZ first order | σZ measured | σX measured | axis ratio measured |
|---|---|---|---|---|
| 2.5 m | 0.040 m | 0.040 m | 0.57 cm | 10.1 (κ = 10.1) |
| 5 m | 0.161 m | 0.160 m | 1.13 cm | 20.0 (κ = 20.0) |
| 10 m | 0.643 m | 0.649 m | 2.28 cm | 40.3 (κ = 40.0) |
Doubling the depth multiplies σZ by 4.0 and σX by 2.0: at the tower's depth the point is known to 0.65 m in depth and 2.3 cm sideways. First-order theory has a limit. Z = f·b/d is convex in d, so a symmetric error in d gives a skewed error in Z: the cloud sits a little beyond P and has a long far tail. The exact spread (an integral over the Gaussian, matched by simulation) exceeds the first-order one by 0.1% at 2.5 m, 0.4% at 5 m and 1.7% at 10 m, about 4(σd/d)². The formula is good while σd/d < 0.1, here for Z below f·b/(10σd) = 15.6 m, and underestimates beyond.
3 · Spend the baseline
The denominator of the law holds two knobs. A larger f buys precision by narrowing the field of view: for the same 120 pixels, doubling f from 44 to 88 px halves σZ and cuts the field of view from 107° to 69°. The other, σZ ∝ 1/b, costs nothing on paper. Why not take the widest baseline? Two prices, measured on the street, where 51 of A's 120 pixels see a surface. Overlap: B stands further right, so it loses what lies beyond its left edge and what a near object now hides; a surface pixel of A is seen if its point lands inside B's image and B's own depth buffer agrees there. Matching: the partner now lies anywhere within f·b/Zmin pixels, with Zmin = 2 m the nearest thing expected (11 px at b = 0.5 m, 44 px at 2 m), and the more candidates, the more look right (§4 describes the matcher). We count the seen pixels whose whole-pixel match lands within one pixel of the truth:
| baseline b | σZ at 10 m | surface pixels seen by B | of those, matched right | usable, of 51 |
|---|---|---|---|---|
| 0.25 m | 1.29 m | 100.0% | 100.0% | 51 |
| 0.5 m | 0.64 m | 98.0% | 92.0% | 46 |
| 1 m | 0.32 m | 80.4% | 80.5% | 33 |
| 2 m | 0.16 m | 52.9% | 55.6% | 15 |
Precision grows in proportion to b (here at σu = 0.1 px); the number of usable points falls gently at first and then fast: from 1 m to 2 m it more than halves.
4 · Find the partner
The pin needs both pixels, and so far the partner was given. The point lies somewhere on A's ray, and B sees that ray as a line. In a plane that line is the whole row of B, so the partner of pixel i is the pixel i − d on the same row, with d between 0 and the search range: a one-dimensional search. (In three dimensions the line is a slanted line of the image, the epipolar line.) Compare a window of colours, half-width h pixels, around i in A with the window around i − d in B, sum the squared differences over the window and the three colour channels, and keep the shift with the lowest cost:
Ci(d) = Σk=−hh ‖A(i + k) − B(i + k − d)‖², d̂ = argmind Ci(d)
Block matching has three behaviours, each visible in the widget below. A pixel on a textured surface has one deep valley at the true disparity. A pixel on sky has no texture: Ci(d) is zero for every shift and nothing is learned. Where stripes repeat, as on the ball, the cost has several valleys one stripe apart; the wider the search range, the more rivals, and one can win. The photographs are noise-free, so whatever goes wrong goes wrong on texture and geometry alone.
Where the match is right, how right? d̂ is a whole number and the truth is not. If true disparities are spread evenly the rounding error is uniform on ±½ px, with standard deviation 1/√12 = 0.29 px: lesson 1's half pixel. Measured on the crate's front face, a surface of constant depth, over 81 baselines from 0.30 to 0.70 m, the whole-pixel matcher is off by 0.30 px rms. Sub-pixel refinement fits a parabola through the costs at d̂ − 1, d̂, d̂ + 1 and takes its vertex: 0.14 px. Since σd = √2·σu, that is σu = 0.21 px for whole pixels and 0.10 px refined, which is why the widget opens at σu = 0.1 px. Refinement halves σZ at every depth but cannot change the law: σu multiplies Z²/(f·b) and leaves the square where it was.
What to try. The page opens on the pin: b = 0.5 m, P at Z = 5 m, σu = 0.1 px. The rays meet at θ = 5.71° with κ = 20.0; the 20 000 trials give σZ = 0.160 m against 0.161 m from the formula and an ellipse with axes in the ratio 20.0 : 1, a needle. Drag Z to 10 m and to 2.5 m for the other rows of the table: 0.649 m and 40.3, then 0.040 m and 10.1. At Z = 10 m double b to 1 m: σZ halves to 0.321 m and B sees 80.4% of A's surface pixels. Set σu = 0.2 px (whole-pixel matching) and b = 0.5 m: at 10 m σZ is 1.37 m against the formula's 1.29 m, and at 14 m 2.92 m against 2.52 m: the first-order law failing, the tail stretching away from the cameras. The dashed 2.5 cm line meets the stereo curve at Z* = 1.97 m for b = 0.5 m and 3.94 m for b = 2 m (§5). Now choose the partner. Pixel 64, on the ball, at b = 0.5 m has one deep valley: the true disparity is 4.31 px, the whole-pixel estimate 4, the refined one 4.28. Pixel 100 sees only sky, and every shift costs zero. At b = 2 m pixel 60, near the ball's left edge, has true disparity 16.7 px but its cheapest shift is 8, because the stripes repeat; pixel 14, on the crate, would have its partner 35.9 px to the left, outside B, and still gets 8, a confident error.
5 · Time the light instead
The Z² comes from the principle, not the engineering: triangulation measures an angle, and depth is the reciprocal of a small angle. A sensor that times light measures range directly. A pulse that returns after Δt has travelled out and back, so r = c·Δt/2 (Li and Ibanez-Guzman, 2020; air's refractive index taken as 1): a timing error of 1 ns is 15.0 cm of range and 100 ps is 1.5 cm, at any range. A continuous-wave time-of-flight camera modulates its light at frequency fm and measures the phase φ of the return, d = c·φ/(4π·fm) (Li, 2014). The phase wraps every 2π, so the range wraps every c/(2fm): 14.99 m at 10 MHz, 7.49 m at 20, 3.00 m at 50 and 1.50 m at 100 MHz. A higher modulation frequency measures phase more finely and wraps sooner: the baseline's bargain again, with ambiguity in place of overlap.
The error of such a sensor is set by timing and signal strength, not by geometry. Datasheets show a few centimetres, growing slowly with range as the signal-to-noise ratio falls: the Velodyne HDL-64E spec sheet quotes ±2 cm typical, and the Ouster OS2 datasheet (2020) a precision of 2.5 cm from 1 to 30 m, 4 cm from 30 to 60 m and 8 cm beyond (one standard deviation, 10% reflectivity target). Equate stereo's Z²σd/(f·b) with a range sensor's error σR and solve for the depth: Z* = √(σR·f·b/σd). For a rig with f = 1000 px (a camera 1280 pixels wide has an f of this order), b = 0.5 m and σd = 0.1 px against σR = 2.5 cm, Z* = 11.2 m. Stereo's error is 2.0, 8.0, 50 and 200 cm at 10, 20, 50 and 100 m (×4 per doubling); the range sensor's is 2.5, 2.5, 4 and 8 cm, from the datasheet brackets above.
Stereo wins close in, where the angle is wide; the range sensor wins far out, where it is not. Neither is better in general: one has a closed-form law in Z² that a longer baseline can fight, the other pays in photons. The widget's toy camera has f = 44 px and σd = √2σu, so at σu = 0.1 px its crossover with the same sensor is at 1.97 m for b = 0.5 m.
6 · A depth map is a cloud, in the sensor's own frame
Either source gives a depth Zi for each pixel i. The pixel's ray and its depth are a point, Pi = ((ui − W/2)·Zi/f, Zi), and A's depth map becomes a cloud of 51 points (the other pixels see only sky). This is the first 3-D representation of the track: an unordered set of points, with no surface (lesson 6) and no colour model (lesson 7). And its coordinates are the sensor's own: X to the right of whichever sensor made the cloud, Z along that sensor's axis, the origin at its centre.
Now move the sensor. Let a second range camera C stand where A would be after sliding by b and turning 15° to the left; both sensors have 2 cm of range noise, cm-level as in §5. Its depth map back-projects, by the same formula, into C's own frame. Switch the widget to two scans. On the left, both clouds on the same axes, as delivered, are two different pictures of the street; their mean nearest-point distance (from each point to the closest point of the other cloud, averaged both ways) is 0.70 m at b = 0.5 m, with 51 points from A and 48 from C. On the right C's cloud is rotated by 15° and moved to where C really stood, and the distance falls to 0.056 m, the floor set by sampling and noise. Drag the baseline: with the true pose the distance never exceeds 0.10 m, while as delivered it stays between 0.58 and 0.88 m; with no slide at all, a pure 15° turn, it is 0.94 m. The pin of §1 stood in the same place: the b on the right-hand side of its equations was told to us.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
We can now say how well a known pair of cameras pins a point: a needle about 2Z/b times longer than wide, with depth error √2·Z²σu/(f·b), which a longer baseline or a range sensor buys back at a price. Either source turns a depth map into a cloud, 51 points for A in the street. But those points live in their sensor's frame: two sensors in different places deliver clouds 0.70 m apart on average when C has slid 0.5 m and turned 15°, and 0.056 m apart once C's true pose is applied. Everything above assumed that pose, and nothing in a pixel or a range supplies it. To merge the clouds we must write down a rotation and a translation, and a rotation cannot be added, averaged or stepped along like a vector. How do we describe a rigid motion so that a computer can estimate it?
Interview prompts
- Derive the depth error of a stereo pair and say why it grows with the square of the distance. (§2 — differentiate Z = f·b/d: ∂Z/∂d = −Z²/(f·b), so σZ = Z²σd/(f·b); the disparity shrinks like 1/Z while its error stays put.)
- Why is the uncertainty of a triangulated point elongated along the rays, and by how much? (§1, §2 — a pixel error turns one ray, moving it sideways by Z·σu/f at P, and the intersection slides along the other ray by that over sin θ; the axes are in the ratio cot(θ/2) ≈ 2Z/b, the condition number.)
- Why does a wider baseline help, and what stops you using a very wide one? (§3 — σZ ∝ 1/b, but overlap shrinks, the search range grows and false matches multiply.)
- Why is the correspondence search one-dimensional, and how can it fail? (§4 — the partner lies on the epipolar line, here the same row; it fails on texture-less pixels, repeated stripes and partners outside the other image.)
- Where does a 2.5 cm range sensor overtake a stereo rig with f = 1000 px, b = 0.5 m, σd = 0.1 px? (§5 — equate Z²σd/(f·b) with σR: Z* = √(σRfb/σd) = 11.2 m.)
- Why can two depth maps not simply be overlaid into one cloud? (§6 — each is in its own sensor's frame; they coincide only after the relative rotation and translation are applied.)
Companion reads: Computer Vision · 05 Multi-view, depth and SLAM (the stereo equation and epipolar geometry in three dimensions).