all_lessons/3D Vision/02 · A second ray pins the pointlesson 2 / 14

A second ray pins the point

Lesson 1 ended on Z = f·b/d and a warning: d is read off a pixel grid and sits in the denominator, so half a pixel moves the tower by metres and the crate by centimetres. This lesson makes the warning exact. Two pixel rays are a 2×2 linear system whose conditioning is the angle between them, about b/Z, so a pixel error costs a depth error that grows as Z² and a sideways error that grows only as Z: a triangulated point is a needle along the rays. Depth maps from either source back-project into clouds that do not fit together until we know where the sensors were.

The thesis, here
Two pixel rays meet in one point, but only because we say where both cameras are, and how sharply they pin it is set by the angle between them, about b/Z. A pixel error σu costs √2·Z²σu/(f·b) in depth and Z·σu/f sideways, and the error region is a needle about 2Z/b times longer than it is wide. A longer baseline, or a sensor that times light, shortens the needle. Neither tells us where the second sensor is.
Linear position
Forced by: A pixel pins down a direction and nothing else: every point along its ray lands on the same pixel, so depth is the one number per pixel that the image destroys. It is also exactly the number a new viewpoint needs, because moving the camera sideways by b shifts a pixel by f·b/Z. So the shift between two views measures depth. How precisely can two rays pin down a point, and what limits that precision?
New idea: two pixel rays are a 2×2 linear system whose conditioning is the triangulation angle θ ≈ b/Z, so pixel noise becomes an error needle: σZ = √2·Z²σu/(f·b) in depth, σX = Z·σu/f sideways. Precision has to be bought with baseline, with a smaller σu, or with a sensor that does not triangulate, and each purchase has a price.
Forces next: Two rays pin a point only because we were told where the second camera is. In practice that is the unknown: a robot, a phone or a scanner moves between measurements, and every depth map lives in its own sensor frame. To compare or merge frames we must write down a rotation and a translation, and a rotation cannot be added, averaged or stepped along like a vector. How do we describe a rigid motion so that a computer can estimate it?
The plan
Six moves. (1) Write the pin as a 2×2 system and read its conditioning off the angle between the rays. (2) Put noise on the pixels: the Z² law and the shape of the error. (3) Spend the baseline: precision against overlap and matching. (4) Find the partner pixel by a one-dimensional search. (5) Time the light instead, and find where it overtakes triangulation. (6) Back-project a depth map into points, and see what they lack.

1 · Two rays, two unknowns

Lesson 1 closed on two numbers. A half-pixel error in d leaves the crate anywhere in an interval 0.3 m long and the tower in one 4.8 m long: four times the distance, 17 times the interval, close to 4² = 16. To see where the square comes from, write down exactly what two rays determine.

Camera A sits at the origin looking along +z, and camera B is A slid right by b, as in lesson 1 (image width W = 120 px, focal length f in pixels). A point P = (X, Z) is described in A's frame. Pixel uA fixes a ray from A with slope kA = (uA − W/2)/f, so X = kAZ; pixel uB fixes a ray from B, X − b = kBZ. Two equations, two unknowns:

X − kAZ = 0,   X − kBZ = b  ⟹  Z = b/(kA − kB) = f·b/(uA − uB) = f·b/d

Lesson 1's formula is the solution of this system, and its denominator is the system's determinant, kA − kB = d/f. The system is singular exactly when d = 0: the rays are parallel and meet only at infinity. "The half pixel sits in the denominator" therefore has a name: a small determinant, an ill-conditioned system.

For two arbitrary cameras the system is [eA −eB]·(tA, tB)ᵀ = oB − oA, with unit ray directions e, camera centres o and distances t along the rays; the engine's intersection routine solves exactly this. Its determinant is sin θ, with θ the angle between the rays, and its singular values are √2·cos(θ/2) and √2·sin(θ/2), so its condition number is

κ = cot(θ/2) ≈ 2/θ

the factor by which a relative error in the ray directions can be amplified in the solution. For a point at depth Z on A's axis, θ = arctan(b/Z) ≈ b/Z and κ ≈ 2Z/b. With f = 44 px and b = 0.5 m:

Depth Z2.5 m5 m10 m40 m
disparity d = f·b/Z8.8 px4.4 px2.2 px0.55 px
angle θ between the rays11.31°5.71°2.86°0.72°
condition number κ = cot(θ/2)10.120.040.0160

Doubling the depth halves the angle and doubles the conditioning. Nothing here involves noise: κ belongs to the geometry alone.

2 · What a pixel error does to the pin

A matcher never reports a pixel position exactly. Say each pixel is wrong by an independent error of standard deviation σu. To first order the error of a result is its gradient times the errors of the inputs. With d = uA − uB and a point on A's axis (uA = W/2):

Z = f·b/d:   ∂Z/∂uA = −Z²/(f·b),   ∂Z/∂uB = +Z²/(f·b)   |   X = (uA − W/2)Z/f:   ∂X/∂uA = Z/f,   ∂X/∂uB = 0

Independent errors add in variance, so

σZ = √2·Z²σu/(f·b),   σX = Z·σu/f,   Cov(X, Z) = −Z³σu²/(f²b)

The √2 is the two pixels: the disparity has error σd = √2·σu, and σZ = Z²σd/(f·b) is the form that stereo papers print (Gallup et al., 2008) and depth-camera guides use, where σd is called the sub-pixel error. Depth error grows as Z² and falls as 1/(f·b); sideways error grows only as Z; the ratio of the two standard deviations is √2·Z/b. The covariance is negative: a wrong uA slides the point along B's ray, which leans by θ from the axis, so the errors in X and Z are tied together and the error region is a tilted needle.

Its shape is the conditioning of §1 made visible. Turning ray A by a small angle ε slides the intersection along ray B by rε/sin θ, with r the distance to P, and turning ray B slides it along ray A by the same. With unit directions eA, eB and equal distances the covariance is (rε/sin θ)²·(eAeAᵀ + eBeBᵀ) = (rε/sin θ)²·M·Mᵀ for M = [eA eB], whose eigenvalues are 1 ± cos θ. The error ellipse has axes in the ratio √((1 + cos θ)/(1 − cos θ)) = cot(θ/2) = κ ≈ 2Z/b: the condition number drawn in the plane. It is √2 times the ratio √2·Z/b of the two standard deviations, because the ellipse is tilted and its axes are not along X and Z. (Equivalently, two pixel-wide wedges cross in a lozenge whose length-to-width ratio tends to κ as the pixels shrink.)

This is first-order theory, so test it. Put P on A's axis, give both pixels independent errors of σu = 0.1 px, triangulate with the general intersection, repeat 20 000 times and measure the cloud (b = 0.5 m):

Depth ZσZ first orderσZ measuredσX measuredaxis ratio measured
2.5 m0.040 m0.040 m0.57 cm10.1 (κ = 10.1)
5 m0.161 m0.160 m1.13 cm20.0 (κ = 20.0)
10 m0.643 m0.649 m2.28 cm40.3 (κ = 40.0)

Doubling the depth multiplies σZ by 4.0 and σX by 2.0: at the tower's depth the point is known to 0.65 m in depth and 2.3 cm sideways. First-order theory has a limit. Z = f·b/d is convex in d, so a symmetric error in d gives a skewed error in Z: the cloud sits a little beyond P and has a long far tail. The exact spread (an integral over the Gaussian, matched by simulation) exceeds the first-order one by 0.1% at 2.5 m, 0.4% at 5 m and 1.7% at 10 m, about 4(σd/d)². The formula is good while σd/d < 0.1, here for Z below f·b/(10σd) = 15.6 m, and underestimates beyond.

3 · Spend the baseline

The denominator of the law holds two knobs. A larger f buys precision by narrowing the field of view: for the same 120 pixels, doubling f from 44 to 88 px halves σZ and cuts the field of view from 107° to 69°. The other, σZ ∝ 1/b, costs nothing on paper. Why not take the widest baseline? Two prices, measured on the street, where 51 of A's 120 pixels see a surface. Overlap: B stands further right, so it loses what lies beyond its left edge and what a near object now hides; a surface pixel of A is seen if its point lands inside B's image and B's own depth buffer agrees there. Matching: the partner now lies anywhere within f·b/Zmin pixels, with Zmin = 2 m the nearest thing expected (11 px at b = 0.5 m, 44 px at 2 m), and the more candidates, the more look right (§4 describes the matcher). We count the seen pixels whose whole-pixel match lands within one pixel of the truth:

baseline bσZ at 10 msurface pixels seen by Bof those, matched rightusable, of 51
0.25 m1.29 m100.0%100.0%51
0.5 m0.64 m98.0%92.0%46
1 m0.32 m80.4%80.5%33
2 m0.16 m52.9%55.6%15

Precision grows in proportion to b (here at σu = 0.1 px); the number of usable points falls gently at first and then fast: from 1 m to 2 m it more than halves.

Road not taken · the widest baseline the vehicle allows
Attractive because σZ ∝ 1/b is free on paper. What breaks is everything around it: B no longer sees what A sees, the search range and the false candidates grow with b, and the same surface looks different from the two ends of a wide baseline, which the window of §4 assumes it does not. The idea returns in a better form when many views replace one wide pair (lessons 5 and 6).

4 · Find the partner

The pin needs both pixels, and so far the partner was given. The point lies somewhere on A's ray, and B sees that ray as a line. In a plane that line is the whole row of B, so the partner of pixel i is the pixel i − d on the same row, with d between 0 and the search range: a one-dimensional search. (In three dimensions the line is a slanted line of the image, the epipolar line.) Compare a window of colours, half-width h pixels, around i in A with the window around i − d in B, sum the squared differences over the window and the three colour channels, and keep the shift with the lowest cost:

Ci(d) = Σk=−hh ‖A(i + k) − B(i + k − d)‖²,   d̂ = argmind Ci(d)

Block matching has three behaviours, each visible in the widget below. A pixel on a textured surface has one deep valley at the true disparity. A pixel on sky has no texture: Ci(d) is zero for every shift and nothing is learned. Where stripes repeat, as on the ball, the cost has several valleys one stripe apart; the wider the search range, the more rivals, and one can win. The photographs are noise-free, so whatever goes wrong goes wrong on texture and geometry alone.

Where the match is right, how right? d̂ is a whole number and the truth is not. If true disparities are spread evenly the rounding error is uniform on ±½ px, with standard deviation 1/√12 = 0.29 px: lesson 1's half pixel. Measured on the crate's front face, a surface of constant depth, over 81 baselines from 0.30 to 0.70 m, the whole-pixel matcher is off by 0.30 px rms. Sub-pixel refinement fits a parabola through the costs at d̂ − 1, d̂, d̂ + 1 and takes its vertex: 0.14 px. Since σd = √2·σu, that is σu = 0.21 px for whole pixels and 0.10 px refined, which is why the widget opens at σu = 0.1 px. Refinement halves σZ at every depth but cannot change the law: σu multiplies Z²/(f·b) and leaves the square where it was.

Pin it, match it, merge it
Camera A (blue) and camera B (amber, slid right by the baseline) look at the street. The pin: a point P on A's axis is triangulated 20 000 times from pixels with independent errors σu; 300 trials are drawn in the world and around P at equal scale, with the measured 1σ ellipse (solid) and the first-order one (dashed); below, σZ against Z. The partner: the matching cost of every shift for one pixel, and how every surface pixel was matched. Two scans: §6.
angle θ, condition number κ
—
σZ measured (first order)
—
σX measured
—
ellipse axis ratio
—
stereo beats a 2.5 cm sensor up to
—
surface pixels B sees
—
of those, matched right
—
true disparity, pixel i
—
estimate: whole · refined
—
nearest-point distance: as delivered · true pose
—
Show the core JS
for (var k = 0; k < n; k++) {
  var r = FL.triangulate(camA, uA0 + su * g[2 * k], camB, uB0 + su * g[2 * k + 1]);
  if (!r.ok || r.zc <= 0 || r.zc > ZMAX) { out.lost++; continue; }
  var c = FL.toCam(camA, r.x, r.z);
  out.x[out.n] = c.xc; out.z[out.n] = c.zc; out.n++;
}
...
var sx = Z * su / f, sz = Math.SQRT2 * Z * Z * su / (f * b), cxz = -Z * Z * Z * su * su / (f * f * b), e = TL.eig2(sx * sx, cxz, sz * sz);
...
for (k = -h; k <= h; k++) {
  var ia = i + k, ib = ia - d;
  if (ia < 0 || ia >= imgA.W || ib < 0 || ib >= imgB.W) { c = NaN; break; }
  c += sq(imgA.r[ia] - imgB.r[ib]) + sq(imgA.g[ia] - imgB.g[ib]) + sq(imgA.b[ia] - imgB.b[ib]);
}
...
var den = cst[d - 1] - 2 * cst[d] + cst[d + 1];
return den > 1e-12 ? d + 0.5 * (cst[d - 1] - cst[d + 1]) / den : d;

What to try. The page opens on the pin: b = 0.5 m, P at Z = 5 m, σu = 0.1 px. The rays meet at θ = 5.71° with κ = 20.0; the 20 000 trials give σZ = 0.160 m against 0.161 m from the formula and an ellipse with axes in the ratio 20.0 : 1, a needle. Drag Z to 10 m and to 2.5 m for the other rows of the table: 0.649 m and 40.3, then 0.040 m and 10.1. At Z = 10 m double b to 1 m: σZ halves to 0.321 m and B sees 80.4% of A's surface pixels. Set σu = 0.2 px (whole-pixel matching) and b = 0.5 m: at 10 m σZ is 1.37 m against the formula's 1.29 m, and at 14 m 2.92 m against 2.52 m: the first-order law failing, the tail stretching away from the cameras. The dashed 2.5 cm line meets the stereo curve at Z* = 1.97 m for b = 0.5 m and 3.94 m for b = 2 m (§5). Now choose the partner. Pixel 64, on the ball, at b = 0.5 m has one deep valley: the true disparity is 4.31 px, the whole-pixel estimate 4, the refined one 4.28. Pixel 100 sees only sky, and every shift costs zero. At b = 2 m pixel 60, near the ball's left edge, has true disparity 16.7 px but its cheapest shift is 8, because the stripes repeat; pixel 14, on the crate, would have its partner 35.9 px to the left, outside B, and still gets 8, a confident error.

5 · Time the light instead

The Z² comes from the principle, not the engineering: triangulation measures an angle, and depth is the reciprocal of a small angle. A sensor that times light measures range directly. A pulse that returns after Δt has travelled out and back, so r = c·Δt/2 (Li and Ibanez-Guzman, 2020; air's refractive index taken as 1): a timing error of 1 ns is 15.0 cm of range and 100 ps is 1.5 cm, at any range. A continuous-wave time-of-flight camera modulates its light at frequency fm and measures the phase φ of the return, d = c·φ/(4π·fm) (Li, 2014). The phase wraps every 2π, so the range wraps every c/(2fm): 14.99 m at 10 MHz, 7.49 m at 20, 3.00 m at 50 and 1.50 m at 100 MHz. A higher modulation frequency measures phase more finely and wraps sooner: the baseline's bargain again, with ambiguity in place of overlap.

The error of such a sensor is set by timing and signal strength, not by geometry. Datasheets show a few centimetres, growing slowly with range as the signal-to-noise ratio falls: the Velodyne HDL-64E spec sheet quotes ±2 cm typical, and the Ouster OS2 datasheet (2020) a precision of 2.5 cm from 1 to 30 m, 4 cm from 30 to 60 m and 8 cm beyond (one standard deviation, 10% reflectivity target). Equate stereo's Z²σd/(f·b) with a range sensor's error σR and solve for the depth: Z* = √(σR·f·b/σd). For a rig with f = 1000 px (a camera 1280 pixels wide has an f of this order), b = 0.5 m and σd = 0.1 px against σR = 2.5 cm, Z* = 11.2 m. Stereo's error is 2.0, 8.0, 50 and 200 cm at 10, 20, 50 and 100 m (×4 per doubling); the range sensor's is 2.5, 2.5, 4 and 8 cm, from the datasheet brackets above.

Stereo wins close in, where the angle is wide; the range sensor wins far out, where it is not. Neither is better in general: one has a closed-form law in Z² that a longer baseline can fight, the other pays in photons. The widget's toy camera has f = 44 px and σd = √2σu, so at σu = 0.1 px its crossover with the same sensor is at 1.97 m for b = 0.5 m.

Road not taken · estimate the depth from one picture
A network can output a plausible Z per pixel with no second camera, baseline or timing hardware. It is a guess from experience, not a measurement with a known error law: it is wrong exactly when the scene is unlike its training data. Lesson 11 takes this road on purpose.

6 · A depth map is a cloud, in the sensor's own frame

Either source gives a depth Zi for each pixel i. The pixel's ray and its depth are a point, Pi = ((ui − W/2)·Zi/f, Zi), and A's depth map becomes a cloud of 51 points (the other pixels see only sky). This is the first 3-D representation of the track: an unordered set of points, with no surface (lesson 6) and no colour model (lesson 7). And its coordinates are the sensor's own: X to the right of whichever sensor made the cloud, Z along that sensor's axis, the origin at its centre.

Now move the sensor. Let a second range camera C stand where A would be after sliding by b and turning 15° to the left; both sensors have 2 cm of range noise, cm-level as in §5. Its depth map back-projects, by the same formula, into C's own frame. Switch the widget to two scans. On the left, both clouds on the same axes, as delivered, are two different pictures of the street; their mean nearest-point distance (from each point to the closest point of the other cloud, averaged both ways) is 0.70 m at b = 0.5 m, with 51 points from A and 48 from C. On the right C's cloud is rotated by 15° and moved to where C really stood, and the distance falls to 0.056 m, the floor set by sampling and noise. Drag the baseline: with the true pose the distance never exceeds 0.10 m, while as delivered it stays between 0.58 and 0.88 m; with no slide at all, a pure 15° turn, it is 0.94 m. The pin of §1 stood in the same place: the b on the right-hand side of its equations was told to us.

What this lesson did not do
It did not say where the second camera is: b and the heading were inputs everywhere, and in practice they are the unknown (lesson 3 describes a pose; lessons 4 and 5 estimate it). It worked in a plane, where two non-parallel rays always meet; in three dimensions they generally pass each other, triangulation becomes a least-squares closest point and the search runs along a slanted line (lesson 5). It assumed small independent Gaussian pixel errors, which false matches (§4) and the error sources of a range sensor are not. And a cloud is only points: no surface, no colour.

Common mistakes / failure modes

"the depth error is about the same at every depth"
It grows as Z²: 0.040 m at 2.5 m, 0.649 m at 10 m, same pixels (§2).
"a triangulated point has an error bar like a ball"
It is a needle along the rays, about 2Z/b times longer than wide (§2).
"the widest baseline is the best"
Precision grows as b, but at b = 2 m only 15 of 51 pixels are matched right (§3).
"sub-pixel matching fixes the far field"
It scales σu, by about 2 here; Z²/(f·b) is untouched (§4).
"a depth map is a point cloud in the world"
It is in the frame of the sensor that made it: two sensors give clouds that disagree by 0.70 m until the pose is applied (§6).

Checkpoint exercise

Try it
A rig has f = 800 px and b = 0.12 m, and its matcher is good to σd = 0.1 px. (a) How accurate is its depth at 5 m? (b) At what depth does σZ reach 10 cm? (c) What baseline would give 5 cm at 20 m with the same f and σd? (d) What is σX at 20 m with that baseline? Answer: (a) σZ = Z²σd/(f·b) = 25·0.1/96 = 2.6 cm. (b) Z = √(σZ·f·b/σd) = √96 = 9.8 m. (c) b = Z²σd/(f·σZ) = 400·0.1/(800·0.05) = 1.0 m. (d) With σu = σd/√2 = 0.071 px, σX = Z·σu/f = 1.8 mm: tiny next to the depth error, the needle again.

Where this points next

We can now say how well a known pair of cameras pins a point: a needle about 2Z/b times longer than wide, with depth error √2·Z²σu/(f·b), which a longer baseline or a range sensor buys back at a price. Either source turns a depth map into a cloud, 51 points for A in the street. But those points live in their sensor's frame: two sensors in different places deliver clouds 0.70 m apart on average when C has slid 0.5 m and turned 15°, and 0.056 m apart once C's true pose is applied. Everything above assumed that pose, and nothing in a pixel or a range supplies it. To merge the clouds we must write down a rotation and a translation, and a rotation cannot be added, averaged or stepped along like a vector. How do we describe a rigid motion so that a computer can estimate it?

Takeaway
Two pixel rays are a 2×2 linear system, X − kAZ = 0, X − kBZ = b, whose determinant is d/f and whose condition number is cot(θ/2) ≈ 2Z/b. A pixel error σu becomes an error needle, σZ = √2·Z²σu/(f·b) in depth and σX = Z·σu/f sideways, whose axis ratio is the condition number; the law holds while σd stays well below d. The baseline buys precision in proportion to b and pays in overlap, search range and false matches; sub-pixel refinement rescales σu and leaves the square. A sensor that times light has a different law, a few centimetres growing slowly, and overtakes stereo beyond Z* = √(σRfb/σd), 11.2 m for a typical rig. A depth map back-projects to a cloud in its sensor's own frame, and two such clouds disagree by 0.70 m until a pose is applied.

Interview prompts

Companion reads: Computer Vision · 05 Multi-view, depth and SLAM (the stereo equation and epipolar geometry in three dimensions).