From points to a surface
Lesson 5 left us with camera poses and a cloud of points. Asked what each ray of a camera nobody used hits, the cloud answers wrongly for 12 to 17% of the rays with four, eight or twenty-four exact depth maps, and for 31% once the depth is noisy. This lesson says what a model needs in order to answer (a function of position whose sign tells inside from outside and whose size is a distance), builds it by letting every noisy depth map vote on every cell of a grid, and combines the votes with maximum-likelihood weights, so that the error falls as 1/√n. The surface that results has no colour.
New idea: the surface is the zero set of a signed-distance field, and the field is the inverse-variance weighted mean of the truncated signed distances that every depth map votes at every cell. Its error falls as 1/√n, and the band in which a vote counts is chosen between the noise and the smallest feature.
Forces next: Fusing noisy depth into a signed distance field gives a surface that answers "where is the first hit along this ray?", with a normal and an inside and an outside. But the held-out photograph is made of colors, not positions, and so far nothing has a color. What does a pixel actually measure, and how do we write down, and score, the image a model predicts?
1 · What a cloud cannot answer
Take four of the ring's 24 depth maps (cameras 90° apart, noise switched off) and back-project every pixel that hit something, as in lesson 2: 153 points, every one exactly on the true surface, each with the colour of its pixel. Now the exam. Twenty-four held-out cameras, interleaved with the ring and never used, photograph the statue scene at 64 pixels each; 886 of their 1,536 pixel rays hit something. Render from the cloud the usual way: every point lands on a pixel and the nearest one wins. The cloud answers "what does this ray hit?" with the point that landed in the ray's pixel, and the answer is right if that point is within 5 cm, along the ray, of the true first hit.
17.2% of the 886 rays are answered wrongly. 75 see through: the pixel holds no point of the front surface, only one of the far side. 74 are answered by a point in front of the true surface: a point is a pixel-sized disc facing the camera, and where the surface is seen at an angle the disc sits off it. 3 get no answer. (A ray that hits nothing but still gets an answer, 34 of the other 650, is not counted as wrong here, though it costs PSNR.) Scored by PSNR (lesson 7 defines it; higher is better) and averaged over the 24 photographs, the cloud gets 18.3 dB. Would more points fix it?
| point cloud (depth noise off unless stated) | points | rays wrong | see through | answered in front |
|---|---|---|---|---|
| 2 views | 62 | 61.2% | 267 | 32 |
| 4 views | 153 | 17.2% | 75 | 74 |
| 8 views | 296 | 11.7% | 6 | 98 |
| 24 views | 884 | 15.9% | 0 | 141 |
| 24 views, noisy depth (3 cm at 6 m) | 884 | 31.3% | 0 | 277 |
Two views leave gaps: 62 points, 267 rays that see through. More points close them, and from eight views on almost no ray sees through, but the wrong answers do not go away; they all become "in front", 141 of the 886 rays, although every point is exactly on the surface. With the stereo sensor of lesson 2 (3 cm of depth error at 6 m) more views make the cloud worse: 23.7% wrong with four views, 31.3% with twenty-four, because the nearest of the many noisy points in a pixel is the one that happens to lie closest to the camera. Nothing in a cloud averages: each point is a separate claim, and nothing says that two of them describe the same place.
So the trouble is not the number of points. A point can say "I am here". A ray needs to be told "the first surface along you is there", and that is a statement about the space between points and about which side of a surface a place is on.
2 · What an answerable model provides
Three queries, because they are what a ray tracer, and so the exam, asks of a scene. Inside or outside? at any point. Where does this ray first hit? the smallest distance at which the ray is inside. Which way does the surface face? the normal at the hit, which shading needs in lesson 7. A function f(x, z) with f > 0 outside and f < 0 inside answers the first by its sign, the second by the first sign change along the ray, and the third by the direction of its gradient ∇f. Which f? Four candidates; the bit and the number are read from one and the same fused field (§5), so only the representation differs.
| what is stored | first hit | normal | field lookups per ray | verdict |
|---|---|---|---|---|
| points (§1) | wrong for 15.9% to 31.3% of the rays | none | - | no inside, no between |
| a bit per cell (the sign of the field) | to one cell: median 2.20 cm, 90% within 4.85 cm | median error 20.1° | 162 | decides inside, loses the distance |
| a triangle mesh | exact | exact | - | needs the points' connectivity, which we lack: an output of the field (§5), not an input |
| a number per cell, the signed distance | sub-cell: median 0.67 cm, 90% within 2.08 cm | median error 3.8° | 60 | answers all three |
A bit tells a ray only that the surface is somewhere in a cell; a number lets it interpolate, and step. If |∇f| ≤ 1 then f is a lower bound on the distance to the nearest surface, so a ray can advance by f without crossing anything: sphere tracing, which Computer Graphics · 03 uses on shapes given as formulas. Here no formula is given: f must be estimated from depth maps. A depth map gives such a distance one ray at a time: for every point on a pixel's ray, how much deeper the first surface lies. What remains is to turn many such noisy measurements, from different cameras, into one function of position.
3 · One depth map votes
Lay a grid over the region, a node every h = 5 cm. Camera v sees node x at axis depth zc(x) (depth as in lesson 1) and its depth map says that the first surface on that pixel is at depth zm. The view's vote on the node is
d = zm − zc(x)
positive for a node in front of the surface (free space), negative behind it, zero on it. Taken at face value the vote is wrong in two ways, and each is decided by a computation.
Far in front. d is measured along this view's ray, not to the nearest surface. A node within 20 cm of a surface can get a vote of +2.3 m from a camera whose ray runs past that surface and hits something far behind it, and averaged with the others that vote would drag the field away from the truth for no reason. So clip at a truncation distance μ: d ← min(d, μ), which says "at least μ in front". Without the clip the statue's surface is 1.8 times as far off (1.59 cm against 0.88 cm RMS with 24 views).
Behind. A depth map says nothing about what lies behind the first surface it sees, so a node on the far side of the statue is hidden from this camera. If it voted "inside" (d = −μ) it would outvote the views that can see it, the field would be negative behind every object and the statue's zero set would vanish (what remains lies up to 0.97 m from any real surface). So the view abstains when d < −μ. Nodes within μ behind the surface do vote, on the assumption that the object is at least μ thick there; §5 is where that assumption is paid for.
A vote is therefore a number in [−μ, +μ] or an abstention, and each view attaches a weight w ≥ 0 to its number. What should the weight be, and how should the numbers be combined?
4 · Averaging the votes
At one node, k views vote d1, …, dk. Write di = d* + εi, with d* the unknown true value and εi independent Gaussian noise of standard deviation σi, the depth noise of that pixel. Which one number D should the node store? The one of maximum likelihood L:
−ln L(D) = ½ Σi (di − D)² / σi² + const ⟹ D = Σ wi di / Σ wi, wi = 1/σi²
The second derivative of −ln L is W = Σ wi, so D has variance 1/W. That is the whole fusion: the field is the inverse-variance weighted mean of the votes, and W is how much the node knows. Four consequences.
- Equal noise. D is the plain mean and its standard deviation is σ/√k. Measured on a node of a flat wall with σ = 3 cm (4000 trials each): 2.98 cm for one view, 1.48 for 4, 0.76 for 16, 0.38 for 64, a log-log slope of −0.50.
- Unequal noise. The stereo sensor of lesson 2 has σZ ∝ Z², so w = (6 m / Z)4: with 3 cm at 6 m a camera at 4.85 m has σ = 1.96 cm and one at 7.15 m 4.26 cm. On the statue the plain mean costs 24% more error (twelve noise realisations; never less than 16%).
- Incremental. D ← (W D + w d)/(W + w), W ← W + w: two numbers per node, however many views, and a new view is added without keeping the old ones. The order does not matter (reversing the 24 views gives the same field to rounding error). This is the cumulative weighted signed distance of Curless and Levoy (1996) for range scans; KinectFusion (Newcombe et al., 2011) runs it on every frame of a depth camera, with w = 1 and a cap on W so that old evidence fades when the scene moves.
- Unobserved. W = 0 means no view has spoken. That node is neither outside nor inside; it is a hole, and a ray must not mistake it for free space.
Why a grid? Because a node is a place. Votes from different cameras about the same place meet there without anyone having to decide which pixel of one photograph matches which pixel of another, the correspondence problem that averaging points or depth maps directly would raise.
What to try. (1) As it opens, four views are fused: the zero set lies 2.09 cm RMS from the statue's outline, 12.0% of that outline was never observed (grey), and 21.4% of the exam's rays are answered wrongly, about the cloud's 23.7%. Drag the views up to 24: 0.88 cm and 1.8%, against the cloud's 31.3%. (2) The amber curve follows the blue σ/√n line down and then flattens towards the grey one, the error with a perfect sensor (0.65 cm at 24 views). That floor belongs to the depth maps' pixels, not to the voxels: a 2.5 cm grid leaves it at 0.70 cm, cameras with twice the pixels lower it to 0.41 cm. The two add in quadrature: √(0.65² + 0.61²) = 0.90 cm, close to the measured 0.88. (3) Slide μ to 4 cm (1.40 cm: noise is clipped) and to 30 cm (2.37 cm; 14.5 cm at the worst place, where neighbours merge). (4) Raise the noise to 8 cm at μ = 10 cm: 3.13 cm against σ/√n = 1.63. A wider band helps, to 1.92 cm at μ = 15, and then hurts again (2.51 cm at 25): this sensor is too noisy for this scene (§5). (5) Choose "the held-out exam" and step the amber camera round the ring: the bottom strip, the error of the copied colours, lights up at silhouettes.
5 · Reading the surface, and choosing the band
The zero set. The surface is where D changes sign. On a grid edge whose end values Da and Db have opposite signs, the crossing lies a fraction Da/(Da − Db) of the way along. Marching squares joins the crossings of each cell: its 16 sign patterns fall into 4 classes under rotation and swapping inside with outside, and only one class (two opposite corners) is ambiguous; the sign of the average of the four corners settles it. In three dimensions the same idea is marching cubes (Lorensen and Cline, 1987): the 256 sign patterns of eight corners form, by our own count, 15 classes under rotation and the inside-outside swap, 14 of them not empty. The paper's text speaks of 14 patterns and its figure numbers them 0 to 14, counting the empty cube; it does not resolve the ambiguous faces, which later work (Nielson and Hamann, 1991) does.
The normal is the unit gradient of D, by central differences one cell either side: median error 3.8° against the true normal. Its length is not 1: at the surface the median of |∇D| is 1.2, because a vote is measured along a camera ray, so it exceeds the perpendicular distance by about 1/cos of the angle between ray and normal. This is why the tracer steps by 0.8 of the field value, not the whole of it.
The ray. Sphere tracing: at each step advance by 0.8 D (at least half a cell, at most μ, since a truncated field cannot vouch for anything farther), stop at the first sign change and bisect it. A node that was never observed is crossed blindly, in steps of μ/2. On the 886 held-out rays that really hit something, the first hit is within 0.67 cm of the truth for the median ray and within 2.08 cm for 90% of them; 2 get no answer, and 1.8% (16 rays) are wrong by the 5 cm criterion of §1, where the cloud was wrong for 31.3%.
The band. A vote counts only within μ of the surface. Noise sets the floor: a vote on the surface is off by σ, and the clip and the abstention must not cut into it, so μ should be at least about 3σ (9 cm here). Geometry sets the ceiling: a vote reaches μ behind the surface, assuming the object is at least that thick there and no other surface is nearer. The narrowest gap in the scene, between the crate's corner and a lobe of the statue, is 22.7 cm, so bands wider than 11.4 cm overlap in it, and from μ = 10 cm the worst error anywhere (last row below) sits at that corner. A thin plate suffers worst: a 4 cm plate has a region D < 0 across it that is 3.2 cm wide at μ = 6 cm, 7.3 cm at 10, 16.6 cm at 20 and 30.8 cm at 30: things thinner than the band swell towards its width.
| band μ (cm), σ = 3 cm, 24 views | 4 | 6 | 10 | 15 | 30 |
|---|---|---|---|---|---|
| statue error, RMS (cm) | 1.40 | 1.05 | 0.88 | 1.31 | 2.37 |
| worst error anywhere (cm) | 6.1 | 4.1 | 11.1 | 14.4 | 14.5 |
6 · The exam, and what the surface still lacks
Run the exam that the cloud failed on the surface (24 views, 3 cm of noise at 6 m, μ = 10 cm). A colourless surface needs a colour for its pixels, and there are two honest choices: the one colour that fits best, the mean colour of the object pixels in the training photographs; or, for each hit point, the colour of the pixel it lands on in the training photograph whose camera is nearest to the held-out one. The two rows that say "true surface" replace the fused surface by the exact one, to find out how much of any shortfall belongs to the geometry.
| what answers the ray, and with what colour | rays answered wrongly | held-out PSNR |
|---|---|---|
| the cloud of 884 points, each with its own colour | 31.3% | 16.6 dB |
| the fused surface, one colour | 1.8% | 18.1 dB |
| the true surface, one colour | 0% | 19.2 dB |
| the fused surface, colours copied from the nearest photograph | 1.8% | 26.9 dB |
| the true surface, colours copied | 0% | 28.6 dB |
Geometry went from 31.3% of the rays wrong to 1.8%, and the score moved by 1.5 dB. The true surface painted with the best single colour reaches only 19.2 dB, so a model with no colour cannot pass this exam whatever its geometry. Copying colours helps (26.9 dB) and does not remove the real limit: the true surface with the same recipe scores 28.6 dB. Where does the rest go? 24 of the 1,536 pixels carry 94% of the squared error of the copied colours, and 23 of those 24 lie on silhouettes (the last on an edge of the crate, between a lit face and a shaded one). At a silhouette one pixel's ray lands on the statue and its neighbour's on the wall behind; in another photograph the same surface point falls in a pixel that is entirely statue or entirely background, depending on where that photograph's grid happens to lie. A pixel is not the colour of a point; it is whatever the ray through it meets first. The surface knows where; nothing in it knows what comes back.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
Noisy depth has become a surface that answers where each ray ends and which way it faces there: 1.8% of the held-out rays are wrong where the cloud missed 31.3%. It cannot say what comes back along the ray. Painted in the best single colour even the exact surface scores only 19.2 dB, and colours copied from photographs leave the error at silhouettes, where a pixel is entirely object or entirely background by the luck of its grid. What does a pixel actually measure, and how do we write down, and score, the image a model predicts?
Interview prompts
- Why can a point cloud not answer "what does this ray hit?", and why do more points not fix it? (§1 — gaps close but the wrong answers become "in front" (15.9% at 884 exact points), and with noisy depth more views are worse: 31.3%.)
- Which three queries must a model of geometry answer, and why store a signed distance rather than a bit or a mesh? (§2 — inside, first hit, normal; a bit gives the hit to a cell and the normal to about 20°, a mesh needs connectivity the points lack, a distance interpolates and lets the ray step.)
- Why is a depth map's vote truncated, and what does each side do? (§3 — clip in front, since d is measured along this ray (1.8 times the error without it); abstain behind, since hidden cells are unknown and voting "inside" erases the statue.)
- Derive the weighted-average fusion. (§4 — Gaussian votes of variance σi²: minimising ½ Σ (di − D)²/σi² gives D = Σ widi/Σ wi, wi = 1/σi², variance 1/W, updated incrementally in any order.)
- How does the fused error fall with the number of views, and why does it stop falling? (§4, widget — the noise part as σ/√n (slope −0.50); the depth maps' pixels leave a floor, 0.65 cm at 24 views, and the two add in quadrature.)
- How do you choose the truncation distance? (§5 — at least about 3σ, so clipping does not bias the vote; at most about half the narrowest gap (11.4 cm), or neighbouring surfaces merge; thin plates swell with μ.)
- The geometry is 98% right, yet the PSNR rises by only 1.5 dB. Why? (§6 — one colour caps even the true surface at 19.2 dB; copied colours reach 26.9 dB, and 24 pixels, nearly all at silhouettes, carry 94% of what is left: a pixel is not the colour of a point.)
Companion reads: Computer Graphics · 03 Geometry, meshes and curves (implicit surfaces and marching cubes as a graphics tool; here the field is estimated, not given) and Computer Vision · 05 Multi-view, depth and SLAM (the stereo depth fused here).