all_lessons/3D Vision/06 · From points to a surfacelesson 6 / 14

From points to a surface

Lesson 5 left us with camera poses and a cloud of points. Asked what each ray of a camera nobody used hits, the cloud answers wrongly for 12 to 17% of the rays with four, eight or twenty-four exact depth maps, and for 31% once the depth is noisy. This lesson says what a model needs in order to answer (a function of position whose sign tells inside from outside and whose size is a distance), builds it by letting every noisy depth map vote on every cell of a grid, and combines the votes with maximum-likelihood weights, so that the error falls as 1/√n. The surface that results has no colour.

The thesis, here
A model of geometry has to answer three questions anywhere in space: is this point inside, where does this ray first hit, and which way does the surface face there. A cloud of points answers none of them, and it cannot average its own noise, because nothing says that points from two views describe the same place. A function whose sign tells inside from outside and whose value is a distance answers all three, and a grid is the common place where the votes of different cameras meet.
Linear position
Forced by: Images alone give camera poses and a sparse set of 3D points, up to one global scale, by adjusting both until the points reproject onto the pixels. That is the first model we fitted by comparing its prediction with the data. Now test it on the job it is for: predict a photograph from a camera we did not use. Points show through one another and nothing exists between them. What must a model of geometry provide to answer "what does this ray hit?", and how is it built from noisy depth?
New idea: the surface is the zero set of a signed-distance field, and the field is the inverse-variance weighted mean of the truncated signed distances that every depth map votes at every cell. Its error falls as 1/√n, and the band in which a vote counts is chosen between the noise and the smallest feature.
Forces next: Fusing noisy depth into a signed distance field gives a surface that answers "where is the first hit along this ray?", with a normal and an inside and an outside. But the held-out photograph is made of colors, not positions, and so far nothing has a color. What does a pixel actually measure, and how do we write down, and score, the image a model predicts?
The plan
Six moves. (1) Run the exam on a point cloud and see what it cannot answer. (2) State what an answerable model provides; choose the signed distance by elimination. (3) Let one depth map vote, and truncate the vote. (4) Average the votes: the maximum-likelihood weights and the 1/√n law. (5) Read the surface out of the field and choose the band. (6) Run the exam on the surface and find what it still lacks.

1 · What a cloud cannot answer

Take four of the ring's 24 depth maps (cameras 90° apart, noise switched off) and back-project every pixel that hit something, as in lesson 2: 153 points, every one exactly on the true surface, each with the colour of its pixel. Now the exam. Twenty-four held-out cameras, interleaved with the ring and never used, photograph the statue scene at 64 pixels each; 886 of their 1,536 pixel rays hit something. Render from the cloud the usual way: every point lands on a pixel and the nearest one wins. The cloud answers "what does this ray hit?" with the point that landed in the ray's pixel, and the answer is right if that point is within 5 cm, along the ray, of the true first hit.

17.2% of the 886 rays are answered wrongly. 75 see through: the pixel holds no point of the front surface, only one of the far side. 74 are answered by a point in front of the true surface: a point is a pixel-sized disc facing the camera, and where the surface is seen at an angle the disc sits off it. 3 get no answer. (A ray that hits nothing but still gets an answer, 34 of the other 650, is not counted as wrong here, though it costs PSNR.) Scored by PSNR (lesson 7 defines it; higher is better) and averaged over the 24 photographs, the cloud gets 18.3 dB. Would more points fix it?

point cloud (depth noise off unless stated)pointsrays wrongsee throughanswered in front
2 views6261.2%26732
4 views15317.2%7574
8 views29611.7%698
24 views88415.9%0141
24 views, noisy depth (3 cm at 6 m)88431.3%0277

Two views leave gaps: 62 points, 267 rays that see through. More points close them, and from eight views on almost no ray sees through, but the wrong answers do not go away; they all become "in front", 141 of the 886 rays, although every point is exactly on the surface. With the stereo sensor of lesson 2 (3 cm of depth error at 6 m) more views make the cloud worse: 23.7% wrong with four views, 31.3% with twenty-four, because the nearest of the many noisy points in a pixel is the one that happens to lie closest to the camera. Nothing in a cloud averages: each point is a separate claim, and nothing says that two of them describe the same place.

So the trouble is not the number of points. A point can say "I am here". A ray needs to be told "the first surface along you is there", and that is a statement about the space between points and about which side of a surface a place is on.

2 · What an answerable model provides

Three queries, because they are what a ray tracer, and so the exam, asks of a scene. Inside or outside? at any point. Where does this ray first hit? the smallest distance at which the ray is inside. Which way does the surface face? the normal at the hit, which shading needs in lesson 7. A function f(x, z) with f > 0 outside and f < 0 inside answers the first by its sign, the second by the first sign change along the ray, and the third by the direction of its gradient ∇f. Which f? Four candidates; the bit and the number are read from one and the same fused field (§5), so only the representation differs.

what is storedfirst hitnormalfield lookups per rayverdict
points (§1)wrong for 15.9% to 31.3% of the raysnone-no inside, no between
a bit per cell (the sign of the field)to one cell: median 2.20 cm, 90% within 4.85 cmmedian error 20.1°162decides inside, loses the distance
a triangle meshexactexact-needs the points' connectivity, which we lack: an output of the field (§5), not an input
a number per cell, the signed distancesub-cell: median 0.67 cm, 90% within 2.08 cmmedian error 3.8°60answers all three

A bit tells a ray only that the surface is somewhere in a cell; a number lets it interpolate, and step. If |∇f| ≤ 1 then f is a lower bound on the distance to the nearest surface, so a ray can advance by f without crossing anything: sphere tracing, which Computer Graphics · 03 uses on shapes given as formulas. Here no formula is given: f must be estimated from depth maps. A depth map gives such a distance one ray at a time: for every point on a pixel's ray, how much deeper the first surface lies. What remains is to turn many such noisy measurements, from different cameras, into one function of position.

3 · One depth map votes

Lay a grid over the region, a node every h = 5 cm. Camera v sees node x at axis depth zc(x) (depth as in lesson 1) and its depth map says that the first surface on that pixel is at depth zm. The view's vote on the node is

d = zm − zc(x)

positive for a node in front of the surface (free space), negative behind it, zero on it. Taken at face value the vote is wrong in two ways, and each is decided by a computation.

Far in front. d is measured along this view's ray, not to the nearest surface. A node within 20 cm of a surface can get a vote of +2.3 m from a camera whose ray runs past that surface and hits something far behind it, and averaged with the others that vote would drag the field away from the truth for no reason. So clip at a truncation distance μ: d ← min(d, μ), which says "at least μ in front". Without the clip the statue's surface is 1.8 times as far off (1.59 cm against 0.88 cm RMS with 24 views).

Behind. A depth map says nothing about what lies behind the first surface it sees, so a node on the far side of the statue is hidden from this camera. If it voted "inside" (d = −μ) it would outvote the views that can see it, the field would be negative behind every object and the statue's zero set would vanish (what remains lies up to 0.97 m from any real surface). So the view abstains when d < −μ. Nodes within μ behind the surface do vote, on the assumption that the object is at least μ thick there; §5 is where that assumption is paid for.

A vote is therefore a number in [−μ, +μ] or an abstention, and each view attaches a weight w ≥ 0 to its number. What should the weight be, and how should the numbers be combined?

4 · Averaging the votes

At one node, k views vote d1, …, dk. Write di = d* + εi, with d* the unknown true value and εi independent Gaussian noise of standard deviation σi, the depth noise of that pixel. Which one number D should the node store? The one of maximum likelihood L:

−ln L(D) = ½ Σi (di − D)² / σi² + const  ⟹  D = Σ wi di / Σ wi,   wi = 1/σi²

The second derivative of −ln L is W = Σ wi, so D has variance 1/W. That is the whole fusion: the field is the inverse-variance weighted mean of the votes, and W is how much the node knows. Four consequences.

Why a grid? Because a node is a place. Votes from different cameras about the same place meet there without anyone having to decide which pixel of one photograph matches which pixel of another, the correspondence problem that averaging points or depth maps directly would raise.

Fuse the depth maps
Left: the world from above (x right, z up), the grid coloured by the fused value D (blue in front of the surface, orange behind, grey never observed), the zero set in black over the true outline, the points of the n depth maps as dots, the n cameras at the rim and the held-out camera in amber. Top right: the statue's error against the number of views, for this sensor, a perfect one (grey) and σ/√n (blue). Bottom right: the votes at the point on the true surface that the amber camera looks at (dot size = weight) and their weighted mean. "The held-out exam" shows that camera's photograph and what each model makes of it; the dB figures are exam scores over all 24 cameras.
statue error (RMS)
-
σ/√n at 6 m
-
error if σ = 0
-
worst error anywhere
-
holes (boundary unobserved)
-
cloud: rays wrong
-
surface: rays wrong
-
PSNR: cloud
-
PSNR: surface, one colour
-
PSNR: surface, colours copied
-
PSNR: true surface, one colour
-
Show the core JS
var dx = X0 + i * H - cam.x, zc = dx * A.fx + dz * A.fz;
...
var d = dep[px] - zc;
if (d < -mu) continue;
if (d > mu) d = mu;
var r = ZR / dep[px], wt = o.uniform ? 1 : r * r * r * r;
a1[j * NG + i] += wt * d; a0[j * NG + i] += wt;
...
D[k] = Wt[k] > 0 ? S[k] / Wt[k] : 0;
...
var D = TS.sample(fld, ox + dx * t, oz + dz * t);
if (D !== D) { prevD = NaN; t += 0.5 * mu; continue; }
prevT = t; prevD = D;
t += Math.min(mu, Math.max(0.8 * D, 0.5 * TS.H));

What to try. (1) As it opens, four views are fused: the zero set lies 2.09 cm RMS from the statue's outline, 12.0% of that outline was never observed (grey), and 21.4% of the exam's rays are answered wrongly, about the cloud's 23.7%. Drag the views up to 24: 0.88 cm and 1.8%, against the cloud's 31.3%. (2) The amber curve follows the blue σ/√n line down and then flattens towards the grey one, the error with a perfect sensor (0.65 cm at 24 views). That floor belongs to the depth maps' pixels, not to the voxels: a 2.5 cm grid leaves it at 0.70 cm, cameras with twice the pixels lower it to 0.41 cm. The two add in quadrature: √(0.65² + 0.61²) = 0.90 cm, close to the measured 0.88. (3) Slide μ to 4 cm (1.40 cm: noise is clipped) and to 30 cm (2.37 cm; 14.5 cm at the worst place, where neighbours merge). (4) Raise the noise to 8 cm at μ = 10 cm: 3.13 cm against σ/√n = 1.63. A wider band helps, to 1.92 cm at μ = 15, and then hurts again (2.51 cm at 25): this sensor is too noisy for this scene (§5). (5) Choose "the held-out exam" and step the amber camera round the ring: the bottom strip, the error of the copied colours, lights up at silhouettes.

Road not taken · connect the dots
Skip the field and build a mesh straight from the cloud, by Delaunay triangulation or by fitting a smooth surface to the points and their normals, as Poisson surface reconstruction (Kazhdan et al., 2006) does. These start from a cloud that has already thrown away what each point knew about the space along its line of sight (free in front, unknown behind) and has no common frame in which repeated measurements of one place meet. The statue's noisy points of §1 lie 1.7 cm RMS from its surface, where the zero set of the fused field lies 0.88 cm from it: a mesh through the points inherits the roughness of single measurements unless something averages them first, and averaging is what the field does. Poisson's method does average, but it needs oriented points, and orienting a normal means knowing which side is outside, which the line of sight of a depth pixel tells us and a bare cloud does not. A third road, learning the field with a network (DeepSDF, Park et al., 2019; Occupancy Networks, Mescheder et al., 2019), keeps the field and replaces the average over views by a prior over shapes (lesson 11).

5 · Reading the surface, and choosing the band

The zero set. The surface is where D changes sign. On a grid edge whose end values Da and Db have opposite signs, the crossing lies a fraction Da/(Da − Db) of the way along. Marching squares joins the crossings of each cell: its 16 sign patterns fall into 4 classes under rotation and swapping inside with outside, and only one class (two opposite corners) is ambiguous; the sign of the average of the four corners settles it. In three dimensions the same idea is marching cubes (Lorensen and Cline, 1987): the 256 sign patterns of eight corners form, by our own count, 15 classes under rotation and the inside-outside swap, 14 of them not empty. The paper's text speaks of 14 patterns and its figure numbers them 0 to 14, counting the empty cube; it does not resolve the ambiguous faces, which later work (Nielson and Hamann, 1991) does.

The normal is the unit gradient of D, by central differences one cell either side: median error 3.8° against the true normal. Its length is not 1: at the surface the median of |∇D| is 1.2, because a vote is measured along a camera ray, so it exceeds the perpendicular distance by about 1/cos of the angle between ray and normal. This is why the tracer steps by 0.8 of the field value, not the whole of it.

The ray. Sphere tracing: at each step advance by 0.8 D (at least half a cell, at most μ, since a truncated field cannot vouch for anything farther), stop at the first sign change and bisect it. A node that was never observed is crossed blindly, in steps of μ/2. On the 886 held-out rays that really hit something, the first hit is within 0.67 cm of the truth for the median ray and within 2.08 cm for 90% of them; 2 get no answer, and 1.8% (16 rays) are wrong by the 5 cm criterion of §1, where the cloud was wrong for 31.3%.

The band. A vote counts only within μ of the surface. Noise sets the floor: a vote on the surface is off by σ, and the clip and the abstention must not cut into it, so μ should be at least about 3σ (9 cm here). Geometry sets the ceiling: a vote reaches μ behind the surface, assuming the object is at least that thick there and no other surface is nearer. The narrowest gap in the scene, between the crate's corner and a lobe of the statue, is 22.7 cm, so bands wider than 11.4 cm overlap in it, and from μ = 10 cm the worst error anywhere (last row below) sits at that corner. A thin plate suffers worst: a 4 cm plate has a region D < 0 across it that is 3.2 cm wide at μ = 6 cm, 7.3 cm at 10, 16.6 cm at 20 and 30.8 cm at 30: things thinner than the band swell towards its width.

band μ (cm), σ = 3 cm, 24 views46101530
statue error, RMS (cm)1.401.050.881.312.37
worst error anywhere (cm)6.14.111.114.414.5

6 · The exam, and what the surface still lacks

Run the exam that the cloud failed on the surface (24 views, 3 cm of noise at 6 m, μ = 10 cm). A colourless surface needs a colour for its pixels, and there are two honest choices: the one colour that fits best, the mean colour of the object pixels in the training photographs; or, for each hit point, the colour of the pixel it lands on in the training photograph whose camera is nearest to the held-out one. The two rows that say "true surface" replace the fused surface by the exact one, to find out how much of any shortfall belongs to the geometry.

what answers the ray, and with what colourrays answered wronglyheld-out PSNR
the cloud of 884 points, each with its own colour31.3%16.6 dB
the fused surface, one colour1.8%18.1 dB
the true surface, one colour0%19.2 dB
the fused surface, colours copied from the nearest photograph1.8%26.9 dB
the true surface, colours copied0%28.6 dB

Geometry went from 31.3% of the rays wrong to 1.8%, and the score moved by 1.5 dB. The true surface painted with the best single colour reaches only 19.2 dB, so a model with no colour cannot pass this exam whatever its geometry. Copying colours helps (26.9 dB) and does not remove the real limit: the true surface with the same recipe scores 28.6 dB. Where does the rest go? 24 of the 1,536 pixels carry 94% of the squared error of the copied colours, and 23 of those 24 lie on silhouettes (the last on an edge of the crate, between a lit face and a shaded one). At a silhouette one pixel's ray lands on the statue and its neighbour's on the wall behind; in another photograph the same surface point falls in a pixel that is entirely statue or entirely background, depending on where that photograph's grid happens to lie. A pixel is not the colour of a point; it is whatever the ray through it meets first. The surface knows where; nothing in it knows what comes back.

What this lesson did not do
The surface has no colour, no light and no idea what a pixel is (lesson 7). Unobserved nodes stay unobserved: with one view 77.6% of the true boundary has a hole, with four 12.0%, with twelve 0.2%; a hole is an honest "unknown", not a guess. The fusion believes the poses it is given: moving each camera by 2 cm and 0.2° raises the statue error from 0.88 to 1.36 cm, and 5 cm with 0.5° to 4.23 cm, so the poses of lessons 4 and 5 must come first. A grid costs memory in proportion to the volume it covers, (extent/h)d nodes in d dimensions, while the surface occupies a thin shell of them (lesson 10). The scene is assumed still (lesson 14), and the sensor's noise Gaussian and independent from pixel to pixel.

Common mistakes / failure modes

"more points will fix it"
More points close the gaps and leave the wrong depths: with every point exactly on the surface, 15.9% of the rays are still wrong at 884 points. With noise, more points are worse (§1).
"a view can vote on every node"
Not behind the first surface it sees. Making hidden nodes vote "inside" deletes the statue; the view must abstain (§3).
"average the depth maps"
Pixels of different cameras do not correspond. Average in a shared grid, with inverse-variance weights: equal weights cost 24% more error (§4).
"a wider band is safer"
Past half the smallest gap, neighbouring surfaces start to merge, and a thin plate swells to the band's width (§5).
"W = 0 means outside"
It means unknown. A ray that crosses unobserved nodes has not seen free space (§4).

Checkpoint exercise

Try it
A node is seen by three cameras. Their depth noise at that node is σ = 2, 3 and 4 cm, and their votes are d = +2.0, −1.0 and +4.0 cm. Compute the weights, the fused value D, the weight W, and the standard deviation of D; then compare with the plain average and its standard deviation. Answer: w = 1/σ² = 0.250, 0.111 and 0.0625 per cm², so W = 0.4236 and D = (0.250·2.0 + 0.111·(−1.0) + 0.0625·4.0)/W = 1.51 cm, with standard deviation 1/√W = 1.54 cm. The plain average is 1.67 cm with standard deviation √(4 + 9 + 16)/3 = 1.80 cm: the most reliable camera carries 59% of the weight, the least reliable 15%.

Where this points next

Noisy depth has become a surface that answers where each ray ends and which way it faces there: 1.8% of the held-out rays are wrong where the cloud missed 31.3%. It cannot say what comes back along the ray. Painted in the best single colour even the exact surface scores only 19.2 dB, and colours copied from photographs leave the error at silhouettes, where a pixel is entirely object or entirely background by the luck of its grid. What does a pixel actually measure, and how do we write down, and score, the image a model predicts?

Takeaway
A cloud of points cannot say what a ray hits: with every point exactly on the surface 15.9% of the held-out rays are answered wrongly, and with noisy depth more views make it worse. A ray needs a function of position whose sign tells inside from outside and whose value is a distance, estimated by letting every depth map vote d = zm − zc at every cell of a grid, clipped in front and withheld behind. For independent Gaussian votes the maximum-likelihood combination is the mean weighted by 1/σ², so the error falls as σ/√n down to a floor set by the depth maps' pixels; the band μ belongs between about 3σ and half the narrowest gap. The wrong rays fall from 31.3% to 1.8%, yet the score rises only from 16.6 to 18.1 dB: the surface has no colour.

Interview prompts

Companion reads: Computer Graphics · 03 Geometry, meshes and curves (implicit surfaces and marching cubes as a graphics tool; here the field is estimated, not given) and Computer Vision · 05 Multi-view, depth and SLAM (the stereo depth fused here).