all_lessons/3D Vision/08 · Making visibility soft: volume renderinglesson 8 / 14

Making visibility soft: volume rendering

Lesson 7 could render a model and score it against a photograph, but it could not improve the model by gradient: a pixel is a cliff in the scene, so the derivative of the loss is zero almost everywhere. This lesson replaces "the first surface the ray meets" by an average over where the ray might stop in a fog of absorbing, glowing particles. The pixel becomes a smooth function of the density and the colour at every point of its ray; a solid surface is the limit of a very dense thin sheet of fog; and the derivative has a closed form with a simple reading, what a sample adds, minus what it hides. The price is a blur whose width is a knob you must set, and a scene that is now a density and a colour at every point of space.

The thesis, here
A pixel does not have to see exactly one surface. Let the ray be stopped, with some probability per metre, by particles that each hand back a colour, and let the pixel be the expected colour. That one change makes the image a smooth function of the scene, contains the hard renderer as the limit of a very dense thin layer, and gives every point along every ray, as far as light reaches, a gradient.
Linear position
Forced by: We can render a model and score it against a photograph, but we cannot yet improve it by gradient: where one surface hides another, a pixel is a cliff in the geometry, so the loss is flat almost everywhere and says nothing about which way to move. What would a renderer look like whose pixels change smoothly with the scene, and still reduce to the hard one when the scene is solid?
New idea: the pixel is the expected colour of whatever stops the ray in a fog: its weights are w(t) = T(t)·σ(t) with transmittance T(t) = exp(−∫σ). A solid is the limit σ → ∞ of a thin sheet, and the derivative with respect to the density at any sample is what that sample adds minus what it hides.
Forces next: Volume rendering makes every pixel a smooth function of the density and color along its ray, and a solid surface is simply its sharp limit. It still leaves the unknown: a density and a color at every point of space, which gradient descent on photographs must find, in a function class flexible enough for fine detail and smooth enough to train. What should that function be?
The plan
Seven moves. (1) Measure what is wrong with "first hit" on the disc of lesson 7. (2) Derive the fog: transmittance, stopping weights, expected colour. (3) Discretise it, and meet alpha compositing. (4) Take the limit and recover solids. (5) Differentiate: ∂C/∂σk. (6) Put it to work on the disc. (7) Find what the softness knob buys and what it costs.

1 · What is wrong with "first hit", in numbers

Lesson 7 left one experiment standing. A disc 1 m in radius is photographed by six cameras on a ring of radius 5 m. The model is the same disc, whose centre is off by δ metres along x, and the loss E(δ) is the mean squared difference between the model's pixels and the photographs. With the hard renderer, a pixel shows the disc's colour if the ray meets the disc and the background if it does not. Sweep δ from −1.6 m to 1.6 m in steps of one millimetre and look at the loss: it takes only 59 distinct values, and in 96.4% of the steps it does not change at all. Its derivative is exactly zero there. A gradient method standing on one of these steps sees a slope of 0, so it does not move, however far the disc is from where it belongs.

The staircase has a cause that anti-aliasing does not remove. Lesson 7 gave the model's own edge a slope by averaging each pixel over its footprint, but that slope lives only in the few pixels that contain the edge. Which surface a ray meets first is a discrete choice, and so is whether it meets a surface at all; the pixel's colour depends on the scene only through that choice. Two more defects follow from it. A surface hidden behind another receives no signal until it has moved in front, and a model with nothing along a ray has no surface to pull towards the photograph. What we want instead is a pixel that depends smoothly on everything along its ray, and that is still the hard pixel when the scene is solid.

2 · A fog that stops light

Replace the surface by a medium. Along a ray, parametrised by the depth t in metres, suppose there are tiny particles, and let σ(t) (per metre) be the probability per metre that a ray at depth t is stopped by one. Let T(t) be the transmittance, the probability that the ray reaches depth t without having been stopped. Across a short step dt it is stopped with probability σ(t) dt, so

T(t + dt) = T(t)·(1 − σ(t) dt)  ⟹  dT/dt = −σ(t) T(t)  ⟹  T(t) = exp(−∫0t σ(s) ds)

The probability that the ray is stopped in the next dt is therefore T(t) σ(t) dt. Write w(t) = T(t) σ(t) for this stopping density: it is a probability distribution over where the ray ends, and it integrates to 1 − T(L) over a ray of length L, the probability of being stopped at all. Now let whatever stops the ray hand back its colour c(t), and let the pixel show the expected colour. A ray that survives to the end sees the background colour cbg:

C = ∫0L T(t) σ(t) c(t) dt + T(L)·cbg

That is the whole renderer (the emission–absorption model; light emitted by the particles along the ray, dimmed by the particles in front of them). Every quantity in it is a smooth function of σ and c. Nothing in it chooses a first surface.

3 · Discretise it: alpha compositing

A computer cannot integrate over a continuum of depths; it evaluates the field at N samples. Let sample i stand for a segment of length δi with constant density σi and colour ci. A ray survives the segment with probability e−σiδi, so the probability of being stopped inside it, given that it got there, is

αi = 1 − exp(−σi δi),   Ti = ∏j<i (1 − αj),   wi = Ti αi,   C = Σi wi ci + Tend·cbg,   Tend = ∏j≤N (1 − αj)

This is exactly the "over" operator of compositing translucent layers front to back, with one difference: graphics picks the opacities αi by hand, whereas here physics hands them over as a density and a length. A four-sample example, with δi = 0.5 m, colours (one channel) c = (0.9, 0.2, 0.9, 0.4) and a white background (1.0):

sample iσi (1/m)αiTiwi = Tiαici
100.0001.0000.0000.9
20.50.2211.0000.2210.2
320.6320.7790.4920.9
480.9820.2870.2810.4

The weights add up to 0.995, and the remaining 0.005 is Tend, the chance of getting through every sample unstopped: weights plus leftover are exactly 1. The pixel is C = 0.605. Read the table as a story: sample 3 is dense and bright, but sample 2 sits in front of it and is dim, so the pixel is a blend; sample 4 is very dense, but most of the ray never reaches it.

4 · Solids are the sharp limit

Make a surface out of fog: a sheet of thickness h and density σ0. Its opacity is α = 1 − e−σ0h, which tends to 1 as σ0 → ∞. A ray reaching the sheet is stopped inside it, the transmittance behind it is 0, and the weights pile up on the first sheet. A red sheet at 1.0 m in front of a blue one at 2.0 m (both 0.2 m thick, white background) gives, as the density grows:

sheet density σ0 (1/m)1101001000
colour error against pure red, |C − red|0.850.130.00000.0000
probability that the red sheet stops the ray0.180.861.001.00

Volume rendering is therefore not a rival of the hard renderer. It contains it: "first hit" is what it computes when the scene consists of infinitely dense thin sheets, and every finite density is a softened version. That is the second half of the baton's question, answered.

5 · Differentiate: what a sample adds, minus what it hides

Raise the density σk of one sample by a little. Two things change. Its own opacity grows, dαk/dσk = δk(1 − αk), so it contributes more of its colour: Tk δk (1 − αk) ck = δk Tk+1 ck. And every sample behind it, and the background, is seen through a thicker layer: each term Ti ∝ (1 − αk) for i > k shrinks by −δk Ti. Adding the two effects:

∂C/∂σk = δk · ( Tk+1 ck − Σi>k wi ci − Tend cbg ),    ∂C/∂ck = wk

The bracket reads "what sample k adds" minus "what it hides". If ck is brighter than the average of what lies behind it, thickening k brightens the pixel; if it is darker, thickening it darkens the pixel. Three consequences matter for what follows.

On the four-sample example of §3 (δi = 0.5 m) the formula gives:

sample kσkckadds: Tk+1ckhides: Σi>k wici + Tendcbg∂C/∂σk
10 (empty)0.90.9000.605+0.147
20.50.20.1560.561−0.203
320.90.2580.118+0.070
480.40.0020.005−0.002

Sample 1 has no density at all and still has a gradient, +0.147: it is bright, and the pixel would be brighter if a layer of it hid the dimmer samples behind. Sample 2 is dim and in front of brighter ones, so thickening it darkens the pixel, more strongly still. Sample 4 is the densest and has almost no gradient: it already stops 98% of the light that reaches it, so thickening it changes nothing, and almost nothing lies behind it to hide (only 29% of the ray gets as far as it). One more fact is exact. If every sample has the same colour c, the bracket equals Tend(c − cbg) for every k: the position of the density along the ray no longer matters, only its total. Geometry is carried by colour contrast along the ray, and by the other cameras whose rays cross this one.

6 · Put it to work: the disc, softly

Rebuild lesson 7's experiment with the soft renderer. The photographs are the same, six of them. The model disc becomes a fog sheet: the density around the circle of radius 1 m, at distance d from the model centre, is a Gaussian of width τ,

σ(d) = κ / (τ √(2π)) · exp(−(d − 1)² / (2τ²)),   κ = 6

so that the optical thickness across the sheet, ∫σ ds = κ, is the same for every τ and the sheet is opaque (1 − e−6 = 0.9975). Then τ is the softness: as τ → 0 the sheet is the hard circle of lesson 7; larger τ spreads the same opacity over a wider band. For each pixel the renderer samples the ray 64 times, composites, and the engine's backward sweep gives ∂C/∂σk; the chain rule through σk(δ) gives dE/dδ.

Soft or hard?
Left: the six cameras, the true disc (outline) and the model, a fog sheet of softness τ centred δ metres to the right of where it belongs. Top right: one camera's photograph, the soft render and the hard render of the model. Middle: the ray of the chosen pixel (the blue line in the left view and the vertical line on the strips): the density σ along it (grey), the transmittance T (blue line) and the stopping weights w (orange). Bottom: the loss E(δ) for the hard renderer (grey) and the soft one (orange). The buttons run 80 steps of Adam on δ with the exact derivative of each renderer.
E soft
—
dE/dδ soft
—
E hard
—
dE/dδ hard
—
ray: weight in its back half
—
last descent ends at
—
Show the core JS
for (k = 0; k < NS; k++) {
  var t = t0 + (k + 0.5) * DT, x = r.ox + r.dx * t, z = r.oz + r.dz * t, d = Math.hypot(x - delta, z), s = pk * Math.exp(-(d - R0) * (d - R0) / (2 * tau * tau));
  sg[k] = s; dl[k] = DT; col.push(ORANGE); ts[k] = t;
  dsd[k] = s * (-(d - R0) / (tau * tau)) * (delta - x) / (d + 1e-12);
}
var fwd = FL.vol.composite(sg, col, dl, BG), out = { C: fwd.C, fwd: fwd, sg: sg, ts: ts };
...
for (var k = n - 1; k >= 0; k--) {
  var Tk1 = f.Ts[k + 1], c = col[k];
  dsig[k] = delta[k] * (dLdC[0] * (Tk1 * c[0] - b0) + dLdC[1] * (Tk1 * c[1] - b1) + dLdC[2] * (Tk1 * c[2] - b2));
  b0 += f.w[k] * c[0]; b1 += f.w[k] * c[1]; b2 += f.w[k] * c[2];
}

What to try. The page opens with the model 1.5 m to the left of the true disc (δ = −1.5 m), softness τ = 0.15 m, and the ray of pixel 36 of one camera. In the ray panel the grey bumps are σ: the ray crosses the model's sheet twice, once going in and once coming out. The blue line is the transmittance; it falls to 0.0025 across the first crossing, so the orange weights sit there and only 0.24% of the stopping probability is left for the second. The back of the sheet is invisible to this pixel: a wall hides its own back. Now the derivative. At this offset the soft loss is 0.100 and its exact slope is −0.038 (negative: moving the model right lowers the loss, which is the correct direction); the hard loss is 0.100 and its slope is exactly 0. Press descend (hard): nothing happens. Press descend (soft): the model walks to within 0.06 m of the true centre. Start further out, δ = −2.8 m, where the hard loss is a flat 0.139 because the silhouettes barely overlap in most views: the soft slope there is only 0.011 in magnitude, yet Adam, which normalises the step by the size of the gradient, still brings the model to within 0.19 m, while the hard descent again stays put. Now slide τ upwards. At τ = 0.45 m the landscape has 2 false valleys, and a descent from −2.8 m ends at −1.3 m; at 0.6 m the landscape has inverted and the descent ends 1.9 m from the truth. Softness is a knob, and it can be turned too far.

Road not taken · keep the hard renderer and smooth its silhouettes
Soft rasterisers and edge-sampling renderers keep "first hit" and add a correction at the boundary of each shape. They give the model's own edge a gradient and that is a real gain for fitting a mesh whose topology is known. They do not give a hidden surface a way to come forward, and a model with nothing along a ray still has no edge to pull. Volume rendering supplies both because every sample on every ray is a variable. The other road, estimating the gradient of the discrete first-hit choice by sampling, gives unbiased but very noisy slopes, and the exact one in §5 costs less.

7 · What the softness costs, and what is still unknown

The widget exposes the trade. At τ = 0.03 the soft sheet behaves like the hard circle and the gradient is exact but local. At τ = 0.15–0.3 m the same opacity is spread over a band and the landscape has one valley at the truth. Beyond that the blur starts to change the answer: the sheet's silhouette grows, the best-fitting position is biased, and false valleys appear. In practice the softness is annealed from large to small while the fit tracks the minimum. A second cost is that nothing here prefers a sharp surface: a fog spread over a thick band explains a photograph almost as well as a thin sheet, so where the views cannot tell them apart nothing removes the haze.

What this lesson did not do
It fitted one number, δ, for a known radius and a known colour. It did not say what object holds a density and a colour for every point of space (lesson 9), how to choose the sample positions along a ray, or how colour may depend on the viewing direction. Nothing in it removes a haze that the views cannot tell from a sheet (lesson 9 meets one). It did not make the renderer fast: 64 samples per ray for 288 rays is already 18 thousand evaluations per loss (lesson 10 returns to this). And the emission–absorption model has no shadows, scattering or reflections: it is a rendering of appearance for fitting, not of light transport.

Common mistakes / failure modes

"volume rendering is for clouds and smoke"
It is a differentiable stand-in for solids too. A solid surface is the limit of a thin sheet as its density grows without bound (§4), and the table there shows the convergence.
"density is opacity"
Opacity is α = 1 − e−σδ: it depends on the length of the segment as well as the density. Doubling the step doubles the optical thickness (§3).
"wi is the probability that the surface is at sample i"
wi is the probability that the ray is stopped at sample i. The weights sum to 1 − Tend; a field that spreads them out is a fog, not a surface (§2, §3).
"the softer the renderer, the easier the optimisation"
Up to a point. Too much softness biases the minimum, creates false valleys and finally inverts the landscape (§6 and §7).
"gradients reach every point on the ray"
They reach as deep as light does. Behind a nearly opaque layer T ≈ 0 and the bracket in §5 vanishes: a wall hides its own back.
"a smooth loss means a good answer"
A smooth loss has a gradient. Whether its minimum is the truth depends on the model: the over-soft sheet of §6 has a smooth loss and the wrong answer.

Checkpoint exercise

Try it
A ray has three samples of length 0.4 m with densities σ = (1, 3, 5) per metre, colours (one channel) (0.2, 0.6, 1.0) and background 0.5. Compute αi, Ti, wi and C. Then use §5 to compute ∂C/∂σ1 and say in words why its sign is what it is. Answer: α = (0.330, 0.699, 0.865), T = (1, 0.670, 0.202), w = (0.330, 0.468, 0.175), Tend = 0.027, so C = 0.535. The derivative is 0.4·(0.670·0.2 − (0.468·0.6 + 0.175·1.0) − 0.027·0.5) = −0.134: sample 1 is darker (0.2) than everything behind it, so thickening it hides brighter light and darkens the pixel.

Where this points next

Volume rendering turns "which surface does the ray meet?" into an expected colour that is smooth in every density and colour along the ray, and a solid is its sharp limit. In the widget the unknown was a single number for a known shape. A scene has no known shape. On a small grid of 64 × 64 nodes, the unknowns are 4096 densities and 12288 colour values, and a 512 × 512 × 512 grid in three dimensions would need 134217728 densities. §5 gives a gradient to every one of them that light reaches, so in principle those can be found from photographs. But a table of independent numbers has no preference for smooth or compact scenes and cannot represent detail finer than its spacing. What should that function be?

Takeaway
The hard renderer's loss is a staircase in the scene: its derivative is zero almost everywhere, and a hidden or missing surface gets no signal. Volume rendering replaces the first hit by the expected colour of whatever stops the ray in a fog of density σ: the transmittance is T = exp(−∫σ), the stopping weights are w = Tσ, and discretised this is alpha compositing with α = 1 − e−σδ. A solid is the limit of a thin sheet as σ → ∞, so nothing of the old renderer is lost. The derivative is ∂C/∂σk = δk(Tk+1ck − Σi>k wici − Tendcbg): what a sample adds minus what it hides. It is nonzero everywhere light reaches, including empty space in front of a surface, which lets a model create density where the photograph wants it. The knob is the softness of the sheet: none and the gradient is local, too much and the minimum is biased or false. What remains is to choose a function that holds a density and a colour for every point.

Interview prompts

Companion reads: Computer Graphics · 07 Light and the rendering equation (the emission–absorption model is a special case of radiance transport), Computer Graphics · 11 Path tracing (the same "expected colour over random paths" idea, with scattering), and Computer Vision · 05 Multi-view, depth and SLAM (the one-screen NeRF overview this lesson derives).