Making visibility soft: volume rendering
Lesson 7 could render a model and score it against a photograph, but it could not improve the model by gradient: a pixel is a cliff in the scene, so the derivative of the loss is zero almost everywhere. This lesson replaces "the first surface the ray meets" by an average over where the ray might stop in a fog of absorbing, glowing particles. The pixel becomes a smooth function of the density and the colour at every point of its ray; a solid surface is the limit of a very dense thin sheet of fog; and the derivative has a closed form with a simple reading, what a sample adds, minus what it hides. The price is a blur whose width is a knob you must set, and a scene that is now a density and a colour at every point of space.
New idea: the pixel is the expected colour of whatever stops the ray in a fog: its weights are w(t) = T(t)·σ(t) with transmittance T(t) = exp(−∫σ). A solid is the limit σ → ∞ of a thin sheet, and the derivative with respect to the density at any sample is what that sample adds minus what it hides.
Forces next: Volume rendering makes every pixel a smooth function of the density and color along its ray, and a solid surface is simply its sharp limit. It still leaves the unknown: a density and a color at every point of space, which gradient descent on photographs must find, in a function class flexible enough for fine detail and smooth enough to train. What should that function be?
1 · What is wrong with "first hit", in numbers
Lesson 7 left one experiment standing. A disc 1 m in radius is photographed by six cameras on a ring of radius 5 m. The model is the same disc, whose centre is off by δ metres along x, and the loss E(δ) is the mean squared difference between the model's pixels and the photographs. With the hard renderer, a pixel shows the disc's colour if the ray meets the disc and the background if it does not. Sweep δ from −1.6 m to 1.6 m in steps of one millimetre and look at the loss: it takes only 59 distinct values, and in 96.4% of the steps it does not change at all. Its derivative is exactly zero there. A gradient method standing on one of these steps sees a slope of 0, so it does not move, however far the disc is from where it belongs.
The staircase has a cause that anti-aliasing does not remove. Lesson 7 gave the model's own edge a slope by averaging each pixel over its footprint, but that slope lives only in the few pixels that contain the edge. Which surface a ray meets first is a discrete choice, and so is whether it meets a surface at all; the pixel's colour depends on the scene only through that choice. Two more defects follow from it. A surface hidden behind another receives no signal until it has moved in front, and a model with nothing along a ray has no surface to pull towards the photograph. What we want instead is a pixel that depends smoothly on everything along its ray, and that is still the hard pixel when the scene is solid.
2 · A fog that stops light
Replace the surface by a medium. Along a ray, parametrised by the depth t in metres, suppose there are tiny particles, and let σ(t) (per metre) be the probability per metre that a ray at depth t is stopped by one. Let T(t) be the transmittance, the probability that the ray reaches depth t without having been stopped. Across a short step dt it is stopped with probability σ(t) dt, so
T(t + dt) = T(t)·(1 − σ(t) dt) ⟹ dT/dt = −σ(t) T(t) ⟹ T(t) = exp(−∫0t σ(s) ds)
The probability that the ray is stopped in the next dt is therefore T(t) σ(t) dt. Write w(t) = T(t) σ(t) for this stopping density: it is a probability distribution over where the ray ends, and it integrates to 1 − T(L) over a ray of length L, the probability of being stopped at all. Now let whatever stops the ray hand back its colour c(t), and let the pixel show the expected colour. A ray that survives to the end sees the background colour cbg:
C = ∫0L T(t) σ(t) c(t) dt + T(L)·cbg
That is the whole renderer (the emission–absorption model; light emitted by the particles along the ray, dimmed by the particles in front of them). Every quantity in it is a smooth function of σ and c. Nothing in it chooses a first surface.
3 · Discretise it: alpha compositing
A computer cannot integrate over a continuum of depths; it evaluates the field at N samples. Let sample i stand for a segment of length δi with constant density σi and colour ci. A ray survives the segment with probability e−σiδi, so the probability of being stopped inside it, given that it got there, is
αi = 1 − exp(−σi δi), Ti = ∏j<i (1 − αj), wi = Ti αi, C = Σi wi ci + Tend·cbg, Tend = ∏j≤N (1 − αj)
This is exactly the "over" operator of compositing translucent layers front to back, with one difference: graphics picks the opacities αi by hand, whereas here physics hands them over as a density and a length. A four-sample example, with δi = 0.5 m, colours (one channel) c = (0.9, 0.2, 0.9, 0.4) and a white background (1.0):
| sample i | σi (1/m) | αi | Ti | wi = Tiαi | ci |
|---|---|---|---|---|---|
| 1 | 0 | 0.000 | 1.000 | 0.000 | 0.9 |
| 2 | 0.5 | 0.221 | 1.000 | 0.221 | 0.2 |
| 3 | 2 | 0.632 | 0.779 | 0.492 | 0.9 |
| 4 | 8 | 0.982 | 0.287 | 0.281 | 0.4 |
The weights add up to 0.995, and the remaining 0.005 is Tend, the chance of getting through every sample unstopped: weights plus leftover are exactly 1. The pixel is C = 0.605. Read the table as a story: sample 3 is dense and bright, but sample 2 sits in front of it and is dim, so the pixel is a blend; sample 4 is very dense, but most of the ray never reaches it.
4 · Solids are the sharp limit
Make a surface out of fog: a sheet of thickness h and density σ0. Its opacity is α = 1 − e−σ0h, which tends to 1 as σ0 → ∞. A ray reaching the sheet is stopped inside it, the transmittance behind it is 0, and the weights pile up on the first sheet. A red sheet at 1.0 m in front of a blue one at 2.0 m (both 0.2 m thick, white background) gives, as the density grows:
| sheet density σ0 (1/m) | 1 | 10 | 100 | 1000 |
|---|---|---|---|---|
| colour error against pure red, |C − red| | 0.85 | 0.13 | 0.0000 | 0.0000 |
| probability that the red sheet stops the ray | 0.18 | 0.86 | 1.00 | 1.00 |
Volume rendering is therefore not a rival of the hard renderer. It contains it: "first hit" is what it computes when the scene consists of infinitely dense thin sheets, and every finite density is a softened version. That is the second half of the baton's question, answered.
5 · Differentiate: what a sample adds, minus what it hides
Raise the density σk of one sample by a little. Two things change. Its own opacity grows, dαk/dσk = δk(1 − αk), so it contributes more of its colour: Tk δk (1 − αk) ck = δk Tk+1 ck. And every sample behind it, and the background, is seen through a thicker layer: each term Ti ∝ (1 − αk) for i > k shrinks by −δk Ti. Adding the two effects:
∂C/∂σk = δk · ( Tk+1 ck − Σi>k wi ci − Tend cbg ), ∂C/∂ck = wk
The bracket reads "what sample k adds" minus "what it hides". If ck is brighter than the average of what lies behind it, thickening k brightens the pixel; if it is darker, thickening it darkens the pixel. Three consequences matter for what follows.
- Every sample in front of the first opaque layer has a gradient. A sample in empty space in front of the surface has σk = 0 and still receives δk(…), a nonzero push to create density there if the photograph wants its colour. The hard renderer cannot do this at all.
- A wall hides its own back. Behind a nearly opaque layer T is near 0, so the bracket is near 0: samples there get no signal. Gradients reach as deep as light does.
- It is cheap. One backward sweep over the samples of a ray, with a running sum for the hidden term, gives all N derivatives; the engine checks it against finite differences.
On the four-sample example of §3 (δi = 0.5 m) the formula gives:
| sample k | σk | ck | adds: Tk+1ck | hides: Σi>k wici + Tendcbg | ∂C/∂σk |
|---|---|---|---|---|---|
| 1 | 0 (empty) | 0.9 | 0.900 | 0.605 | +0.147 |
| 2 | 0.5 | 0.2 | 0.156 | 0.561 | −0.203 |
| 3 | 2 | 0.9 | 0.258 | 0.118 | +0.070 |
| 4 | 8 | 0.4 | 0.002 | 0.005 | −0.002 |
Sample 1 has no density at all and still has a gradient, +0.147: it is bright, and the pixel would be brighter if a layer of it hid the dimmer samples behind. Sample 2 is dim and in front of brighter ones, so thickening it darkens the pixel, more strongly still. Sample 4 is the densest and has almost no gradient: it already stops 98% of the light that reaches it, so thickening it changes nothing, and almost nothing lies behind it to hide (only 29% of the ray gets as far as it). One more fact is exact. If every sample has the same colour c, the bracket equals Tend(c − cbg) for every k: the position of the density along the ray no longer matters, only its total. Geometry is carried by colour contrast along the ray, and by the other cameras whose rays cross this one.
6 · Put it to work: the disc, softly
Rebuild lesson 7's experiment with the soft renderer. The photographs are the same, six of them. The model disc becomes a fog sheet: the density around the circle of radius 1 m, at distance d from the model centre, is a Gaussian of width τ,
σ(d) = κ / (τ √(2π)) · exp(−(d − 1)² / (2τ²)), κ = 6
so that the optical thickness across the sheet, ∫σ ds = κ, is the same for every τ and the sheet is opaque (1 − e−6 = 0.9975). Then τ is the softness: as τ → 0 the sheet is the hard circle of lesson 7; larger τ spreads the same opacity over a wider band. For each pixel the renderer samples the ray 64 times, composites, and the engine's backward sweep gives ∂C/∂σk; the chain rule through σk(δ) gives dE/dδ.
What to try. The page opens with the model 1.5 m to the left of the true disc (δ = −1.5 m), softness τ = 0.15 m, and the ray of pixel 36 of one camera. In the ray panel the grey bumps are σ: the ray crosses the model's sheet twice, once going in and once coming out. The blue line is the transmittance; it falls to 0.0025 across the first crossing, so the orange weights sit there and only 0.24% of the stopping probability is left for the second. The back of the sheet is invisible to this pixel: a wall hides its own back. Now the derivative. At this offset the soft loss is 0.100 and its exact slope is −0.038 (negative: moving the model right lowers the loss, which is the correct direction); the hard loss is 0.100 and its slope is exactly 0. Press descend (hard): nothing happens. Press descend (soft): the model walks to within 0.06 m of the true centre. Start further out, δ = −2.8 m, where the hard loss is a flat 0.139 because the silhouettes barely overlap in most views: the soft slope there is only 0.011 in magnitude, yet Adam, which normalises the step by the size of the gradient, still brings the model to within 0.19 m, while the hard descent again stays put. Now slide τ upwards. At τ = 0.45 m the landscape has 2 false valleys, and a descent from −2.8 m ends at −1.3 m; at 0.6 m the landscape has inverted and the descent ends 1.9 m from the truth. Softness is a knob, and it can be turned too far.
7 · What the softness costs, and what is still unknown
The widget exposes the trade. At τ = 0.03 the soft sheet behaves like the hard circle and the gradient is exact but local. At τ = 0.15–0.3 m the same opacity is spread over a band and the landscape has one valley at the truth. Beyond that the blur starts to change the answer: the sheet's silhouette grows, the best-fitting position is biased, and false valleys appear. In practice the softness is annealed from large to small while the fit tracks the minimum. A second cost is that nothing here prefers a sharp surface: a fog spread over a thick band explains a photograph almost as well as a thin sheet, so where the views cannot tell them apart nothing removes the haze.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
Volume rendering turns "which surface does the ray meet?" into an expected colour that is smooth in every density and colour along the ray, and a solid is its sharp limit. In the widget the unknown was a single number for a known shape. A scene has no known shape. On a small grid of 64 × 64 nodes, the unknowns are 4096 densities and 12288 colour values, and a 512 × 512 × 512 grid in three dimensions would need 134217728 densities. §5 gives a gradient to every one of them that light reaches, so in principle those can be found from photographs. But a table of independent numbers has no preference for smooth or compact scenes and cannot represent detail finer than its spacing. What should that function be?
Interview prompts
- Derive the transmittance T(t) and the expected colour of a ray in the emission–absorption model. (§2 — survival over a short step multiplies by 1 − σ dt, which solves to exp(−∫σ); the colour is ∫ Tσc dt plus T(L) times the background.)
- What is αi in terms of density, and why does the exponential appear? (§3 — α = 1 − e−σδ is the probability of being stopped inside a segment given arrival; survival through a segment is e−σδ.)
- In what sense is a hard surface a limit of a volume? (§4 — a thin sheet with σ → ∞ has α → 1, the weights pile onto the first sheet, and C becomes the first colour; the table shows the error go to zero.)
- Derive ∂C/∂σk and interpret its terms. (§5 — increasing σk adds δkTk+1ck and hides everything behind it by δk(Σ wici + Tendcbg).)
- Why can a density model create an object from nothing while a hard-surface model cannot? (§5, Road — empty-space samples in front of the surface still have a nonzero gradient; a hard model has no surface to receive one.)
- Why is the first-hit renderer's derivative zero almost everywhere, and does anti-aliasing fix it? (§1, Road — the pixel depends on the scene through a discrete choice; anti-aliasing smooths the boundary of each shape only, not the order of hidden surfaces nor missing ones.)
- What goes wrong if the fog is made very soft? (§6, §7 — the sheet grows outward by about τ, the best-fit position is biased, false valleys appear and the landscape can invert.)
Companion reads: Computer Graphics · 07 Light and the rendering equation (the emission–absorption model is a special case of radiance transport), Computer Graphics · 11 Path tracing (the same "expected colour over random paths" idea, with scattering), and Computer Vision · 05 Multi-view, depth and SLAM (the one-screen NeRF overview this lesson derives).