Radiance fields: a scene as a function
Lesson 8 made a pixel a smooth function of the density and the colour at each point of its ray, and left the unknown: those two functions, everywhere in space. This lesson makes the unknown finite and trainable. Store the scene as a function of position with a few thousand parameters (a 40 × 40 grid here, a neural network in NeRF), render every photographed pixel through lesson 8's fog, compare, and send the error back to the parameters: lesson 5's explain–compare–update loop with a dense renderer. It derives what the function may depend on, why it needs detail at many scales, and where to sample it. From twelve photographs the field predicts cameras it never saw at 20.4 dB; from three it fits its photographs and ghosts the rest; and 83% of its samples could not change a pixel.
New idea: a scene is a function F(x, d) → (σ, c) with finitely many parameters, trained by pushing the photographs backwards through lesson 8's renderer. Such a function is a radiance field. The renderer is the only link between it and the pictures, so the loop decides what it may depend on, how much detail it needs and where it is sampled.
Forces next: A radiance field trained on photographs can reproduce views it never saw, but only after a long training run, and each pixel costs hundreds of evaluations of a large network, almost all of them in empty space. Where is the waste, and what representation spends its effort only where the scene is?
1 · From a function to numbers
Lesson 8's pixel needs, at every sample of its ray, a density σ ≥ 0 per metre and a colour c ∈ [0, 1]3: in Flatland, two functions of position. The photographs offer 12 × 64 × 3 = 2,304 numbers. A function has infinitely many, and even a coarse grid over the scene (40 × 40 nodes of four numbers each, below) has 6,400. The data cannot pin the unknowns one by one, so what we assume about the function does much of the work. Lesson 5's loop needs finitely many numbers θ to differentiate with respect to, hence a rule that turns θ into the two functions; we ask three things of it.
- Shared. A value serves every ray through its point: where two rays meet, two cameras constrain the same numbers (lesson 2's triangulation, with a field in place of a point). A separate table for each ray, 768 rays × 48 samples × 4 numbers = 147,456 for twelve cameras, reproduces every photograph and says nothing about any other ray.
- Defined everywhere and differentiable in θ. The rays of a camera we did not use pass through any point of the box, and lesson 8's gradient has to flow on into the parameters.
- Related between neighbours. One ray's information has to spread a little, and detail has to have a scale.
Three candidates; the first two were trained by the loop of §2 for 150 steps on twelve data cameras and scored on those cameras and on the 24 held-out ones, the exam of lesson 7.
| Rule θ ↦ field | Numbers | Data / held-out (dB) | What it costs |
|---|---|---|---|
| independent cells: a point takes the value of its nearest node | 6,400 | 22.4 / 19.1 | blocky: the colour jumps at every cell border |
| bilinear: linear in each direction between four nodes | 6,400 | 24.4 / 20.4 | detail has a scale, the node spacing (§4) |
| a network of position, as in NeRF | its weights | §4 runs a small one | compact and smooth; every sample is a forward pass |
We run the grid, which trains in seconds. It has 40 × 40 nodes, 0.174 m apart over the square −3.4 to 3.4 m (lesson 8 priced a 64 × 64 grid, 4,096 densities; this one is smaller so that training takes seconds). Each node holds a raw density s and three raw colour values: 6,400 unknowns in all. The value at a point is Σj φj(x) θj, where the four nodes around x have weights φj (products of its fractional distances along the two axes: non-negative, adding to 1) and every other node has weight 0. The density is σ = softplus(s) = ln(1 + es), positive and smooth; the colour is the sigmoid of its raw value, between 0 and 1. The field starts nearly empty and grey: s = −5 everywhere, σ = 0.0067 per metre. A network replaces the table by weights and each lookup by a forward pass; the rest of the loop is unchanged.
2 · Close the loop
Take one pixel of one data camera and run lesson 5's three stages. Explain. Cast its ray, clip it to the box, cut the clipped part into N = 48 equal bins and draw one sample at a random place in each, afresh at every step (stratified sampling, as in NeRF (Mildenhall et al., 2020); here the samples are about as dense as the nodes, so it changes the exam little: 20.2 dB with samples at the bin centres, 20.4 dB with random ones). Read σk and ck at each sample from the field and composite as in lesson 8: C = Σk wk ck + Tend cbg. Compare. The loss L is the mean of (C − p)2 over every pixel of every data camera and the three channels, p being the photographed colour. Update. Move the parameters against the gradient of L.
The gradient is lesson 8's, carried two links further. Let e = ∂L/∂C be the pixel's error signal (three numbers, one per channel) and ŝ(xk) = Σj φj(xk) sj the interpolated raw density at sample k, so that σk = softplus(ŝ). The chain rule gives
∂L/∂sj = Σpixels Σk e · δk ( Tk+1 ck − Σi>k wi ci − Tend cbg ) · sigmoid(ŝ(xk)) · φj(xk)
and for a raw colour rj (red; green and blue alike) ∂L/∂rj = Σ er wk ck(1 − ck) φj(xk). Three factors multiply in the first: what the sample adds minus what it hides (lesson 8, dotted with e), the slope of the softplus, which is the sigmoid, and φj, the share of the sample that node j answers for. So every node near a ray before it is stopped is pushed, not only the node at the surface: density can be created in the air in front of a surface if the photograph wants that colour. And a node behind an opaque layer meets T ≈ 0 and gets almost no signal: the data hardly say what is there (lesson 8: a wall hides its own back).
A hand-derived gradient is easy to get wrong, so check it as lesson 8 did: central finite differences of a separately written loss agree with the analytic gradient to 8.5 digits of its largest entry, on all 196 raw numbers of a 7 × 7 field. The update is Adam, which steps each parameter against a running average of its gradient divided by that gradient's running root-mean-square, so even a tiny gradient gets a step of about the full size; the step is 0.15 on the raw values (NeRF: 5×10⁻⁴ falling to 5×10⁻⁵, 4096 rays per batch (Mildenhall et al., 2020)). Each step here uses every pixel of every data camera: 768 rays and 36,864 evaluations of the field for twelve cameras.
3 · What the function may depend on
A pixel records the radiance that leaves a point toward the camera, which for a glossy surface depends on the direction d of the ray. So let the function take both, F(x, d) → (σ, c), and ask whether the density may depend on d too.
Suppose it may. For every pixel of every data camera put a thin sheet of fog on that pixel's ray, of optical thickness 6 (opacity 0.9975, lesson 8), coloured with the pixel's photographed colour and switched on only for rays travelling in exactly that direction. Every data pixel is reproduced, at 56.8 dB. Every other ray, including every ray of every held-out camera, passes through and sees the background: the exam score is 4.6 dB, the empty scene's. Shuffle the pixels of the photographs and the construction fits them just as well. A near-perfect fit certifies nothing: a density that may differ with the direction can paint any photographs onto the cameras and know nothing about the scene.
The remedy is to take the freedom away: the density is a function of position alone, σ(x), the same for every ray through x. A sheet in front of one camera now stands in front of every camera that looks through it, and blocks them; the only way to explain all the photographs is a scene they share. On the same pixels, real and shuffled, 150 steps:
| Model class | Fits the photographs | Fits the same pixels shuffled | Held-out cameras |
|---|---|---|---|
| σ(x, d), c(x, d) free: one sheet per pixel | 56.8 dB | 56.8 dB | 4.6 dB |
| σ(x), c(x) on the grid, 12 cameras | 24.4 dB | 12.1 dB | 20.4 dB |
| the same grid, 3 cameras | 32.9 dB | 21.8 dB | 9.2 dB |
A class that fits scrambled photographs as well as real ones has learned nothing by fitting, and the sheets are such a class. The shared grid is far from it with twelve cameras; with three it fits scrambled pixels at 21.8 dB (§5).
Colour may keep the direction, since real surfaces change with the viewpoint, and NeRF (Mildenhall et al., 2020) lets it: its network computes σ from position alone and adds the direction only in the layers that produce the colour; dropping it costs 3.35 dB on its synthetic scenes (27.66 against 31.01 dB). That freedom has to be limited too: a colour that may change arbitrarily fast with direction could paint the photographs onto an opaque shell around the scene. NeRF's direction encoding has L = 4 frequencies, against L = 10 for position (§4): the colour may vary with direction, but slowly. In Flatland every surface is diffuse (lesson 7), so c depends on position only, and so does our field.
4 · Detail needs frequency
A smooth function class has a resolution built in. For the grid it is the node spacing h = 0.174 m: between nodes the function is a straight line, so a wave of period P is held with a worst-case error, as a share of its amplitude, of
e = 1 − cos(π h / P)
(worst when two nodes straddle a crest: both read A cos(πh/P), the line between them is flat, and the crest A is lost). The statue's five stripes (mean radius 1.15 m) have period 2π·1.15/5 = 1.45 m, 8.3 spacings, and lose up to 7.1%; the ball's seven (radius 0.5 m), 2π·0.5/7 = 0.45 m, 2.6 spacings, up to 66%. Two spacings per period is the limit: nodes on the zero crossings see nothing.
A finer grid holds more, at N2 numbers in the plane and N3 in space, but the photographs cannot pin it: one pixel at 6 m spans 6/56 = 0.107 m. Twelve cameras, 150 steps:
| Grid | Numbers | Data cameras | Held-out cameras |
|---|---|---|---|
| 20 × 20 | 1,600 | 20.7 dB | 19.2 dB |
| 40 × 40 | 6,400 | 24.4 dB | 20.4 dB |
| 80 × 80 | 25,600 | 27.0 dB | 19.6 dB |
From 40 to 80 nodes the fit to the data gains 2.6 dB and the exam loses 0.8 dB: detail beyond what the photographs resolve is memorised, not learned.
A network has no node spacing. Its resolution is whatever its weights learn, and gradient descent learns slow structure first. One dimension shows it: the target is a sum of five sine waves with 1, 2, 4, 8 and 16 periods across p ∈ [−1, 1], sampled at 256 points; a network of two hidden layers of 32 ReLU units is trained by Adam for 1000 steps. NeRF's remedy (Mildenhall et al., 2020) is to feed it not p but γ(p) = (sin 20πp, cos 20πp, …, sin 2L−1πp, cos 2L−1πp), sines and cosines at geometrically spaced frequencies, 1, 2, …, 2L−1 periods across [−1, 1], so that the fast components are inputs and the network has only to combine them. (Equation 4 of the paper gives 2L numbers per coordinate: 60 for a position at L = 10 and 24 for a direction at L = 4; its code also appends the coordinates themselves, 63 and 27.) A component's error is the distance between its amplitude and phase in the fit and in the target, as a share of the target's (0 is perfect, 1 is nothing there); the last column is the rms error of the whole fit as a share of the target's rms.
| Input to the network | Weights | 4 periods | 8 periods | 16 periods | rms error |
|---|---|---|---|---|---|
| p | 1,153 | 0.09 | 0.52 | 0.77 | 49% |
| γ(p), L = 2 | 1,249 | 0.00 | 0.01 | 0.07 | 6.4% |
| γ(p), L = 4 | 1,377 | 0.00 | 0.00 | 0.00 | 1.7% |
Fed the coordinate, the network holds the slow components and misses half or more of the two fastest; with L = 4 every component is within 0.3%, and L = 5 does no better (2.6% rms). (L = 4 stops at 8 periods; the 16-period wave is twice that frequency, a product of two inputs: sin 2a = 2 sin a cos a.) NeRF uses L = 10 for position, with coordinates scaled to [−1, 1]. Its ablation gives 28.77 dB without γ, 30.59 with L = 5, 31.01 with L = 10 and 30.81 with L = 15, and its heuristic is that the gain stops once 2L exceeds the highest frequency in the images, about 1024 in its data (Mildenhall et al., 2020). The grid has one scale of detail built in; γ builds in a ladder of them.
What to try. The page opens with twelve data cameras, a 40 × 40 grid and the photographs as taken, on a nearly empty grey field: both scores are 4.9 dB. Press train 150 steps (or train 50 steps three times: the same result). The statue, the crate and the ball appear out of nothing, striped, with a grey inside that no ray sees. The data cameras score 24.4 dB and the held-out cameras 20.4 dB: above the nearest stored photograph, 14.0 dB (the dashed line, passed at step 40), and above the exact surface painted in one colour, 19.2 dB (lesson 7). Press again: 27.3 and 21.2 dB, mostly a better fit to the data. Now set data cameras to 3 and train 150 steps: 32.9 dB on the data and 9.2 on the held-out cameras, below the nearest stored photograph (9.7 dB), with wedges of haze along the data rays. Switch photographs to shuffled: three cameras are still fitted at 21.8 dB, twelve only at 12.1. Back at twelve cameras, as taken, try the other grids (§4's table). Three readouts belong to §6: after 150 steps 83.2% of the samples have weight below 1/255, and 16 even samples score 19.1 dB against 19.7 dB for 8 even plus 8 drawn from their weights.
5 · The exam: how many views?
Fit on n data cameras evenly spaced on the ring, score the same 24 held-out cameras, and add the share of the density that sits in the air, more than 0.3 m (almost two node spacings) outside every true object:
| Data cameras | Fit to the data | Held-out | Nearest stored photograph | Density in the air |
|---|---|---|---|---|
| 3 | 32.9 dB | 9.2 dB | 9.7 dB | 61% |
| 6 | 27.1 dB | 14.1 dB | 11.4 dB | 30% |
| 12 | 24.4 dB | 20.4 dB | 14.0 dB | 8.5% |
| 24 | 23.2 dB | 22.2 dB | 14.4 dB | 3.1% |
Read it by what a ray tells the field. A pixel is a weighted mix of everything its ray crosses: it says that something along the ray has this colour, not where, and nothing about points off the ray (lesson 1). Only rays of different directions through a point can say where, as in lesson 2, and three cameras give each point at most three. A count shows how little that fixes: three photographs are 3 × 64 × 3 = 576 numbers and the grid has 6,400 parameters, so to first order at least 5,824 directions of the field can change without changing any photograph, and the data cannot choose among them. The loop fits what the photographs say; the rest is what its pushes along the rays, rescaled by Adam, leave behind: the haze of the picture, 61% of the density in the air. The field reproduces its three photographs (32.9 dB) and predicts the others at 9.2 dB, below copying the nearest one; the gap between fit and exam is 23.7 dB, against 4.1 dB with twelve cameras and 1.0 dB with 24.
The shuffled pixels of §3 show the same thing from the other side: the field's ability to fit nonsense falls as cameras are added (21.8, 12.1 and 10.3 dB at 3, 12 and 24) while it keeps fitting the truth. Training does not repair what the data leave open: from 150 to 300 steps the twelve-camera field gains 2.9 dB on the data and moves on the exam from 20.4 to 21.2 dB.
6 · The bill
The field works; now count what it cost. One pass over the data is 768 rays × 48 samples = 36,864 evaluations of the field, and the 150 steps of the main run are 5,529,600. Where do they go? Take each sample's stopping weight wk (lesson 8): a sample with wk < 1/255 moves its pixel by less than one grey level of an 8-bit image, whatever its colour. In the trained field 83.2% of the samples are below that line: 79.1 of every hundred because their own opacity 1 − e−σδ is under 1/255 (empty space: they stop nothing whatever light reaches them), 4.1 because what stands in front of them dims them. A ray carries on average 8.1 samples that matter, and 215 of the 768 rays, those that cross only background, carry none.
A uniform pass cannot skip them: it does not know where they are. NeRF's hierarchical sampling (Mildenhall et al., 2020) reuses a by-product of lesson 8's composite. The weights wk of a coarse pass add up to at most 1; normalised, they are a probability distribution along the ray. Read them as a density over the bins and draw the fine samples from it by inverting its cumulative sum (the last lines of the listing). NeRF takes 64 coarse and 128 fine samples per ray, 256 queries in all (64 through its coarse network, 64 + 128 = 192 through its fine one), and replacing that by 256 uniform samples costs about a decibel (30.06 against 31.01 dB). Here, in the field of the main run, on the 24 held-out cameras:
| Samples per ray | All evenly spaced | Half evenly spaced, half drawn from their weights |
|---|---|---|
| 16 | 19.1 dB | 19.7 dB |
| 24 | 19.7 dB | 20.1 dB |
| 32 | 20.2 dB | 20.3 dB |
| 48 | 20.4 dB | — |
Drawn samples give 16 the score of 24 even ones, and 32 nearly the score of 48: a third fewer evaluations. A third, and not the 83% that carry nothing, because the first pass is uniform and pays for the empty space once. At NeRF's scale the arithmetic is 4096 rays per step and 256 queries per ray, 1,048,576 queries per step, for 100 to 300 thousand steps, about one to two days on one V100: 1.05 to 3.15 × 1011 queries for one scene, each a pass through a network of 8 layers of 256 units. Rendering one 800 × 800 image takes 640,000 rays × 256 = 163,840,000 queries, about 30 seconds on that GPU (Mildenhall et al., 2020). The waste is not an accident of the grid or of the network: both are defined everywhere, so both are paid for everywhere.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
The field is what lesson 8's question asked for: a function with a few thousand numbers, found from pictures by gradient steps. From twelve photographs it predicts cameras it never saw at 20.4 dB, above the nearest stored photograph (14.0) and above the exact surface in one colour (19.2); from three it ghosts, and no training repairs that. What it cannot do is be cheap. Fitting twelve 64-pixel photographs took 5,529,600 evaluations of the field, and in every pass 83% of the samples carry a weight under 1/255, almost all of them in empty space. NeRF pays the same bill at another scale: 256 network queries per ray, 4096 rays per step, a day or two on one V100, and the waste in every query. Where is the waste, and what representation spends its effort only where the scene is?
Interview prompts
- Why must the density of a radiance field not depend on the viewing direction, while the colour may? (§3 — otherwise a sheet per pixel, switched on only along its own ray, reproduces any photographs and predicts nothing; real surfaces do change colour with direction, but slowly.)
- Derive the gradient of the loss with respect to the raw density of a grid node. (§2 — three factors: lesson 8's adds-minus-hides bracket, the softplus slope sigmoid(ŝ), and the bilinear weight of the node.)
- Does a node behind an opaque surface learn anything? (§2 — the factor T ≈ 0 makes the bracket vanish, so the photographs say nothing about it: a wall hides its own back.)
- Why does a network fed raw coordinates fit fine detail badly, and what does γ do? (§4 — it learns slow components first; sines and cosines at frequencies 2kπ hand it the fast ones at the input.)
- A finer grid fits the training photographs better. Why is that not a better model? (§4 — the photographs cannot resolve the extra detail, so the grid memorises it: the data fit improves and the held-out score drops.)
- Why does a radiance field fitted to three views score 33 dB on its photographs and 9 on the rest? (§5 — each ray constrains the field only along itself; 576 numbers cannot fix 6,400 parameters, at least 5,824 directions change no photograph, and the loop leaves haze along the rays there.)
- How does hierarchical sampling work, and what does it leave wasted? (§6 — the weights of a coarse pass are a distribution along the ray; the fine samples are drawn from it by inverting the cumulative sum; the coarse pass is uniform, so it still pays for the empty space.)
Companion reads: Computer Vision · 05 Multi-view, depth and SLAM (the one-screen NeRF overview this lesson derives), Computer Graphics · 06 Sampling and anti-aliasing (the Nyquist limit behind §4 and the jittered samples of §2) and Computer Graphics · 11 Monte Carlo path tracing (importance sampling, the idea of §6).