all_lessons/3D Vision/09 · Radiance fields: a scene as a functionlesson 9 / 14

Radiance fields: a scene as a function

Lesson 8 made a pixel a smooth function of the density and the colour at each point of its ray, and left the unknown: those two functions, everywhere in space. This lesson makes the unknown finite and trainable. Store the scene as a function of position with a few thousand parameters (a 40 × 40 grid here, a neural network in NeRF), render every photographed pixel through lesson 8's fog, compare, and send the error back to the parameters: lesson 5's explain–compare–update loop with a dense renderer. It derives what the function may depend on, why it needs detail at many scales, and where to sample it. From twelve photographs the field predicts cameras it never saw at 20.4 dB; from three it fits its photographs and ghosts the rest; and 83% of its samples could not change a pixel.

The thesis, here
A scene is a function that returns, at any point, how much it stops light and what colour it hands back. Photographs speak about that function only through lesson 8's renderer, so we find it the way lesson 5 found camera poses: render what we have photographs of, compare, and move the parameters along the gradient. What the function may depend on, how much detail it holds and where it is sampled are then three decisions about what makes that loop find the scene instead of merely fitting its pictures.
Linear position
Forced by: Volume rendering makes every pixel a smooth function of the density and color along its ray, and a solid surface is simply its sharp limit. It still leaves the unknown: a density and a color at every point of space, which gradient descent on photographs must find, in a function class flexible enough for fine detail and smooth enough to train. What should that function be?
New idea: a scene is a function F(x, d) → (σ, c) with finitely many parameters, trained by pushing the photographs backwards through lesson 8's renderer. Such a function is a radiance field. The renderer is the only link between it and the pictures, so the loop decides what it may depend on, how much detail it needs and where it is sampled.
Forces next: A radiance field trained on photographs can reproduce views it never saw, but only after a long training run, and each pixel costs hundreds of evaluations of a large network, almost all of them in empty space. Where is the waste, and what representation spends its effort only where the scene is?
The plan
Six moves. (1) Count the unknowns and turn the function into numbers. (2) Close the loop: sample, composite, compare, back-propagate. (3) Decide what the function may depend on. (4) Give it detail. (5) Run the exam, and watch it fail when the views are few. (6) Add up the bill.

1 · From a function to numbers

Lesson 8's pixel needs, at every sample of its ray, a density σ ≥ 0 per metre and a colour c ∈ [0, 1]3: in Flatland, two functions of position. The photographs offer 12 × 64 × 3 = 2,304 numbers. A function has infinitely many, and even a coarse grid over the scene (40 × 40 nodes of four numbers each, below) has 6,400. The data cannot pin the unknowns one by one, so what we assume about the function does much of the work. Lesson 5's loop needs finitely many numbers θ to differentiate with respect to, hence a rule that turns θ into the two functions; we ask three things of it.

Three candidates; the first two were trained by the loop of §2 for 150 steps on twelve data cameras and scored on those cameras and on the 24 held-out ones, the exam of lesson 7.

Rule θ ↦ fieldNumbersData / held-out (dB)What it costs
independent cells: a point takes the value of its nearest node6,40022.4 / 19.1blocky: the colour jumps at every cell border
bilinear: linear in each direction between four nodes6,40024.4 / 20.4detail has a scale, the node spacing (§4)
a network of position, as in NeRFits weights§4 runs a small onecompact and smooth; every sample is a forward pass

We run the grid, which trains in seconds. It has 40 × 40 nodes, 0.174 m apart over the square −3.4 to 3.4 m (lesson 8 priced a 64 × 64 grid, 4,096 densities; this one is smaller so that training takes seconds). Each node holds a raw density s and three raw colour values: 6,400 unknowns in all. The value at a point is Σj φj(x) θj, where the four nodes around x have weights φj (products of its fractional distances along the two axes: non-negative, adding to 1) and every other node has weight 0. The density is σ = softplus(s) = ln(1 + es), positive and smooth; the colour is the sigmoid of its raw value, between 0 and 1. The field starts nearly empty and grey: s = −5 everywhere, σ = 0.0067 per metre. A network replaces the table by weights and each lookup by a forward pass; the rest of the loop is unchanged.

2 · Close the loop

Take one pixel of one data camera and run lesson 5's three stages. Explain. Cast its ray, clip it to the box, cut the clipped part into N = 48 equal bins and draw one sample at a random place in each, afresh at every step (stratified sampling, as in NeRF (Mildenhall et al., 2020); here the samples are about as dense as the nodes, so it changes the exam little: 20.2 dB with samples at the bin centres, 20.4 dB with random ones). Read σk and ck at each sample from the field and composite as in lesson 8: C = Σk wk ck + Tend cbg. Compare. The loss L is the mean of (C − p)2 over every pixel of every data camera and the three channels, p being the photographed colour. Update. Move the parameters against the gradient of L.

The gradient is lesson 8's, carried two links further. Let e = ∂L/∂C be the pixel's error signal (three numbers, one per channel) and ŝ(xk) = Σj φj(xk) sj the interpolated raw density at sample k, so that σk = softplus(ŝ). The chain rule gives

∂L/∂sj = Σpixels Σk e · δk ( Tk+1 ck − Σi>k wi ci − Tend cbg ) · sigmoid(ŝ(xk)) · φj(xk)

and for a raw colour rj (red; green and blue alike) ∂L/∂rj = Σ er wk ck(1 − ck) φj(xk). Three factors multiply in the first: what the sample adds minus what it hides (lesson 8, dotted with e), the slope of the softplus, which is the sigmoid, and φj, the share of the sample that node j answers for. So every node near a ray before it is stopped is pushed, not only the node at the surface: density can be created in the air in front of a surface if the photograph wants that colour. And a node behind an opaque layer meets T ≈ 0 and gets almost no signal: the data hardly say what is there (lesson 8: a wall hides its own back).

A hand-derived gradient is easy to get wrong, so check it as lesson 8 did: central finite differences of a separately written loss agree with the analytic gradient to 8.5 digits of its largest entry, on all 196 raw numbers of a 7 × 7 field. The update is Adam, which steps each parameter against a running average of its gradient divided by that gradient's running root-mean-square, so even a tiny gradient gets a step of about the full size; the step is 0.15 on the raw values (NeRF: 5×10⁻⁴ falling to 5×10⁻⁵, 4096 rays per batch (Mildenhall et al., 2020)). Each step here uses every pixel of every data camera: 768 rays and 36,864 evaluations of the field for twelve cameras.

3 · What the function may depend on

A pixel records the radiance that leaves a point toward the camera, which for a glossy surface depends on the direction d of the ray. So let the function take both, F(x, d) → (σ, c), and ask whether the density may depend on d too.

Suppose it may. For every pixel of every data camera put a thin sheet of fog on that pixel's ray, of optical thickness 6 (opacity 0.9975, lesson 8), coloured with the pixel's photographed colour and switched on only for rays travelling in exactly that direction. Every data pixel is reproduced, at 56.8 dB. Every other ray, including every ray of every held-out camera, passes through and sees the background: the exam score is 4.6 dB, the empty scene's. Shuffle the pixels of the photographs and the construction fits them just as well. A near-perfect fit certifies nothing: a density that may differ with the direction can paint any photographs onto the cameras and know nothing about the scene.

The remedy is to take the freedom away: the density is a function of position alone, σ(x), the same for every ray through x. A sheet in front of one camera now stands in front of every camera that looks through it, and blocks them; the only way to explain all the photographs is a scene they share. On the same pixels, real and shuffled, 150 steps:

Model classFits the photographsFits the same pixels shuffledHeld-out cameras
σ(x, d), c(x, d) free: one sheet per pixel56.8 dB56.8 dB4.6 dB
σ(x), c(x) on the grid, 12 cameras24.4 dB12.1 dB20.4 dB
the same grid, 3 cameras32.9 dB21.8 dB9.2 dB

A class that fits scrambled photographs as well as real ones has learned nothing by fitting, and the sheets are such a class. The shared grid is far from it with twelve cameras; with three it fits scrambled pixels at 21.8 dB (§5).

Colour may keep the direction, since real surfaces change with the viewpoint, and NeRF (Mildenhall et al., 2020) lets it: its network computes σ from position alone and adds the direction only in the layers that produce the colour; dropping it costs 3.35 dB on its synthetic scenes (27.66 against 31.01 dB). That freedom has to be limited too: a colour that may change arbitrarily fast with direction could paint the photographs onto an opaque shell around the scene. NeRF's direction encoding has L = 4 frequencies, against L = 10 for position (§4): the colour may vary with direction, but slowly. In Flatland every surface is diffuse (lesson 7), so c depends on position only, and so does our field.

4 · Detail needs frequency

A smooth function class has a resolution built in. For the grid it is the node spacing h = 0.174 m: between nodes the function is a straight line, so a wave of period P is held with a worst-case error, as a share of its amplitude, of

e = 1 − cos(π h / P)

(worst when two nodes straddle a crest: both read A cos(πh/P), the line between them is flat, and the crest A is lost). The statue's five stripes (mean radius 1.15 m) have period 2π·1.15/5 = 1.45 m, 8.3 spacings, and lose up to 7.1%; the ball's seven (radius 0.5 m), 2π·0.5/7 = 0.45 m, 2.6 spacings, up to 66%. Two spacings per period is the limit: nodes on the zero crossings see nothing.

A finer grid holds more, at N2 numbers in the plane and N3 in space, but the photographs cannot pin it: one pixel at 6 m spans 6/56 = 0.107 m. Twelve cameras, 150 steps:

GridNumbersData camerasHeld-out cameras
20 × 201,60020.7 dB19.2 dB
40 × 406,40024.4 dB20.4 dB
80 × 8025,60027.0 dB19.6 dB

From 40 to 80 nodes the fit to the data gains 2.6 dB and the exam loses 0.8 dB: detail beyond what the photographs resolve is memorised, not learned.

A network has no node spacing. Its resolution is whatever its weights learn, and gradient descent learns slow structure first. One dimension shows it: the target is a sum of five sine waves with 1, 2, 4, 8 and 16 periods across p ∈ [−1, 1], sampled at 256 points; a network of two hidden layers of 32 ReLU units is trained by Adam for 1000 steps. NeRF's remedy (Mildenhall et al., 2020) is to feed it not p but γ(p) = (sin 20πp, cos 20πp, …, sin 2L−1πp, cos 2L−1πp), sines and cosines at geometrically spaced frequencies, 1, 2, …, 2L−1 periods across [−1, 1], so that the fast components are inputs and the network has only to combine them. (Equation 4 of the paper gives 2L numbers per coordinate: 60 for a position at L = 10 and 24 for a direction at L = 4; its code also appends the coordinates themselves, 63 and 27.) A component's error is the distance between its amplitude and phase in the fit and in the target, as a share of the target's (0 is perfect, 1 is nothing there); the last column is the rms error of the whole fit as a share of the target's rms.

Input to the networkWeights4 periods8 periods16 periodsrms error
p1,1530.090.520.7749%
γ(p), L = 21,2490.000.010.076.4%
γ(p), L = 41,3770.000.000.001.7%

Fed the coordinate, the network holds the slow components and misses half or more of the two fastest; with L = 4 every component is within 0.3%, and L = 5 does no better (2.6% rms). (L = 4 stops at 8 periods; the 16-period wave is twice that frequency, a product of two inputs: sin 2a = 2 sin a cos a.) NeRF uses L = 10 for position, with coordinates scaled to [−1, 1]. Its ablation gives 28.77 dB without γ, 30.59 with L = 5, 31.01 with L = 10 and 30.81 with L = 15, and its heuristic is that the gain stops once 2L exceeds the highest frequency in the images, about 1024 in its data (Mildenhall et al., 2020). The grid has one scale of detail built in; γ builds in a ladder of them.

Fit the statue from its photographs
Left: the scene from above, each cell of the field in its colour with its opacity, the true outlines as lines, data cameras blue, one held-out camera amber. Right: that camera's photograph, the field's prediction, their difference, and the 48 samples along its centre ray: density σ (grey), transmittance T (blue), weight w (orange). Bottom: exam score against steps on the data cameras (blue) and the 24 held-out cameras (amber); dashed, the nearest stored photograph. Changing the cameras, the grid or the photographs restarts from an empty field; 150 steps take a few seconds.
data cameras
—
held-out cameras
—
steps
—
field evaluations
—
samples with w < 1/255
—
density in the air
—
held-out, 16 even samples
—
held-out, 8 + 8 importance
—
Show the core JS
/* explain: one sample in each of the N bins, the field read through bilinear weights, then lesson 8's composite */
for (var m = 0; m < n; m++) {
  var t = t0 + (m + (rng ? rng() : 0.5)) * dt, x = ox + dx * t, z = oz + dz * t;
  this._stencil(x, z, m, idx, wt);
  ...
  for (var k = 0; k < 4; k++) {
    var id = idx[q + k], w = wt[q + k];
    s += w * this.s[id]; r += w * this.c[3 * id]; g += w * this.c[3 * id + 1]; b += w * this.c[3 * id + 2];
  }
  sig[m] = softplus(s); col[m] = [sigmoid(r), sigmoid(g), sigmoid(b)]; delta[m] = dt;
}
var f = FL.vol.composite(sig, col, delta, this.bg);
/* compare, and send the error back through the composite, the squashing and the interpolation weights */
var e0 = f.C[0] - gt[0], e1 = f.C[1] - gt[1], e2 = f.C[2] - gt[2];
var dL = [2 * e0 * wgt, 2 * e1 * wgt, 2 * e2 * wgt];
var gr = FL.vol.compositeGrad(sig, col, delta, this.bg, dL, f);
for (var mm = 0; mm < n; mm++) {
  var ds = gr.dsig[mm] * sigmoid(sraw[mm]);
  ...
  for (var kk = 0; kk < 4; kk++) {
    var ii = idx[qq + kk], ww = wt[qq + kk];
    this.gS[ii] += ww * ds; this.gC[3 * ii] += ww * dc0; this.gC[3 * ii + 1] += ww * dc1; this.gC[3 * ii + 2] += ww * dc2;
  }
}
/* update: Adam on every node */
this.ms[k] = b1 * this.ms[k] + (1 - b1) * g; this.vs[k] = b2 * this.vs[k] + (1 - b2) * g * g;
this.s[k] -= lr * (this.ms[k] / c1) / (Math.sqrt(this.vs[k] / c2) + 1e-8);
/* where to sample: draw the fine samples from the weights of a first pass, by inverting their cumulative sum */
for (k = 0; k < nc; k++) { w.push(f.w[k] + 1e-5); tot += w[k]; }
for (k = 0; k < nc; k++) cdf.push(cdf[k] + w[k] / tot);
for (j = 0; j < nf; j++) {
  u = (j + 0.5) / nf;
  while (b < nc - 1 && cdf[b + 1] < u) b++;
  fine.push(t0 + (b + (u - cdf[b]) / (cdf[b + 1] - cdf[b])) * dt);
}

What to try. The page opens with twelve data cameras, a 40 × 40 grid and the photographs as taken, on a nearly empty grey field: both scores are 4.9 dB. Press train 150 steps (or train 50 steps three times: the same result). The statue, the crate and the ball appear out of nothing, striped, with a grey inside that no ray sees. The data cameras score 24.4 dB and the held-out cameras 20.4 dB: above the nearest stored photograph, 14.0 dB (the dashed line, passed at step 40), and above the exact surface painted in one colour, 19.2 dB (lesson 7). Press again: 27.3 and 21.2 dB, mostly a better fit to the data. Now set data cameras to 3 and train 150 steps: 32.9 dB on the data and 9.2 on the held-out cameras, below the nearest stored photograph (9.7 dB), with wedges of haze along the data rays. Switch photographs to shuffled: three cameras are still fitted at 21.8 dB, twelve only at 12.1. Back at twelve cameras, as taken, try the other grids (§4's table). Three readouts belong to §6: after 150 steps 83.2% of the samples have weight below 1/255, and 16 even samples score 19.1 dB against 19.7 dB for 8 even plus 8 drawn from their weights.

Road not taken · store the rays
The tempting alternative never builds a scene: keep the photographs and answer a new camera from the stored ones. Copying the nearest stored photograph scores 14.0 dB (lesson 7), the field 20.4 dB. Whatever such a method blends, it has no density, so nothing knows where along a ray a thing lies, and storing enough rays to close the gap is expensive: NeRF's weights are 5 MB, the light-field method LLFF needs over 15 GB for one of the same synthetic scenes, more than 3,000 times as much (Mildenhall et al., 2020). Painting colours copied from the data onto a known surface reaches 34.6 dB (lesson 7), but needs the surface that lesson 6 had to build from depth; the field finds geometry and colour from pixels alone.

5 · The exam: how many views?

Fit on n data cameras evenly spaced on the ring, score the same 24 held-out cameras, and add the share of the density that sits in the air, more than 0.3 m (almost two node spacings) outside every true object:

Data camerasFit to the dataHeld-outNearest stored photographDensity in the air
332.9 dB9.2 dB9.7 dB61%
627.1 dB14.1 dB11.4 dB30%
1224.4 dB20.4 dB14.0 dB8.5%
2423.2 dB22.2 dB14.4 dB3.1%

Read it by what a ray tells the field. A pixel is a weighted mix of everything its ray crosses: it says that something along the ray has this colour, not where, and nothing about points off the ray (lesson 1). Only rays of different directions through a point can say where, as in lesson 2, and three cameras give each point at most three. A count shows how little that fixes: three photographs are 3 × 64 × 3 = 576 numbers and the grid has 6,400 parameters, so to first order at least 5,824 directions of the field can change without changing any photograph, and the data cannot choose among them. The loop fits what the photographs say; the rest is what its pushes along the rays, rescaled by Adam, leave behind: the haze of the picture, 61% of the density in the air. The field reproduces its three photographs (32.9 dB) and predicts the others at 9.2 dB, below copying the nearest one; the gap between fit and exam is 23.7 dB, against 4.1 dB with twelve cameras and 1.0 dB with 24.

The shuffled pixels of §3 show the same thing from the other side: the field's ability to fit nonsense falls as cameras are added (21.8, 12.1 and 10.3 dB at 3, 12 and 24) while it keeps fitting the truth. Training does not repair what the data leave open: from 150 to 300 steps the twelve-camera field gains 2.9 dB on the data and moves on the exam from 20.4 to 21.2 dB.

6 · The bill

The field works; now count what it cost. One pass over the data is 768 rays × 48 samples = 36,864 evaluations of the field, and the 150 steps of the main run are 5,529,600. Where do they go? Take each sample's stopping weight wk (lesson 8): a sample with wk < 1/255 moves its pixel by less than one grey level of an 8-bit image, whatever its colour. In the trained field 83.2% of the samples are below that line: 79.1 of every hundred because their own opacity 1 − e−σδ is under 1/255 (empty space: they stop nothing whatever light reaches them), 4.1 because what stands in front of them dims them. A ray carries on average 8.1 samples that matter, and 215 of the 768 rays, those that cross only background, carry none.

A uniform pass cannot skip them: it does not know where they are. NeRF's hierarchical sampling (Mildenhall et al., 2020) reuses a by-product of lesson 8's composite. The weights wk of a coarse pass add up to at most 1; normalised, they are a probability distribution along the ray. Read them as a density over the bins and draw the fine samples from it by inverting its cumulative sum (the last lines of the listing). NeRF takes 64 coarse and 128 fine samples per ray, 256 queries in all (64 through its coarse network, 64 + 128 = 192 through its fine one), and replacing that by 256 uniform samples costs about a decibel (30.06 against 31.01 dB). Here, in the field of the main run, on the 24 held-out cameras:

Samples per rayAll evenly spacedHalf evenly spaced, half drawn from their weights
1619.1 dB19.7 dB
2419.7 dB20.1 dB
3220.2 dB20.3 dB
4820.4 dB—

Drawn samples give 16 the score of 24 even ones, and 32 nearly the score of 48: a third fewer evaluations. A third, and not the 83% that carry nothing, because the first pass is uniform and pays for the empty space once. At NeRF's scale the arithmetic is 4096 rays per step and 256 queries per ray, 1,048,576 queries per step, for 100 to 300 thousand steps, about one to two days on one V100: 1.05 to 3.15 × 1011 queries for one scene, each a pass through a network of 8 layers of 256 units. Rendering one 800 × 800 image takes 640,000 rays × 256 = 163,840,000 queries, about 30 seconds on that GPU (Mildenhall et al., 2020). The waste is not an accident of the grid or of the network: both are defined everywhere, so both are paid for everywhere.

What this lesson did not do
It fitted one scene from enough views, and did not say what to do when they run out: three cameras gave 9.2 dB and ghosts that no training repairs (lesson 11). It did not make the work cheap: the bill of §6 is lesson 10's. Everything here is diffuse, so colour never had to depend on direction: that half of §3 is argued, not run. A pixel is treated as a line, so rendering at another resolution aliases (Mip-NeRF (Barron et al., 2021) casts a cone and encodes its volume instead of a point), and the scene is bounded (Mip-NeRF 360 (Barron et al., 2022) contracts space, samples with a small proposal network and adds a distortion regulariser). And the poses are given: from the generator for synthetic scenes, from COLMAP for real ones in NeRF (Mildenhall et al., 2020).

Common mistakes / failure modes

"a NeRF is a neural network"
The idea is the function and the loop. A network is one way to store the function; the grid here does the same job and trains in seconds. The network is compact and costs a forward pass per sample (§1).
"if it fits the photographs, it has learned the scene"
Sheets switched on by direction fit them to 57 dB and predict 4.6; the grid fits three cameras of scrambled pixels at 21.8 dB. Only held-out cameras say what was learned (§3, §5).
"the density can depend on direction too"
Then each camera can have its own fog and no scene is needed. Density is a function of position; only the colour may see the direction, and only slowly (§3).
"a finer grid, or more frequencies, always helps"
From 40 to 80 nodes the data fit gains 2.6 dB and the exam loses 0.8: the photographs do not resolve the extra detail (§4).
"more samples per ray means a better picture"
Where they go matters more than how many: 16 samples, half drawn from the weights, score 19.7 dB, as 24 even ones do; 48 even samples score 20.4 with only 8 carrying weight (§6).

Checkpoint exercise

Try it
A one-dimensional field has two nodes, 1 m apart, with raw densities sA = −1 and sB = 2. A sample lies a quarter of the way from A to B. (a) What are the interpolated raw density, the density σ there and the slope of the softplus? (b) Lesson 8's bracket and the error signal give ∂L/∂σ = −0.3 at this sample. How much of that push does each node receive, and in what ratio? Answer: (a) the interpolated value is 0.75·(−1) + 0.25·2 = −0.25; σ = ln(1 + e−0.25) = 0.576 per metre; the slope of the softplus is the sigmoid, 0.438. (b) ∂L/∂sA = −0.3 · 0.438 · 0.75 = −0.099 and ∂L/∂sB = −0.3 · 0.438 · 0.25 = −0.033. The nodes share the push in the ratio of their weights: B receives 0.33 of what A receives.

Where this points next

The field is what lesson 8's question asked for: a function with a few thousand numbers, found from pictures by gradient steps. From twelve photographs it predicts cameras it never saw at 20.4 dB, above the nearest stored photograph (14.0) and above the exact surface in one colour (19.2); from three it ghosts, and no training repairs that. What it cannot do is be cheap. Fitting twelve 64-pixel photographs took 5,529,600 evaluations of the field, and in every pass 83% of the samples carry a weight under 1/255, almost all of them in empty space. NeRF pays the same bill at another scale: 256 network queries per ray, 4096 rays per step, a day or two on one V100, and the waste in every query. Where is the waste, and what representation spends its effort only where the scene is?

Takeaway
A scene is a function of position (and of direction, for colour) that returns a density and a colour, and a parametrisation turns "find a function" into "find numbers": a 40 × 40 grid with bilinear interpolation here, a network in NeRF. The loop is lesson 5's with lesson 8's renderer: sample each pixel's ray, composite, compare, and send the error back through the composite, the squashing and the interpolation weights to the nodes. The density must depend on position alone, or sheets painted differently from each side fit any photographs and explain nothing; detail must be built in, as a grid spacing or as input frequencies; samples are drawn at random places and, in NeRF, from the weights of a first pass. With twelve cameras the field scores 20.4 dB on cameras it never saw; with three it fits its photographs at 32.9 dB and the others at 9.2, because each ray constrains the field only along itself. And 83% of its samples have a weight under 1/255: the cost is paid wherever the function is defined, which is everywhere.

Interview prompts

Companion reads: Computer Vision · 05 Multi-view, depth and SLAM (the one-screen NeRF overview this lesson derives), Computer Graphics · 06 Sampling and anti-aliasing (the Nyquist limit behind §4 and the jittered samples of §2) and Computer Graphics · 11 Monte Carlo path tracing (importance sampling, the idea of §6).