all_lessons/3D Vision/10 · Speed: grids, hashes and Gaussian splatslesson 10 / 14

Speed: grids, hashes and Gaussian splats

Lesson 9 fitted a scene from twelve photographs with 5,529,600 evaluations of its field, and 83% of them could not change a pixel. A surface touches about Nd−1 of the Nd cells of a grid, so cost can follow the scene instead of the box: skip the empty cells with an occupancy mask, share the occupied ones through a hash table read by a tiny decoder, or drop the grid and keep Gaussians that project to Gaussians and composite by lesson 8's rule. Forty Gaussians (360 numbers) score 22.1 dB on 24 cameras they never saw. None of this makes the photographs sufficient: from three views the Gaussians score 29.5 dB on their own photographs and 14.2 on the rest.

The thesis, here
A scene fills a sliver of the space that holds it: a surface at resolution N touches about Nd−1 of Nd cells. A representation that pays per cell pays for air; one that pays per occupied cell, or per object, pays for the scene. Speed is therefore a decision about where numbers and arithmetic are spent, and it changes nothing about what the photographs leave undetermined.
Linear position
Forced by: A radiance field trained on photographs can reproduce views it never saw, but only after a long training run, and each pixel costs hundreds of evaluations of a large network, almost all of them in empty space. Where is the waste, and what representation spends its effort only where the scene is?
New idea: spend numbers and arithmetic only where the scene is: skip the empty cells, hash the occupied ones into a small table, or store the scene as soft primitives that carry their own extent. Each makes cost follow the surface, not the box; none gives the photographs more to say.
Forces next: Grids and Gaussians fit one scene in minutes and draw it in real time, but only by being shown that scene from many views. Give them three views, or one, and the unseen part is unconstrained: they score well on the photographs they were fitted to and fail the exam. Where can the missing information come from, if not from more photographs?
The plan
Seven moves. (1) Itemise the bill of lesson 9. (2) Count what a scene occupies. (3) Skip the empty cells with a mask. (4) Hash the occupied ones into a small table. (5) Replace the grid by Gaussians: project, composite, differentiate, grow. (6) Fit them and price them. (7) Show them three views.

1 · The bill, itemised

Lesson 9's run was 150 steps over 768 rays of 48 samples: 36,864 evaluations of the field per step. A sample whose stopping weight w = Tα (lesson 8) is below 1/255 moves its pixel by less than one grey level, whatever its colour, and 83.2% of the samples are below that line. Of every hundred, 79.1 have an opacity α under 1/255 on their own: they stop nothing. Another 4.1 are visible but dimmed by what stands in front of them. Only 1.7 lie behind the point where the transmittance T has fallen under 1/255 (they are among the 83.2), so ending a ray once the light has gone would save almost nothing: the waste is samples that stop nothing.

What is paidLesson 9's gridAt scaleWhy
field evaluations5,529,600, 83% idleNeRF: 256 queries per ray, 4096 rays per step (Mildenhall et al., 2020)a ray samples the whole box
grid nodes stored1,600 in the planea 3-D grid: N3, 1,073,741,824 at N = 1024a grid is defined everywhere
arithmetic per evaluation4 table readsNeRF: 591,488 multiply–addsone network serves all of space

The last figure counts the layers of NeRF's network: with the paper's position and direction encodings of 60 and 24 numbers (its code also feeds the raw coordinates, 63 and 27), 60·256 + 3·2562 + (256 + 60)·256 + 3·2562 + 256 + 2562 + (256 + 24)·128 + 128·3 multiply–adds. A frame of 640,000 rays at 256 queries is 163,840,000 queries and 96.9 trillion multiply–adds; the paper reports about 30 seconds on a V100. Three bills, one cause: the representation is defined, and so paid for, everywhere. What would make cost follow the scene? Count what the scene occupies.

2 · What a scene occupies

Call a cell of lesson 9's grid touched when the true surface passes within half a cell diagonal of its centre. On its 39 × 39 cells the statue, the crate and the ball touch 112 of 1,521: 7.4%. Quarter the cell size and the count grows 4×, not 16×: at 156 × 156 cells, 450 of 24,336, 1.8%. The touched fraction times N stays at 2.9 from N = 20 to N = 156. A curve is about N cells long on a grid that has N2.

In space a surface is a sheet and the grid has N3 cells. A sphere of diameter half the box has area πN2/4 cells2, and with the same rule its shell is √3 cells thick: 4π(N/4)2·√3 = 1.36 N2 cells. Counted cell by cell, the touched fraction times N rises from 1.27 at N = 32 to 1.36 at N = 256, approaching 1.360. At N = 128 that is 1.05% of the cells; at N = 1024 it is one cell in 753, so 99.87% of a dense grid at that resolution stores air.

So capacity should follow the surface (§3 to §5).

Road not taken · a bigger network, or a tree
A wider network fits more detail per query and costs more per query: the multiply–adds above grow with the square of the width, and nothing in them knows that four samples in five are in air. An octree stores only occupied cells, the right count, at the price of a pointer walk per sample and a structure rebuilt as the scene forms; §3 and §4 reach the same counts with a bit mask and one expression.

3 · Skip the empty: an occupancy mask

Keep lesson 9's field and add one bit per cell: occupied if the field's density at the cell's centre or corners exceeds τ = 0.1 per metre. A ray evaluates only the samples that fall in occupied cells. (Instant-NGP keeps one for NeRF: 1283 one-bit cells, refreshed every 16 iterations, 256 KiB (Müller et al., 2022).) The mask is read from the field it speeds up, and that field starts empty, σ = 0.0067 per metre everywhere. A mask taken at step 16, when no node has reached 0.1 per metre (the largest is 0.078), has 0 cells: every ray skips everything, nothing changes, and a mask that is not refreshed holds the field at the empty scene's 4.65 dB for ever. So the run begins with every cell occupied, pays full price for 32 steps, and refreshes the mask every 16 steps from then on.

Run (12 cameras, 150 steps)Field evaluationsPer step, after warm-upCells occupiedData / held-out (dB)
lesson 9, no mask5,529,60036,8641,52124.4 / 20.4
learned mask (warm-up 32, refresh 16)2,446,444 (44.2%)10,736 (29.1%)394 (25.9%)24.7 / 20.5

The run needs 44% of the evaluations and scores the same. The mask is conservative (it keeps 25.9% of the cells where the surface touches 7.4%), yet what it drops was never needed: applied to lesson 9's finished field at render time it moves the scores by +0.4 and +0.3 dB, because the haze it removes was error. The field still stores 6,400 numbers, and an occupied sample still costs a full evaluation.

4 · Hash the sparse

To stop storing empty cells, give a node (i, j) no place of its own: a hash sends it to one of T slots (T is the table size here, not a transmittance), h(i, j) = (i·π1 xor j·π2) mod T, with π1 = 1 and π2 = 2,654,435,761 (Instant-NGP's constants, Müller et al., 2022). The table needs no keys, pointers or list of what is occupied, and its size is a budget. The price: two nodes may land in one slot, a collision, and then share a value and a gradient.

How often? With M occupied nodes in T slots, a node shares its slot with another occupied one with probability 1 − (1 − 1/T)M−1. On a 64 × 64-cell grid the surface touches 181 cells and 350 nodes; at T = 256 the formula gives 74.5%, the real hash 74.0%. The empty nodes matter more. Lesson 9's grid has 212 occupied nodes and 1,388 empty ones; at T = 512 every occupied node shares its slot with at least one empty node (2.7 per slot on average), and empty nodes are trained too: each ray through air pushes its slots toward zero density while the occupied node pushes up.

Representation (12 cameras, 150 steps)NumbersData / held-out (dB)
lesson 9's grid6,40024.4 / 20.4
the same grid, trained and rendered only near the true surface (an idealised mask, 382 cells)6,40022.7 / 20.0
one hashed level, T = 512, no mask2,04812.8 / 11.1
the same, with that mask2,04820.1 / 18.4
three hashed levels (16, 32, 64 cells; T = 256 each; 2 features per slot), a decoder, no mask1,71626.6 / 21.0

One hashed level without a mask collapses: the occupied nodes lose the tug-of-war to the empty ones in their slots. A mask that keeps the empty nodes out of the table repairs it, and the table then scores 2.6 dB below the unhashed masked grid, with 21.7% of its occupied nodes still sharing a slot at T = 512.

Instant-NGP's remedy is several levels. Level l has its own resolution Nl and a table of T slots (214 to 224), each holding F = 2 features; a level whose nodes fit in T is indexed directly, finer ones are hashed without collision handling; a point interpolates the features at its cell's corners on each of the 16 levels, concatenates them and hands the vector to a small network. A collision then costs little, because two occupied nodes that share a slot at the finest level are confused only if they share one at the other levels too. Count it on the 64 × 64 grid, where the 350 occupied nodes form 61,075 pairs: 177 collide at the finest level (random slots predict 239), 27 of those collide again at the parent level, and 0 of the 27 at the grandparent: no pair of occupied nodes collides at all three levels. The levels are not independent hashes (independent ones would leave 0.7 of the 177 at the second level, not 27), and here each level halves the cells of the one below. Three levels with a decoder of 16 hidden units hold 1,716 numbers, 27% of lesson 9's 6,400, and fit the data at 26.6 dB and the held-out cameras at 21.0; over three initialisations the fit ranges from 23.7 to 26.7 dB, never near the one-level collapse.

The decoder is where the arithmetic goes. Instant-NGP's NeRF has a density network with one hidden layer and a colour network with two, all 64 wide (Müller et al., 2022): 32 input features (16 levels × 2) enter the density network, 32·64 + 64·16; its 16 outputs and 16 direction coefficients enter the colour network, 32·64 + 64·64 + 64·3; together 9,408 multiply–adds per query, 62.9 times fewer than NeRF's, plus 256 table reads (16 levels × 8 corners × 2 features). Memory is capped by the budget: one dense level of 2,048 cells a side would hold 2,0493 nodes, 513 times the 224 slots a level gets. The paper reports NeRF-competitive quality after 15 s on one RTX 3090, for synthetic scenes and a fully fused implementation, not as a general guarantee.

5 · Primitives that carry their own extent: Gaussians

Everything so far keeps lesson 9's way of working: a ray asks the representation for its value at sample points (a gather), and the cost is samples times cost per sample, each made smaller by a mask or a hash. The alternative is to scatter: let each piece of the scene say which pixels it touches, so that the work is the number of (piece, pixel) pairs and empty space costs nothing at all. What must a piece be? By lesson 8 a pixel must be a smooth function of the scene, so a piece needs a soft edge and an extent (a point has none and its pixel is a cliff again); and it must project onto the image in closed form, so that its footprint costs one exponential, not a ray integral.

PieceSoft extentClosed-form footprintWhat goes wrong
voxelno: a box with an edgeyesa grid again: cells to allocate, an occupied set to know (§3, §4)
point or surfel, no opacitynoyesholes between points, a cliff at each (lesson 8)
mesh trianglenoyestopology fixed in advance (lesson 6); occlusion a cliff
Gaussianyes: smooth, negligible beyond about 5 standard deviations (sd)yes: a linear map of a Gaussian is a Gaussianits opacity is a fitted number, not a fog (below)

Only the Gaussian passes both tests. It has a mean μ, a covariance Σ, an opacity o and a colour c; its density is G(x) = exp(−½ (x − μ)TΣ−1(x − μ)). The covariance is stored as Σ = R S STRT (scales S, rotation R), which stays positive semi-definite whatever gradient steps do to the parameters (the choice of 3D Gaussian splatting, 3DGS (Kerbl et al., 2023)). In Flatland that is 9 numbers per Gaussian (position 2, log-scales 2, angle 1, opacity logit 1, colour 3).

Project. Put a camera at the point e, with unit vectors right and forward as the rows of Q. A point x has camera coordinates (xc, zc) = Q(x − e) and lands on the pixel u = W/2 + f xc/zc. Near the Gaussian's mean the map is close to linear, with Jacobian J = (f/zc, −f xc/zc2), so to first order the Gaussian becomes one on the image line, with mean m = W/2 + f xc/zc and variance

v = J Q Σ QT JT + 0.3    (px2; the 0.3 stops a far Gaussian from shrinking below half a pixel)

3DGS writes this step as Σ′ = J W Σ WT JT, with W the viewing transformation (not the image width) and the third row and column dropped (Kerbl et al., 2023, after Zwicker et al., 2001). Test it against the exact geometry. A Gaussian of scale 0.2 m at 5 m, with f = 56 px, has v = (56·0.2/5)2 = 5.02 px2, an sd of 2.24 px. Integrating its density along every pixel's own ray gives 5.07 (0.97% more: the Jacobian is exact only at the mean) and a profile equal to a Gaussian within 0.2% of its peak; two more placements, one off-axis and one anisotropic and rotated, agree within 1.5%. The footprint is 4.48 px wide in sd at 2.5 m and 1.12 px at 10 m: it shrinks as 1/z, and so does the work.

Composite. Sort the Gaussians by depth. Pixel u sees Gaussian k with opacity αk = ok exp(−(u − mk)2/(2vk)) (dropped beyond 5 sd), the transmittance is Tk = ∏j<k(1 − αj), and C = Σk Tkαkck + Tendcbg (Tend is what is left after the last Gaussian): lesson 8's composite with the opacities handed over directly (σδ = −ln(1 − α); the engine's composite reproduces lesson 8's to 10−4). The work is the number of pairs with α above 10−5; a pixel with no Gaussian over it costs nothing. A caution: o is not a fog's thickness. A fog of optical thickness τ at the Gaussian's centre gives 1 − exp(−τG) where its profile is G, which is (1 − e−τ)G only for small τ: at G = e−1/2 the two differ by 1.9% for τ = 0.1 and 31% for τ = 3. So o is a number to fit.

Differentiate. Raising αk adds the splat's colour and hides what is behind it a little more: lesson 8's bracket with α for σ, the light that arrives times what the splat adds minus the colour behind it.

∂C/∂αk = Tk (ck − Bk),   Bk = (Σi>k Tiαici + Tendcbg) / Tk+1

Then α = oG gives ∂α/∂o = G, ∂α/∂m = α(u − m)/v and ∂α/∂v = α(u − m)2/(2v2), and the engine carries these back through the five-to-two map (x, z, two log-scales, angle) → (m, v) by central differences. The whole gradient agrees with central differences of a separately written loss to 8.5 digits of its largest entry.

Grow. Here the Gaussians start at points the cameras saw (first hits of evenly chosen pixels, 8 cm of noise, the pixel's colour; a stand-in for lesson 5's reconstruction), 0.3 m wide and half opaque. Too few cannot draw the detail, and gradient descent cannot create more. 3DGS adds them where the picture is wrong: every 100 iterations, Gaussians whose mean view-space position gradient exceeds 0.0002 are cloned (small) or split in two with scales divided by 1.6 (large), and nearly transparent ones are removed (Kerbl et al., 2023). The engine's version: the 40% with the largest mean position gradient are densified, a clone if the largest scale is at most 0.35 m and a split otherwise, and opacities under 0.02 are pruned.

Fit Gaussians to the statue
Left: the scene from above, the true outlines and each Gaussian's 2σ ellipse (amber outline: added by the last densify); data cameras blue, one held-out camera amber. Right: that camera's photograph, the render, their difference, and the exam score against steps (green ticks: densify). Bottom: work for one camera, lesson 9's field samples against splat–pixel pairs. The sliders refit from scratch (400 Adam steps); the buttons continue the fit.
Gaussians
—
numbers stored
—
pairs per camera
—
steps
—
data cameras
—
24 held-out cameras
—
this held-out camera
—
fit minus exam
—
Show the core JS
var j0 = cam.f / zc, j1 = -cam.f * xc / (zc * zc);
out.v = j0 * j0 * b11 + 2 * j0 * j1 * b12 + j1 * j1 * b22 + dil;
out.m = cam.W / 2 + cam.f * xc / zc;
...
for (var j = 0; j < n; j++) {
  var k = ord[j], d = u - pr.m[k], v = pr.v[k];
  var g = Math.exp(-d * d / (2 * v)), a = pr.o[k] * g, clamped = 0;
  var w = T * a;
  C0 += w * pr.col[3 * k]; ...
  T *= (1 - a);
}
...
var dLda = dL0 * (Tj * c0 - b0 / (1 - a)) + dL1 * (Tj * c1 - b1 / (1 - a)) + dL2 * (Tj * c2 - b2 / (1 - a));
dO[kk] += dLda * g;
dM[kk] += dLdg * g * d / vv;
dV[kk] += dLdg * g * d * d / (2 * vv * vv);
b0 += w * c0; ...

What to try. The page opens with 40 Gaussians, twelve data cameras and 400 steps done. The ellipses lie along the outlines of the statue, the crate and the ball; the air and the interiors hold none. They hold 360 numbers where lesson 9's grid held 6,400, do 459 pair evaluations per camera where its field took 3,072, and score 24.7 dB on the data cameras and 22.1 on the held-out ones (the grid: 24.4 and 20.4). Slide Gaussians at the start from 8 to 96 to fill the table of §6. Set it to 24 (22.8 and 20.8 dB) and press densify three times: the count becomes 33, 46 and 64, the new Gaussians show in amber, and the scores climb to 25.9 and 22.4 dB at 467 pairs. Set Gaussians at the start back to 40 and data cameras to 3: 29.5 dB on the data and 14.2 on the held-out cameras. Camera 0 is 6.8° from a data camera and scores 18.6 dB, camera 4 53° and 15.1 dB. Press train 100 steps six times: the data cameras reach 33.7 dB and the held-out cameras stay at 14.1.

6 · Fit them, and price them

The widget fits by Adam on lesson 9's loss, 400 steps from the starting points. Budget first (9 numbers per Gaussian):

Gaussians at the start81624406496
data / held-out (dB)17.3 / 17.021.3 / 19.822.8 / 20.824.7 / 22.125.2 / 21.625.8 / 22.2
pairs per camera2482693454597041,011

The fit keeps improving; the exam stops at about 40 Gaussians while the work doubles. Against lesson 9's field the 40-Gaussian fit stores 17.8 times fewer numbers, does 6.7 times fewer evaluations per camera than the field's 3,072 samples (1.9 times fewer than the masked field's 895), and scores higher on the exam. A pair costs an exponential and a few multiplies, a field sample four reads and a softplus: the units are comparable, not equal.

Growth, six initialisations each, 700 steps in all: fit 400 steps from 24 Gaussians, then three rounds of densify and 100 steps.

RuleGaussians at the endData / held-out (dB)Pairs per camera
no growth2423.4 / 21.1320
three rounds, chosen by position gradient6426.0 / 22.4463
three rounds, chosen at random6426.4 / 22.4481
64 from the start6426.2 / 22.0648

The grown set has a loss 1.75 times lower than the set that stays at 24, as any extra capacity would give. The comparison at equal size is with 64 from the start: the loss differs by 2% and the grown set ends with 29% fewer pairs. Choosing where to add by the position gradient is no better than choosing at random: at the same 64 the loss is 2% lower for random, and lower in 4 of the six initialisations. The toy cannot say whether the rule matters at the scale of a real scene.

A 3-D Gaussian holds 3 (position) + 3 (scales) + 4 (a quaternion) + 1 (opacity) + 48 (colour: spherical harmonics of degree 3, 16 coefficients per channel) = 59 numbers, 236 bytes in 32 bits; the paper's scenes hold one to five million, and its 734 MB of stored model would hold 3.1 million at that size. Its rasteriser splits the screen into 16 × 16 tiles (256 pixels), instances each Gaussian in every tile it overlaps under a key of tile and depth, sorts everything once on the GPU, and lets one thread block per tile blend front to back, stopping when the opacity saturates (Kerbl et al., 2023): the widget's pairs, at scale, with no ray cast.

On the Mip-NeRF 360 scenes (average)TrainingRenderStored modelPSNR (dB)
Mip-NeRF 36048 h0.06 FPS8.6 MB27.69†
Instant-NGP, base5 min 37 s11.7 FPS13 MB25.30
3D Gaussian splatting, 30K steps41 min 33 s134 FPS734 MB27.21

Figures from Table 1 of the 3DGS paper (Kerbl et al., 2023): an A6000 throughout, except that Mip-NeRF 360's 48 h counts a four-GPU node's 12 hours and † marks quality quoted from its own paper. Against Mip-NeRF 360 the splats train 69 times faster, render 2,233 times faster and store 85 times more; against Instant-NGP they train 7.4 times slower, render 11.5 times faster and score 1.9 dB higher. The paper claims real-time rendering, at least 30 fps at 1080p; its 134 FPS above is at the benchmark's image sizes, about 1 to 1.6 thousand pixels wide.

7 · What speed did not buy

Masked grids, hashed tables and Gaussians are now cheap to fit and to draw. Take photographs away from them. Gaussians, 400 steps from 40, the same 24 held-out cameras:

Data cameras12631
fit to the data (dB)24.727.929.536.1
held-out cameras (dB)22.118.714.29.9

The fit improves as the views disappear, because fewer photographs are easier to reproduce, and the exam falls. At three views the other representations do the same: lesson 9's grid scores 32.9 dB on its photographs and 9.2 on the exam, the three-level hashed field 31.6 and 9.5, and copying the nearest photograph scores 9.7. More training does not repair it: 600 more steps take the Gaussians from 29.5 to 33.7 dB on the data and leave the exam at 14.1, and four times lesson 9's steps take the grid to 43.2 dB and 9.0. The error sits where no camera looked: the six held-out cameras nearest a data camera (within 8.3°) average 17.8 dB, the six farthest (at least 51.7° from every data camera) 11.5.

Counting parameters does not settle it. Forty Gaussians are 360 numbers against 576 photographed values, yet the exam fails. They start on points the three cameras saw, so the unseen sides hold none; a Gaussian outside every data camera's view would get exactly zero gradient and stay where it started, and one hidden behind others gets almost none. Nothing in three photographs says what the other cameras will see, and whatever part of a representation no photograph can move is arbitrary. Speed made fitting cheap and left that part where it was.

What this lesson did not do
The widget is two-dimensional and diffuse: a Gaussian has one colour, where 3DGS stores 48 spherical-harmonic numbers so that colour can change with direction. The tiled radix-sort rasteriser is described, not run; the loss is squared error, where 3DGS mixes L1 with D-SSIM (λ = 0.2); the starting points are the generator's first hits, not lesson 5's reconstruction. Aliasing under zoom is Mip-Splatting's problem (Yu et al., 2024), surface geometry rather than a cloud is 2D Gaussian splatting's (Huang et al., 2024). The hashed and masked fields are the plane's analogue of Instant-NGP, not its CUDA. The missing views are lesson 11's.

Common mistakes / failure modes

"a hash table is what makes it fast"
The hash caps the memory. Much of the speed is the tiny decoder (63 times fewer multiply–adds per query) and the skipped empty space (§3, §4).
"hash collisions corrupt the scene"
It depends on who collides: one level without a mask collapses to 12.8 dB, three levels with a decoder reach 26.6 (§4).
"compute the occupancy mask once, at the start"
At step 16 it has 0 cells, and a frozen empty mask holds the run at 4.65 dB (§3).
"Gaussian splatting is NeRF with ellipsoids"
It scatters pieces onto pixels where NeRF gathers along rays, and its opacity is a fitted number: 31% off a fog's for a thick one (§5).

Checkpoint exercise

Try it
(a) A Gaussian of scale 0.2 m is seen by a camera with f = 56 px at 5 m. Compute its image variance v (leave out the 0.3 px2), its sd, and the width of the 6-sd span it covers; then say what the sd is at 10 m. (b) A hash table has T = 1024 slots and holds 300 occupied nodes. What share of them collide with another occupied node? Answer: (a) v = (56·0.2/5)2 = 5.02 px2, sd 2.24 px, a span of 13.4 px; at 10 m the variance is 1.25 px2, a quarter, so the sd halves. (b) 1 − (1 − 1/1024)299 = 25.3%.

Where this points next

Cost now follows the scene: a mask cuts lesson 9's evaluations to 44%, a hash with a decoder keeps the quality in 27% of the numbers and 63 times fewer multiply–adds per query than NeRF, and 360 Gaussians reach 22.1 dB on cameras they never saw. From three photographs the grid scores 9.2 dB on the exam, the hashed field 9.5 and the Gaussians 14.2; longer training raises the fit and leaves the exam alone. A faster representation of the visible half says nothing about the half no photograph shows. Where can the missing information come from, if not from more photographs?

Takeaway
Lesson 9's field pays for the box: 83% of its samples are idle, its numbers fill empty cells, and NeRF spends 591,488 multiply–adds on each query. A surface touches about Nd−1 of Nd cells, so cost can follow the scene in three ways. An occupancy mask skips empty cells (44% of the evaluations here, same scores) but must start all-occupied. A hash table holds a fixed number of slots: collisions with empty nodes wreck one level, several levels read by a tiny decoder tolerate them, and the decoder is where Instant-NGP's arithmetic goes. Gaussians scatter onto pixels instead of gathering along rays: a projected Gaussian is a Gaussian, the composite is lesson 8's, the gradient is what a splat adds minus what it hides, and densification grows the set during the fit (in the toy it matched starting large, and where it adds mattered no more than chance). All three leave the photographs as insufficient as before: from three views the fit is high and the exam is not.

Interview prompts

Companion reads: Computer Graphics · 04 Rasterization (scatter: each primitive finds its pixels) and Computer Vision · 05 Multi-view, depth and SLAM (the one-screen NeRF and 3DGS overview).