Sampling the unseen: generative 3D
Lesson 12 ended with a network that, shown the front of an object that an oval and a trefoil both explain, answers with their average, a shape the world almost never makes. Squared error asks for the average, so the question changes from "what is the best guess?" to "what could it be?": a draw from the posterior, an object that fits the photographs and looks like what the world makes. A diffusion model makes draws from regression steps, and 3D systems differ in where it lives: single views, sets of views drawn together, or the scene. A draw is plausible, not true, and describes one still moment.
New idea: return a draw from the posterior over scenes instead of its mean: an object that fits the photographs and is typical of the world. A diffusion model makes draws from regression steps; systems differ in where the distribution lives.
Forces next: We can now reconstruct a scene, learn what scenes look like, and invent the parts nobody saw. But the scene holds still. In the world things move and so do we, and a model that explains the past cannot say what happens next or what would happen if we pushed. What does it take to represent a scene that changes, and then to predict it?
1 · The mean of two answers
Lesson 12's world has two kinds of object, ovals and trefoils, and a camera sees only the arc within 25° of its viewing direction: sixteen radii with 4.6 cm of noise. For a quarter of the objects the world makes (24.8%, 372 of 1,500) the front fits both kinds about equally well: the posterior probability of each kind is between 35% and 65%. Take front A, the first of lesson 12's widget: lesson 11's closed form, run once with each kind's prior, gives an oval mo and a trefoil mt that both fit it. In lesson 12's plane of scores, elongation and triangularity (an object is the unit circle plus those amounts of the world's two traits), they sit at (1.94, 0.27) and (0.08, 1.95). Distances below are the rms difference of two outlines' radii over the directions the camera did not face, in centimetres (mean radius 115 cm); the two explanations are 25.2 cm apart.
Why does a network trained by squared error answer with something between them? Take any answer θ̂ that depends only on the photographs y, and average its squared error over everything the photographs allow (E is over θ given y). Write θ − θ̂ = (θ − m) + (m − θ̂) with m = E[θ | y]. The cross term averages to zero, because θ − m averages to zero given y while m − θ̂ is fixed given y:
E‖θ − θ̂‖² = E‖θ − m‖² + ‖m − θ̂‖²
The first term is the same for every answer and the second is zero only at θ̂ = m. The best squared-error answer is the posterior mean, and a network trained long enough on enough examples approaches it. For two kinds, m = womo + wtmt, the two explanations averaged with the posterior probabilities of their kinds. On front A these are 50.1% and 49.9%, so the average lies halfway: 12.6 cm from the oval and 12.6 cm from the trefoil, at scores (1.01, 1.11).
The data cannot object to it. Its residual on the sixteen readings is 4.1 cm rms over the balanced fronts, inside the 4.6 cm of noise: a residual is linear in the outline, so the average's residual vector is the weighted average of the two explanations' and no longer than the weighted average of their lengths. The world can object. A real oval has triangularity near 0 and a real trefoil has elongation near 0, so h = min(|elongation|, |triangularity|), the object's hybrid score, says how much of the other kind it carries. Of 4,000 objects the world makes, 95.2% have h ≤ 0.5 and we call such an object valid; the median h of all 4,000 is 0.15. The average on front A has h = 1.01; over all the balanced fronts its median h is 0.64 and only 24% of the averages are valid. It agrees with the photographs and is not an object the world makes: the complaint that opened this lesson, with a number on it.
On a front that only one kind explains, the posterior has one hill, its mean lies on the hill, and the network is right: that is "where is the car". The network solves the problem it was given, and the problem is the question.
2 · Ask for a draw
The question that fits is "what could it be?", whose answer is the whole posterior p(θ | y), not one number computed from it. A draw from it has both properties we want. It agrees with the photographs, because the posterior is built from the likelihood, and it is a valid member of the world's kinds, because it is built from the prior.
Lesson 11's posterior had one hill. Here the prior is a mixture of two kinds, an oval and a trefoil, each a Gaussian N(μc, Λc) over the ten numbers θ with weight πc = ½. Let r be the radii minus 1, so r = Aθ + noise with lesson 11's design matrix A and noise variance σ². By total probability the posterior is a mixture of the per-kind posteriors, each weighted by the probability of its kind given the readings, and that is Bayes' rule on how well the kind predicted them (given a kind, r is Gaussian: the prior pushed through the linear measurement, plus noise):
p(θ | r) = Σc wc N(θ; mc, Sc), wc ∝ πc N(r; Aμc, AΛcAᵀ + σ²I)
Here mc and Sc are lesson 11's posterior mean and covariance with kind c's prior. To draw: pick a kind with probability wc, then return mc + Lξ, with ξ standard normal noise and LLT = Sc (a Cholesky factor). Four things could be returned for the same front. On the balanced fronts of §1 (twenty draws per front for the random answers):
| answer | error against the truth | valid (h ≤ 0.5) | two answers to one front differ by |
|---|---|---|---|
| the average (best squared-error answer) | 12.5 cm | 24% | 0, the same each time |
| the average plus noise (a Gaussian with the posterior's mean and covariance) | 17.4 cm | 63% | 17.1 cm |
| the explanation of the likelier kind | 15.1 cm | 99.7% | 0, the same each time |
| a draw from the posterior | 17.7 cm | 96% | 17.3 cm |
The average errs least and is invalid three times in four. Noise spread around it helps little: it puts its mass between the kinds, and more than a third of its answers are still not objects. The likelier kind's explanation is valid (99.7%) but is always the same object, and the wrong kind about as often as the posterior says. Only the draw is valid (96%, the world's own rate being 95.2%) and varied: two draws for one front differ by 17.3 cm.
A draw is not a better estimate. A draw θ′ from the posterior is independent of the truth given the photographs, so the same decomposition gives E‖θ − θ′‖² = E‖θ − m‖² + E‖θ′ − m‖² = 2 tr Cov(θ | y) (tr is the trace, the total variance): twice the average's squared error, so the rms error grows by a factor of √2, 17.7 against 12.5 cm, a ratio of 1.41. Validity and variety are bought with accuracy.
3 · A draw when the posterior cannot be written
The toy's posterior is a two-term formula. A real scene has millions of unknowns and no formula; what exists are examples, and regression, which returns means. A diffusion model turns regression into a sampler (Computer Vision 17 derives it in 2D). Mix an example x with noise, z = a x + s ε, where ε is standard Gaussian noise and a² + s² = 1: s = 0 is the example, s = 1 pure noise, and here a = cos(πt/2), s = sin(πt/2) for a noise level t from 0 to 1. Train a network ε̂(z, t) to predict ε by squared error. By §1 its minimiser is the conditional mean E[ε | z], and that is the score of the noised density pt: differentiating pt(z) = ∫ p(x) N(z; a x, s²I) dx under the integral gives ∇ log pt(z) = E[−ε/s | z], so
ε̂(z, t) = E[ε | z] = −s ∇ log pt(z)
In the toy no network is needed. Take x distributed as §2's posterior (θ in units of 0.05), a Gaussian mixture with weights wc. Then z is one too, with components N(a mc, Cc), Cc = a²Sc + s²I, and the minimiser is exact: ε̂ = s Σc rc(z) Cc−1(z − a mc), where rc(z) ∝ wc N(z; a mc, Cc) is the probability of kind c given z. To draw, start from pure noise and walk the level down in 60 steps from t = 0.99: the noise estimate gives an estimate of the clean sample, x̂ = (z − s ε̂)/a, and the state moves to the next, lower level by mixing the same two estimates in the new proportions, z′ = a′x̂ + s′ε̂.
Why does this not average, like §1? Every step is a regression, but only a small step is taken toward its answer. At the first step the clean-sample estimate is the average, and 0% of the paths have a valid one. The state keeps the noise estimate, so the starting noise leans each path a little to one side of the middle, the next estimate leans further, and the choice of kind is made as the noise falls: the estimate is a valid object on 55% of the paths after 12 of the 60 steps and on 91% after 24. Each estimate is still an average, over the clean samples that could have produced the state, and late in the walk those form one hill. Two thousand draws on front A are 50.3% ovals against a weight of 50.1%; on front C, where the oval is four times likelier, 80.2% against 79.9%; their spread is within 3% of the closed form's. The sampler only ever calls the denoiser; here it was written down from the mixture, a real system learns it from examples, and the posterior is never written.
4 · Lifting a 2D model by optimisation
A diffusion model of views can be trained on images alone, and DreamFusion (Poole et al., 2022) turns one into a 3D generator with no 3D data. Hold a scene θ (a NeRF), render a random view x = g(θ), noise it, and change θ until the frozen 2D model finds the view likely: the loss is the diffusion loss L = ½ w(t)‖ε̂(z) − ε‖², with a weight w(t) of the noise level and z = a x + s ε, differentiated with respect to θ. The chain rule gives
∂L/∂θ = a w(t) (ε̂ − ε)ᵀ · (∂ε̂/∂z) · (∂x/∂θ)
The middle factor is the Jacobian of the network, expensive and badly conditioned at low noise. DreamFusion drops it (and folds a into w): ∇θ LSDS = Et,ε[ w(t) (ε̂ − ε) ∂x/∂θ ], score distillation sampling, one forward pass of the frozen network per step. What does that gradient descend? Since ε = −s ∇ log qt(z) for the noised render qt = N(a x, s²I), we have ε̂ − ε = s(∇ log qt − ∇ log pt), a difference of two scores, and averaged over ε it is (s/a) ∇x KL(qt ‖ pt). So SDS descends a KL divergence from a point, the render of one scene. The entropy of qt does not depend on x, so what falls is −Eq[log pt]: the point is pulled to the highest part of pt, never spread over it. The optimiser is mode seeking. DreamFusion reports it: its samples tend to lack diversity, the results are often oversaturated and oversmoothed, and it needs a classifier-free guidance weight of 100 because the objective is mode seeking and oversmooths at small weights.
The toy can do all of it exactly. The scene is θ ∈ ℝ10; twelve cameras 30° apart each read nine radii across ±50°, a linear render like lesson 11's camera; the 2D prior of a view is the exact diffusion model of what that view reads, conditional on the photograph (the posterior's image under the view) and, for guidance, unconditional (the world's own). Start from a round blob and take 200 Adam steps at 0.03, each with two random views, a random noise level and noise, and w(t) = s² as in DreamFusion; with guidance ω the denoiser is ε̂u + ω(ε̂c − ε̂u). The identity above checks by quadrature in one dimension to six digits.
Then climb. On front A, where the kinds are equally likely, 0 of 200 climbs end as ovals, and the same at guidance 3, 10 and 100, although the posterior gives the oval half the weight. Two climbs differ by 2.1 cm, two draws by 19.2 cm. The cause is geometry, not probability: the climber goes up the hill nearest its start. The trefoil explanation is 11.8 cm from the round start and the oval 21.4 cm, and on all 372 balanced fronts the trefoil is nearer: of one climb from each of 60 of them, 59 end as the trefoil. Start instead from each of 120 real objects and every climb on front A ends on the kind of the explanation nearest its start. The random views and noise only jitter the gradient: at the endpoints its expected value is down to 6.7% of its starting size.
Front C, where the oval is four times likelier, shows what guidance does (200 climbs each):
| front C | end as ovals | endpoints valid |
|---|---|---|
| posterior draws | 79.9% | 95% |
| climb, ω = 1 | 50.5% | 62% |
| climb, ω = 3 | 100% | 100% |
| climb, ω = 10 | 100% | 100% |
| climb, ω = 100 | 100% | 61% |
Without guidance the endpoints are blurs: only 62% are valid, and the oval wins half the time, not the posterior's 80%. A little guidance sharpens them onto the likelier kind, every time. A lot of it overshoots: at 100 only 61% are valid. Over 120 fronts: 92.5% valid at ω = 3, 58.8% at ω = 100. These are the toy's versions of two things DreamFusion reports, oversmoothing at small guidance weights and oversaturated results, the price of drawing by climbing. Later work names the Janus problem, in which the most canonical view of an object (a face) appears in other views (Hong, Ahn and Kim, 2023; DreamFusion's ablation without view-dependent prompts gives a multi-faced dog), and ProlificDreamer (Wang et al., 2023) repairs the oversaturation, oversmoothing and low diversity with variational score distillation, usable at guidance 7.5.
5 · Making views agree, and drawing the scene itself
Zero-1-to-3 (Liu et al., 2023) fine-tunes Stable Diffusion into a model of what an object looks like from another camera, given one image, trained on renderings of Objaverse (over 800,000 models). Asking such a model for each view separately and reconstructing the views with lessons 8 to 10 looks like a way to a 3D scene; the toy shows how that fails. The distribution of one view is the posterior's marginal: a mixture with the same weights. Draw four views of front A independently and each is an oval with probability w = 50.1%, so all four agree on the kind with probability w⁴ + (1 − w)⁴ = 12.5%. Otherwise one view shows an oval's flank beside a trefoil's, and no object of either kind explains them all: a reconstruction either averages them, which is §1 again, or fails.
The cure is to draw the views together, as views of one scene. In the toy that is drawing θ and rendering it from every camera: the four views then agree 100% of the time. MVDream (Shi et al., 2023) inflates the self-attention layers of Stable Diffusion to attend across four views and replaces the 2D prior of score distillation, for more consistent and stable lifting (about 2 hours per asset on a V100); CAT3D (Gao et al., 2024) trains a multi-view latent diffusion model that generates many consistent novel views from any number of observed ones, then fits a NeRF to the observed and generated views with lessons 8 to 10, in as little as one minute.
The third place for the distribution is the scene itself, and the toy's reverse diffusion of §3 is already that. LRM (Hong et al., 2023) maps one image to a triplane NeRF in about 5 seconds with a 500-million-parameter transformer trained on about a million objects, but with image-reconstruction losses, and returns one scene per image: a regression, with §1's problem. TRELLIS (Xiang et al., 2024) draws: a structured latent of about 20,000 active voxels on a 64³ grid, each with a local latent, made by two rectified-flow transformers (voxels, then latents) in about 10 seconds, trained on about 500,000 assets. In 2025 Hunyuan3D 2.0 paired a flow-based shape transformer with a multi-view texture model, TRELLIS.2 scaled flow matching to 4 billion parameters on a sparse voxel structure, and SAM 3D reported a win rate of at least 5:1 in human preference tests; preference judges plausibility, which §6 separates from truth.
| the distribution lives on | what is drawn | 3D data needed | price | how it fails |
|---|---|---|---|---|
| single views | one scene, climbed by an optimiser (DreamFusion) | none | 15,000 iterations, about 1.5 h on a four-chip TPUv4 machine | one mode; oversmoothed or oversaturated; Janus |
| sets of views | views drawn jointly: a prior for the optimiser (MVDream) or views to reconstruct from (CAT3D) | multi-view renderings (MVDream: Objaverse) | about 2 h per asset (MVDream), as little as a minute (CAT3D) | a reconstruction inherits what the views get wrong |
| the scene | a latent of the scene (TRELLIS; LRM returns one answer) | 5·105 to 106 objects | seconds (LRM 5, TRELLIS 10) | what the data never contained cannot be drawn (lesson 11's warning) |
What to try. The page opens on front A with 24 posterior draws: amber ovals and violet trefoils, 42% of them ovals (the posterior says 50.1%), 100% valid, two of them 21.0 cm apart on average, and the kind changing 8 times in 23 consecutive pairs. Drag the draws slider up to 48: the bundle fills in around the oval and the trefoil and the spread stays near 20 cm (19.8 cm): more draws do not average out. Choose the average: one blue outline halfway between the two bundles, 12.6 cm from either explanation, outside both green bands (h = 1.01), with an error of 11.4 cm against 19.6 cm for the draws. Choose reverse diffusion: a similar split (58% ovals) and thin paths from the middle that bend onto an arm. Choose score climbing: every outline is a trefoil (0% ovals), the bundle shrinks to 2.3 cm, and raising the guidance does not bring the oval back. On front C draws are 75% ovals; climbing at ω = 1, 3 and 100 gives 71%, 100% and 54% valid. Choose views drawn independently: 17% of the 24 outlines have four views of one kind (closed form 12.5%), and the rest mix an oval's flank with a trefoil's. On front D, where only a trefoil fits, the average is a real object (h = 0.01, error 5.2 cm) and every answer agrees: regression is right when the posterior has one hill.
6 · What a draw is, and what it is not
A draw answers "what could it be?", not "what is it?". On the balanced fronts a draw has the true kind 51% of the time, a coin flip (the likelier kind's explanation has it 58%): the plausibility that makes a draw useful says nothing about whether it is the truth. A draw is also fresh each time you ask: request the back of front A in two consecutive frames and the kind changes in 50.0% of the pairs, 2w(1 − w) in closed form (49.7% measured over 8,000 draws). Within one still scene the cure is to draw once and render the draw from every camera, or to draw the views together (§5): sharing one draw is what makes the answers agree, and it works because nothing moves between the cameras.
Everything so far in this track has held still. In a world that moves, a back drawn at one frame describes the object at that frame. A new draw for every frame flickers, and a frozen draw does not move with the object: a perfect back, left where it was, is 8.3 cm wrong on front A after the statue turns 10° and 15.8 cm after 20°. What is missing is a representation that carries what was drawn forward and a model of how it changes: a scene with time in it.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
A draw gave the unseen part of a scene a valid answer where the average gave none (96% of draws against 24% of averages), for a factor √2 in error, and a diffusion model makes draws from regression in small steps. But a draw describes one still moment: it is the true kind 51% of the time, asking again changes the kind in 50% of consecutive pairs, and a draw left in place goes stale as the object turns. What does it take to represent a scene that changes, and then to predict it?
Interview prompts
- Why does a network trained by squared error return the posterior mean, and when is that the wrong answer? (§1 — E‖θ − θ̂‖² splits into the posterior variance plus ‖m − θ̂‖²; with two kinds the mean lies between them, valid in 24% of the balanced fronts.)
- A draw has twice the expected squared error of the mean. Why use it? (§2 — E‖θ − θ′‖² = 2 tr Cov, √2 in error; the draw is valid and varied, the mean is neither.)
- A diffusion model is trained by regression. Why does it not return the average? (§3 — it returns the average at every level but moves only a small step toward it; the state leans to one side as the noise falls.)
- Derive the score distillation gradient and say what it minimises. (§4 — drop the network's Jacobian from the diffusion-loss gradient; its expectation is (s/a)∇ KL(q ‖ p), a KL from a point, minimised at a mode.)
- Why does DreamFusion need guidance 100, and what does that cost? (§4 — mode seeking oversmooths at small weights; large ones oversaturate; in the toy 100% valid at ω = 3 and 61% at 100.)
- Why does drawing a new back for every frame of a video flicker? (§6 — draws are independent, so consecutive draws change kind with probability 2w(1 − w), 50.0% on a coin-flip front.)
Companion reads: Computer Vision · 17 Generative vision (diffusion and classifier-free guidance in 2D) and World Models · 05 One future is a lie (the same mean-versus-draw problem for futures, and mixture densities).