all_lessons/3D Vision/13 · Sampling the unseen: generative 3Dlesson 13 / 14

Sampling the unseen: generative 3D

Lesson 12 ended with a network that, shown the front of an object that an oval and a trefoil both explain, answers with their average, a shape the world almost never makes. Squared error asks for the average, so the question changes from "what is the best guess?" to "what could it be?": a draw from the posterior, an object that fits the photographs and looks like what the world makes. A diffusion model makes draws from regression steps, and 3D systems differ in where it lives: single views, sets of views drawn together, or the scene. A draw is plausible, not true, and describes one still moment.

The thesis, here
A network trained by squared error answers with the posterior mean, and where the unseen part has two kinds of answer the mean is neither. The fix is to change the question: ask for a draw from the posterior, an object that agrees with the photographs and is typical of the world. A draw can be made without writing the posterior down, by a denoiser that is regression asked only for small steps.
Linear position
Forced by: A trained network returns one answer per input, the average of everything compatible with what it saw. That is fine for "where is the car" and wrong for "what is behind the mug": the average of two plausible backs is a back that cannot exist. How do we produce a scene that agrees with the data and is also a valid sample of the kinds of scenes there are?
New idea: return a draw from the posterior over scenes instead of its mean: an object that fits the photographs and is typical of the world. A diffusion model makes draws from regression steps; systems differ in where the distribution lives.
Forces next: We can now reconstruct a scene, learn what scenes look like, and invent the parts nobody saw. But the scene holds still. In the world things move and so do we, and a model that explains the past cannot say what happens next or what would happen if we pushed. What does it take to represent a scene that changes, and then to predict it?
The plan
Six moves. (1) Squared error returns the mean; measure what the mean is worth. (2) Change the question to a draw, write the posterior for two kinds, and price the draw. (3) Draw when the posterior cannot be written: noise and small regression steps. (4) Lift a 2D model into 3D by optimisation, and see what it finds. (5) Make several views agree, or draw the scene itself. (6) Ask what a draw is, and what it is not.

1 · The mean of two answers

Lesson 12's world has two kinds of object, ovals and trefoils, and a camera sees only the arc within 25° of its viewing direction: sixteen radii with 4.6 cm of noise. For a quarter of the objects the world makes (24.8%, 372 of 1,500) the front fits both kinds about equally well: the posterior probability of each kind is between 35% and 65%. Take front A, the first of lesson 12's widget: lesson 11's closed form, run once with each kind's prior, gives an oval mo and a trefoil mt that both fit it. In lesson 12's plane of scores, elongation and triangularity (an object is the unit circle plus those amounts of the world's two traits), they sit at (1.94, 0.27) and (0.08, 1.95). Distances below are the rms difference of two outlines' radii over the directions the camera did not face, in centimetres (mean radius 115 cm); the two explanations are 25.2 cm apart.

Why does a network trained by squared error answer with something between them? Take any answer θ̂ that depends only on the photographs y, and average its squared error over everything the photographs allow (E is over θ given y). Write θ − θ̂ = (θ − m) + (m − θ̂) with m = E[θ | y]. The cross term averages to zero, because θ − m averages to zero given y while m − θ̂ is fixed given y:

E‖θ − θ̂‖² = E‖θ − m‖² + ‖m − θ̂‖²

The first term is the same for every answer and the second is zero only at θ̂ = m. The best squared-error answer is the posterior mean, and a network trained long enough on enough examples approaches it. For two kinds, m = womo + wtmt, the two explanations averaged with the posterior probabilities of their kinds. On front A these are 50.1% and 49.9%, so the average lies halfway: 12.6 cm from the oval and 12.6 cm from the trefoil, at scores (1.01, 1.11).

The data cannot object to it. Its residual on the sixteen readings is 4.1 cm rms over the balanced fronts, inside the 4.6 cm of noise: a residual is linear in the outline, so the average's residual vector is the weighted average of the two explanations' and no longer than the weighted average of their lengths. The world can object. A real oval has triangularity near 0 and a real trefoil has elongation near 0, so h = min(|elongation|, |triangularity|), the object's hybrid score, says how much of the other kind it carries. Of 4,000 objects the world makes, 95.2% have h ≤ 0.5 and we call such an object valid; the median h of all 4,000 is 0.15. The average on front A has h = 1.01; over all the balanced fronts its median h is 0.64 and only 24% of the averages are valid. It agrees with the photographs and is not an object the world makes: the complaint that opened this lesson, with a number on it.

On a front that only one kind explains, the posterior has one hill, its mean lies on the hill, and the network is right: that is "where is the car". The network solves the problem it was given, and the problem is the question.

2 · Ask for a draw

The question that fits is "what could it be?", whose answer is the whole posterior p(θ | y), not one number computed from it. A draw from it has both properties we want. It agrees with the photographs, because the posterior is built from the likelihood, and it is a valid member of the world's kinds, because it is built from the prior.

Lesson 11's posterior had one hill. Here the prior is a mixture of two kinds, an oval and a trefoil, each a Gaussian N(μc, Λc) over the ten numbers θ with weight πc = ½. Let r be the radii minus 1, so r = Aθ + noise with lesson 11's design matrix A and noise variance σ². By total probability the posterior is a mixture of the per-kind posteriors, each weighted by the probability of its kind given the readings, and that is Bayes' rule on how well the kind predicted them (given a kind, r is Gaussian: the prior pushed through the linear measurement, plus noise):

p(θ | r) = Σc wc N(θ; mc, Sc),   wc ∝ πc N(r; Aμc, AΛcAᵀ + σ²I)

Here mc and Sc are lesson 11's posterior mean and covariance with kind c's prior. To draw: pick a kind with probability wc, then return mc + Lξ, with ξ standard normal noise and LLT = Sc (a Cholesky factor). Four things could be returned for the same front. On the balanced fronts of §1 (twenty draws per front for the random answers):

answererror against the truthvalid (h ≤ 0.5)two answers to one front differ by
the average (best squared-error answer)12.5 cm24%0, the same each time
the average plus noise (a Gaussian with the posterior's mean and covariance)17.4 cm63%17.1 cm
the explanation of the likelier kind15.1 cm99.7%0, the same each time
a draw from the posterior17.7 cm96%17.3 cm

The average errs least and is invalid three times in four. Noise spread around it helps little: it puts its mass between the kinds, and more than a third of its answers are still not objects. The likelier kind's explanation is valid (99.7%) but is always the same object, and the wrong kind about as often as the posterior says. Only the draw is valid (96%, the world's own rate being 95.2%) and varied: two draws for one front differ by 17.3 cm.

A draw is not a better estimate. A draw θ′ from the posterior is independent of the truth given the photographs, so the same decomposition gives E‖θ − θ′‖² = E‖θ − m‖² + E‖θ′ − m‖² = 2 tr Cov(θ | y) (tr is the trace, the total variance): twice the average's squared error, so the rms error grows by a factor of √2, 17.7 against 12.5 cm, a ratio of 1.41. Validity and variety are bought with accuracy.

3 · A draw when the posterior cannot be written

The toy's posterior is a two-term formula. A real scene has millions of unknowns and no formula; what exists are examples, and regression, which returns means. A diffusion model turns regression into a sampler (Computer Vision 17 derives it in 2D). Mix an example x with noise, z = a x + s ε, where ε is standard Gaussian noise and a² + s² = 1: s = 0 is the example, s = 1 pure noise, and here a = cos(πt/2), s = sin(πt/2) for a noise level t from 0 to 1. Train a network ε̂(z, t) to predict ε by squared error. By §1 its minimiser is the conditional mean E[ε | z], and that is the score of the noised density pt: differentiating pt(z) = ∫ p(x) N(z; a x, s²I) dx under the integral gives ∇ log pt(z) = E[−ε/s | z], so

ε̂(z, t) = E[ε | z] = −s ∇ log pt(z)

In the toy no network is needed. Take x distributed as §2's posterior (θ in units of 0.05), a Gaussian mixture with weights wc. Then z is one too, with components N(a mc, Cc), Cc = a²Sc + s²I, and the minimiser is exact: ε̂ = s Σc rc(z) Cc−1(z − a mc), where rc(z) ∝ wc N(z; a mc, Cc) is the probability of kind c given z. To draw, start from pure noise and walk the level down in 60 steps from t = 0.99: the noise estimate gives an estimate of the clean sample, x̂ = (z − s ε̂)/a, and the state moves to the next, lower level by mixing the same two estimates in the new proportions, z′ = a′x̂ + s′ε̂.

Why does this not average, like §1? Every step is a regression, but only a small step is taken toward its answer. At the first step the clean-sample estimate is the average, and 0% of the paths have a valid one. The state keeps the noise estimate, so the starting noise leans each path a little to one side of the middle, the next estimate leans further, and the choice of kind is made as the noise falls: the estimate is a valid object on 55% of the paths after 12 of the 60 steps and on 91% after 24. Each estimate is still an average, over the clean samples that could have produced the state, and late in the walk those form one hill. Two thousand draws on front A are 50.3% ovals against a weight of 50.1%; on front C, where the oval is four times likelier, 80.2% against 79.9%; their spread is within 3% of the closed form's. The sampler only ever calls the denoiser; here it was written down from the mixture, a real system learns it from examples, and the posterior is never written.

4 · Lifting a 2D model by optimisation

A diffusion model of views can be trained on images alone, and DreamFusion (Poole et al., 2022) turns one into a 3D generator with no 3D data. Hold a scene θ (a NeRF), render a random view x = g(θ), noise it, and change θ until the frozen 2D model finds the view likely: the loss is the diffusion loss L = ½ w(t)‖ε̂(z) − ε‖², with a weight w(t) of the noise level and z = a x + s ε, differentiated with respect to θ. The chain rule gives

∂L/∂θ = a w(t) (ε̂ − ε)ᵀ · (∂ε̂/∂z) · (∂x/∂θ)

The middle factor is the Jacobian of the network, expensive and badly conditioned at low noise. DreamFusion drops it (and folds a into w): ∇θ LSDS = Et,ε[ w(t) (ε̂ − ε) ∂x/∂θ ], score distillation sampling, one forward pass of the frozen network per step. What does that gradient descend? Since ε = −s ∇ log qt(z) for the noised render qt = N(a x, s²I), we have ε̂ − ε = s(∇ log qt − ∇ log pt), a difference of two scores, and averaged over ε it is (s/a) ∇x KL(qt ‖ pt). So SDS descends a KL divergence from a point, the render of one scene. The entropy of qt does not depend on x, so what falls is −Eq[log pt]: the point is pulled to the highest part of pt, never spread over it. The optimiser is mode seeking. DreamFusion reports it: its samples tend to lack diversity, the results are often oversaturated and oversmoothed, and it needs a classifier-free guidance weight of 100 because the objective is mode seeking and oversmooths at small weights.

The toy can do all of it exactly. The scene is θ ∈ ℝ10; twelve cameras 30° apart each read nine radii across ±50°, a linear render like lesson 11's camera; the 2D prior of a view is the exact diffusion model of what that view reads, conditional on the photograph (the posterior's image under the view) and, for guidance, unconditional (the world's own). Start from a round blob and take 200 Adam steps at 0.03, each with two random views, a random noise level and noise, and w(t) = s² as in DreamFusion; with guidance ω the denoiser is ε̂u + ω(ε̂c − ε̂u). The identity above checks by quadrature in one dimension to six digits.

Then climb. On front A, where the kinds are equally likely, 0 of 200 climbs end as ovals, and the same at guidance 3, 10 and 100, although the posterior gives the oval half the weight. Two climbs differ by 2.1 cm, two draws by 19.2 cm. The cause is geometry, not probability: the climber goes up the hill nearest its start. The trefoil explanation is 11.8 cm from the round start and the oval 21.4 cm, and on all 372 balanced fronts the trefoil is nearer: of one climb from each of 60 of them, 59 end as the trefoil. Start instead from each of 120 real objects and every climb on front A ends on the kind of the explanation nearest its start. The random views and noise only jitter the gradient: at the endpoints its expected value is down to 6.7% of its starting size.

Front C, where the oval is four times likelier, shows what guidance does (200 climbs each):

front Cend as ovalsendpoints valid
posterior draws79.9%95%
climb, ω = 150.5%62%
climb, ω = 3100%100%
climb, ω = 10100%100%
climb, ω = 100100%61%

Without guidance the endpoints are blurs: only 62% are valid, and the oval wins half the time, not the posterior's 80%. A little guidance sharpens them onto the likelier kind, every time. A lot of it overshoots: at 100 only 61% are valid. Over 120 fronts: 92.5% valid at ω = 3, 58.8% at ω = 100. These are the toy's versions of two things DreamFusion reports, oversmoothing at small guidance weights and oversaturated results, the price of drawing by climbing. Later work names the Janus problem, in which the most canonical view of an object (a face) appears in other views (Hong, Ahn and Kim, 2023; DreamFusion's ablation without view-dependent prompts gives a multi-faced dog), and ProlificDreamer (Wang et al., 2023) repairs the oversaturation, oversmoothing and low diversity with variational score distillation, usable at guidance 7.5.

5 · Making views agree, and drawing the scene itself

Zero-1-to-3 (Liu et al., 2023) fine-tunes Stable Diffusion into a model of what an object looks like from another camera, given one image, trained on renderings of Objaverse (over 800,000 models). Asking such a model for each view separately and reconstructing the views with lessons 8 to 10 looks like a way to a 3D scene; the toy shows how that fails. The distribution of one view is the posterior's marginal: a mixture with the same weights. Draw four views of front A independently and each is an oval with probability w = 50.1%, so all four agree on the kind with probability w⁴ + (1 − w)⁴ = 12.5%. Otherwise one view shows an oval's flank beside a trefoil's, and no object of either kind explains them all: a reconstruction either averages them, which is §1 again, or fails.

The cure is to draw the views together, as views of one scene. In the toy that is drawing θ and rendering it from every camera: the four views then agree 100% of the time. MVDream (Shi et al., 2023) inflates the self-attention layers of Stable Diffusion to attend across four views and replaces the 2D prior of score distillation, for more consistent and stable lifting (about 2 hours per asset on a V100); CAT3D (Gao et al., 2024) trains a multi-view latent diffusion model that generates many consistent novel views from any number of observed ones, then fits a NeRF to the observed and generated views with lessons 8 to 10, in as little as one minute.

The third place for the distribution is the scene itself, and the toy's reverse diffusion of §3 is already that. LRM (Hong et al., 2023) maps one image to a triplane NeRF in about 5 seconds with a 500-million-parameter transformer trained on about a million objects, but with image-reconstruction losses, and returns one scene per image: a regression, with §1's problem. TRELLIS (Xiang et al., 2024) draws: a structured latent of about 20,000 active voxels on a 64³ grid, each with a local latent, made by two rectified-flow transformers (voxels, then latents) in about 10 seconds, trained on about 500,000 assets. In 2025 Hunyuan3D 2.0 paired a flow-based shape transformer with a multi-view texture model, TRELLIS.2 scaled flow matching to 4 billion parameters on a sparse voxel structure, and SAM 3D reported a win rate of at least 5:1 in human preference tests; preference judges plausibility, which §6 separates from truth.

the distribution lives onwhat is drawn3D data neededpricehow it fails
single viewsone scene, climbed by an optimiser (DreamFusion)none15,000 iterations, about 1.5 h on a four-chip TPUv4 machineone mode; oversmoothed or oversaturated; Janus
sets of viewsviews drawn jointly: a prior for the optimiser (MVDream) or views to reconstruct from (CAT3D)multi-view renderings (MVDream: Objaverse)about 2 h per asset (MVDream), as little as a minute (CAT3D)a reconstruction inherits what the views get wrong
the scenea latent of the scene (TRELLIS; LRM returns one answer)5·105 to 106 objectsseconds (LRM 5, TRELLIS 10)what the data never contained cannot be drawn (lesson 11's warning)
One front, five ways to answer
Left: the object from above. The camera at 60° read the green arc (dots: its sixteen radii); black dashed is the truth, blue the average, amber an oval draw, violet a trefoil draw. Right: the plane of scores (elongation across, triangularity up): grey dots are 300 real objects, green bands are where h ≤ 0.5, circles the two explanations, the cross the average, the black dot the truth, thin lines the paths of the clean-sample estimate (reverse diffusion) or of the climber. views stitches each outline from four independently drawn views.
ovals among the answers
—
hybrid score h, median
—
valid answers
—
error against the truth
—
two answers differ by
—
consecutive answers change kind
—
four views agree on the kind
—
h of the average
—
Show the core JS
for (k = 0; k < L13.TDDIM; k++) {
  epsHat(M.P[k], D, z, M.a[k], M.s[k], eps);
  for (i = 0; i < D; i++) x0[i] = (z[i] - M.s[k] * eps[i]) / M.a[k];
  for (i = 0; i < D; i++) z[i] = M.a[k + 1] * x0[i] + M.s[k + 1] * eps[i];
}
...
epsHat(cond[vi][ti], NV, z, a, s, ec);
if (om !== 1) { epsHat(vs[vi].unc[ti], NV, z, a, s, eu); for (i = 0; i < NV; i++) ec[i] = eu[i] + om * (ec[i] - eu[i]); }
for (j = 0; j < D; j++) { var gs = 0; for (i = 0; i < NV; i++) gs += (ec[i] - eps[i]) * A[i * D + j]; g[j] += s * s * gs / V; }

What to try. The page opens on front A with 24 posterior draws: amber ovals and violet trefoils, 42% of them ovals (the posterior says 50.1%), 100% valid, two of them 21.0 cm apart on average, and the kind changing 8 times in 23 consecutive pairs. Drag the draws slider up to 48: the bundle fills in around the oval and the trefoil and the spread stays near 20 cm (19.8 cm): more draws do not average out. Choose the average: one blue outline halfway between the two bundles, 12.6 cm from either explanation, outside both green bands (h = 1.01), with an error of 11.4 cm against 19.6 cm for the draws. Choose reverse diffusion: a similar split (58% ovals) and thin paths from the middle that bend onto an arm. Choose score climbing: every outline is a trefoil (0% ovals), the bundle shrinks to 2.3 cm, and raising the guidance does not bring the oval back. On front C draws are 75% ovals; climbing at ω = 1, 3 and 100 gives 71%, 100% and 54% valid. Choose views drawn independently: 17% of the 24 outlines have four views of one kind (closed form 12.5%), and the rest mix an oval's flank with a trefoil's. On front D, where only a trefoil fits, the average is a real object (h = 0.01, error 5.2 cm) and every answer agrees: regression is right when the posterior has one hill.

Road not taken · let the network predict a distribution
The tempting repair keeps regression and widens what it outputs. Predict a mean and a covariance: that is the "average plus noise" row of §2, and only 63% of its answers are valid. Train several networks and ask them all: each minimises squared error, so each approaches the same mean (§1), and the members differ by optimisation noise, not by kind. Predict a weight, a mean and a covariance per kind (a mixture density network): trained by likelihood it can match §2's posterior, and it is the right tool when the kinds are few and known (World Models 05 uses it for futures). The back of a real mug has no list of kinds, so what has to be learned is a way of drawing, not a list of answers.

6 · What a draw is, and what it is not

A draw answers "what could it be?", not "what is it?". On the balanced fronts a draw has the true kind 51% of the time, a coin flip (the likelier kind's explanation has it 58%): the plausibility that makes a draw useful says nothing about whether it is the truth. A draw is also fresh each time you ask: request the back of front A in two consecutive frames and the kind changes in 50.0% of the pairs, 2w(1 − w) in closed form (49.7% measured over 8,000 draws). Within one still scene the cure is to draw once and render the draw from every camera, or to draw the views together (§5): sharing one draw is what makes the answers agree, and it works because nothing moves between the cameras.

Everything so far in this track has held still. In a world that moves, a back drawn at one frame describes the object at that frame. A new draw for every frame flickers, and a frozen draw does not move with the object: a perfect back, left where it was, is 8.3 cm wrong on front A after the statue turns 10° and 15.8 cm after 20°. What is missing is a representation that carries what was drawn forward and a model of how it changes: a scene with time in it.

What this lesson did not do
It drew ten-number outlines with the exact denoiser in place of a trained network, so the errors of a learned score (finite data, lesson 11's out-of-family warning) were not measured. The per-view priors were marginals of one exact posterior; a real 2D model knows nothing of how views relate. It drew shape only: no appearance, no topology. "Valid" (h ≤ 0.5) detects a blend of the two kinds, not every way an object can be wrong: a featureless circle passes. The repairs of SDS were cited, not run, and every figure for a real system is as its authors report it. A draw describes one moment, and lesson 14 sets the scene moving.

Common mistakes / failure modes

"the average of the plausible answers is the safest answer"
Safest in squared error (12.5 cm against 17.7 for a draw), and a valid object in only 24% of the balanced fronts (§1, §2).
"a diffusion model escapes regression"
It is regression at every noise level, used for small steps: the estimate is the average at the first step (0% valid) and an average over one hill once the state has leaned to a side (§3).
"score distillation draws samples from the 2D model"
It climbs: 0 of 200 climbs on front A ended as the oval, though the posterior is 50.1% oval (§4).
"draw each view separately and fuse them"
Four independent views agree on the kind 12.5% of the time (§5).
"ask again and you get the same scene"
Every request is a new draw: consecutive draws change kind in 50.0% of pairs on a coin-flip front (§6).

Checkpoint exercise

Try it
The back of an object is at radius 0.8 with probability 0.3 and at 1.2 with probability 0.7. (a) Which answer minimises the expected squared error, and what is that error? (b) What is the expected squared error of a draw? (c) How often do two consecutive independent draws differ? (d) Is the answer of (a) a possible back, and how far is it from the nearer one? Answer: (a) the mean, 0.3·0.8 + 0.7·1.2 = 1.08, with expected squared error 0.3·0.28² + 0.7·0.12² = 0.0336. (b) Twice that, 0.0672, because a draw adds the posterior's own spread. (c) 2·0.3·0.7 = 0.42. (d) No: a back is 0.8 or 1.2 and 1.08 is 0.12 from the nearer.

Where this points next

A draw gave the unseen part of a scene a valid answer where the average gave none (96% of draws against 24% of averages), for a factor √2 in error, and a diffusion model makes draws from regression in small steps. But a draw describes one still moment: it is the true kind 51% of the time, asking again changes the kind in 50% of consecutive pairs, and a draw left in place goes stale as the object turns. What does it take to represent a scene that changes, and then to predict it?

Takeaway
A network trained by squared error returns the posterior mean, E‖θ − θ̂‖² = E‖θ − m‖² + ‖m − θ̂‖², and where the unseen part has two kinds of answer the mean lies between them: it fits the photographs and is a valid object in only 24% of the ambiguous cases. The remedy is a draw from the posterior, valid 96% of the time and varied, for a factor √2 in error. A denoiser makes draws without writing the posterior: regression at every noise level, asked only for small steps. Climbing a 2D model (score distillation) descends a KL divergence from a point and finds the nearest mode, not a draw: 0 of 200 ovals on a coin-flip front, and the guidance that sharpens it leaves 61% valid at 100. Independent views disagree (four agree 12.5% of the time), so views are drawn jointly, or the scene itself. A draw is plausible, not true, and new at every request.

Interview prompts

Companion reads: Computer Vision · 17 Generative vision (diffusion and classifier-free guidance in 2D) and World Models · 05 One future is a lie (the same mean-versus-draw problem for futures, and mixture densities).