all_lessons/3D Vision/11 · When the views run out: learning across sceneslesson 11 / 14

When the views run out: learning across scenes

Lesson 10 made a scene quick to fit and left it as hungry for views as ever: show a grid or a set of Gaussians three views and it explains those three photographs and guesses the rest. What is missing has a precise form, directions in the scene along which two different scenes give the same photographs, and more views help only if they look at the unseen side. The other source is other scenes: a stack of fully measured scenes becomes a prior, and the unseen side gets a best guess and an error bar in closed form, honest only inside the family the prior was learned from and impossible to store for a real scene.

The thesis, here
When the photographs leave part of a scene undetermined, nothing in them can settle it, and the one place information exists in bulk is other scenes. Write what other scenes teach as a probability distribution over scenes, and the unseen part gets a best guess and an error bar that is widest exactly where the cameras did not look. The distribution is only as good as its family of scenes, and for real scenes it must live in a trained function, not in a matrix.
Linear position
Forced by: Grids and Gaussians fit one scene in minutes and draw it in real time, but only by being shown that scene from many views. Give them three views, or one, and the unseen part is unconstrained: they score well on the photographs they were fitted to and fail the exam. Where can the missing information come from, if not from more photographs?
New idea: information from outside the data enters as a prior learned over many scenes: with a linear measurement and a Gaussian prior the posterior is closed form, as tight as the photographs where they look and as wide as the prior where they do not. Its error bar is honest only for scenes like the ones it was learned from.
Forces next: The missing information comes from a prior learned over many scenes, and a prior has to live in a network. But a 3D input is not an image: a point cloud is an unordered set, a voxel grid is almost entirely empty, a LiDAR sweep has a different length every time. What must a network respect to learn from 3D data, and what can it then answer?
The plan
Seven moves. (1) Count what the photographs fix and exhibit two scenes they cannot tell apart. (2) Cross out the ways to get more information, leaving other scenes. (3) Turn a stack of scenes into a prior and solve for the posterior. (4) Compare hand-made and learned priors by error and by the honesty of their error bars. (5) Take the prior out of its family. (6) Place today's networks on the same ladder. (7) Price the closed form at real size, and find that the prior must live in a function.

1 · What the photographs cannot say

Lesson 10 ended on a fit that is exact where the scene was shown and arbitrary where it was not. Count why. Lesson 9's grid field has 40 × 40 nodes of four numbers, 6400 unknowns; three photographs of 64 pixels and three channels are 576 numbers, so to first order at least 5824 directions of the field can change without changing any photograph, and the photographs do not choose among them. Plain gradient descent on a linear problem leaves the starting value along every such direction (§2). Adam, which lesson 9 used, divides each step by its parameter's own gradient scale and does not: behind a wall the gradient shrinks by the factor T (lesson 8) but is not zero, and in the three-camera fit all 1600 nodes moved, the node with the faintest gradient (1,180 times below the median node's) by 3.3 in raw density, the median node by 2.4. So the unseen part of a fitted field is its starting value plus whatever the rays that did reach it pushed it by, each push rescaled by Adam: the optimiser's doing, not the photographs'.

To see this structure with ten unknowns instead of 6400, shrink the object. A star-shaped object has an outline that is a radius R(φ) in every direction φ (seen from above, counter-clockwise from +x). In units of its mean radius (1.15 m for the statue):

R(φ) = 1 + Σk=1..5 ( ak cos kφ + bk sin kφ ),   θ = (a1, b1, …, a5, b5) ∈ ℝ10

The ten numbers θ are the scene. (Five harmonics already forbid sharp corners, which is a prior built into the toy.) A depth camera at azimuth α reads R in nine directions across the arc it faces, |φ − α| ≤ 50°; beyond that the surface is too slanted to read. Take the 1 away from the nine readings and call the result r; it is linear in θ:

r = Aθ + n,   row of A for direction φ: (cos φ, sin φ, cos 2φ, sin 2φ, …, cos 5φ, sin 5φ),   n ~ N(0, σ²I), σ = 0.01 (1.2 cm)

One camera gives nine equations in ten unknowns, so the problem is ill-posed: A has a null space, a unit vector v with Av = 0. Add 0.2v to the statue's θ. The nine readings change by less than 10−12, and the outline behind the statue moves by up to 34 cm. These are two different statues and one photograph.

With noise, the null space blurs into a spectrum. The singular values si of A say how far the readings move when θ moves one unit along direction i; a direction with si < σ is lost in the noise: the readings fix it only to ±σ/si, more than a whole radius. One camera leaves 3 of the ten directions like that; each further camera, 30° along a path past the statue, brings the count down to 2, 1 and 0. But the back of the statue stays unread: 72% of the outline with one camera, 47% with four.

Now run the exam of this track in the toy: predict the radius in the directions no camera read, scoring the error in centimetres averaged over 1000 objects from the world of §4. Any answer that fits the readings is as good as any other as far as the data go. The plainest rule, the smallest θ that fits (in effect a unit Gaussian on each of the ten numbers, §3; the widget calls it none), is off by 38 cm, 2.8 times worse than guessing a plain circle (13 cm). The photographs fit and the exam fails, and no amount of fitting changes that, because what is missing is not in the data.

2 · Where else could information come from?

The photographs cannot choose, so something else must. Four candidates, priced in the toy (σ = 0.01; averages over 1000 objects from the world of §4).

SourceWhat it costsWhat the toy says
more photographs of this scenea camera position eachFour cameras walking past one side, no prior: 27.5 cm, 47% of the outline unread. Two cameras on opposite sides do fix the toy (4.3 cm) if you can stand behind it.
a smoothness penalty chosen by handnothing9.6 cm with one camera, and an error bar that is right by luck (§4).
stop the optimiser earlynothingGradient descent from zero moves along direction i at a rate set by si2, and not at all along the null direction. With step 1/s12, 100 steps carry direction 7 (s7 = 0.035) 0.8% of the way: the starting value is the answer there, a prior chosen by accident.
photographs of other scenesone fully measured scene each, paid oncethe rest of the lesson

Only the last asks the right question: what do scenes look like? Lessons 1 to 10 turn plentiful views of one scene into its parameters θ. Run that machine on 400 objects of the world and keep the 400 θ's. Each is a fully measured scene, and together they say which unseen backs go with which seen fronts.

3 · From a stack of scenes to a constraint

Summarise the stack by its mean μ and covariance Λ and take the Gaussian with that mean and covariance as the prior, θ ~ N(μ, Λ). It is the simplest distribution that remembers which numbers go together, and the one for which the algebra closes. Bayes' rule revises it by the readings:

log p(θ | r) = − ‖r − Aθ‖² / 2σ² − (θ − μ)T Λ−1 (θ − μ) / 2 + const

The right side is quadratic in θ, so the posterior is Gaussian, and completing the square gives its covariance S and mean m:

S = ( Λ−1 + ATA / σ² )−1,   m = S ( Λ−1μ + ATr / σ² )

Read it twice. The mean m is the least-squares fit with a penalty (θ − μ)TΛ−1(θ − μ) (the MAP estimate), which makes it expensive to move along directions the stack rarely varied. And S says what is known: the data add ATA/σ² to the precision S−1, large along directions the cameras see and zero on the null space, where S−1v = Λ−1v and the answer is whatever the prior said. The error bar on the radius R = 1 + hTθ in direction φ, with h = (cos φ, sin φ, …), is √(hTSh): narrow where cameras looked, as wide as the prior elsewhere.

The twin of §1 shows the mechanism. The photographs cannot rank the statue and its twin; the penalty can.

penalty (θ − μ)TΛ−1(θ − μ)unit Gaussiansmooth (§4)learned (§4)
the statue0.0313.02.2
its twin0.0729.8237.5

A prior does not remove the ambiguity (the twin still fits to 10−12). It ranks the scenes that fit equally well, and a stack of other scenes ranks them by how much they resemble those scenes.

4 · Hand-made against learned

Where do μ and Λ come from? A hand-made prior says rounded shapes are likelier: μ = 0, harmonic k with standard deviation 0.12/k. A learned prior takes the sample mean and covariance of the stack. The toy world makes its objects from three traits (elongated, three-lobed, lopsided) plus a little noise, and Λ has to find that out from examples. It does: its eigenvalues are 0.019, 0.012 and 0.006, then seven near 0.00015. Three directions carry 97.3% of the variance, and along the other seven the objects differ by only 1.3 to 1.5 cm. The ten numbers of an object have three degrees of freedom, which a smoothness prior cannot know.

Score the three priors on 1000 objects drawn from that world, one camera, σ = 0.01:

priorerror on the unseen sideerror bar it claims (1σ)truth inside its 2σ band
unit Gaussian (none)37.8 cm146 cm99.8%
smooth, hand-made9.6 cm10.1 cm94.6%
learned from 400 scenes3.9 cm3.8 cm95.0%

The learned prior is 2.5 times more accurate than the smooth one, and its bar equals its error with 95% of the truth inside the band, which is what a 2σ band of a Gaussian should hold: it is calibrated. The smooth prior is calibrated at one camera but not at four, where its bar has shrunk from 10.1 to 6.5 cm while its error has not moved (9.6 to 9.8 cm), so the band holds only 82%. A hand-made prior can be wrong about how much it knows, though it stays the right choice when there are no scenes to learn from.

How many scenes does the learned prior need? Averaged over twenty stacks, with one camera: 5 scenes err by 9.3 cm, close to the smooth prior, and their bars hold only 27% of the truth (a covariance estimated from 5 objects believes the world has four dimensions); 20 scenes give 4.5 cm and 86%, and 100 give 3.9 cm and 94%, as 400 do.

The unseen side, with and without a prior
Top left: the object from above. Black is the truth (dashed where no camera looked), green dots are readings, blue is the posterior mean with its ±2σ band, violet outlines are three scenes drawn from the posterior. Right: the same curves unrolled over φ. Bottom left: the singular values of A against the noise σ (red). The readouts compare this object with 1000 from the same world.
unseen share of outline
—
directions below the noise
—
unseen error, this object
—
error bar (1σ), unseen
—
inside 2σ band, seen
—
inside 2σ band, unseen
—
typical error, 1000 objects
—
inside band, 1000 objects
—
Show the core JS
L11.fit = function (phis, sigma, prior) {
  var m = phis.length, A = SH.design(phis), At = SH.mT(A, m, D), P = SH.mmul(At, D, m, A, D), s2 = sigma * sigma, i;
  for (i = 0; i < D * D; i++) P[i] = prior.LamInv[i] + P[i] / s2;
  var S = SH.minv(P, D), K = SH.mmul(S, D, D, At, m);
  for (i = 0; i < K.length; i++) K[i] /= s2;
  var m0 = SH.mmul(S, D, D, SH.mmul(prior.LamInv, D, D, prior.mu, 1), 1);
  return { m: m, A: A, S: S, K: K, m0: m0 };
};
...
  for (i = 0; i < D; i++) for (j = 0; j < fit.m; j++) out[i] += fit.K[i * fit.m + j] * r[j];
...
    for (i = 0; i < D; i++) for (k = 0; k < D; k++) v += H[j * D + i] * S[i * D + k] * H[j * D + k];
    out[j] = Math.sqrt(Math.max(v, 0));

What to try. The page opens with one camera, no prior and the statue: 72% of the outline unseen, 3 of the ten bars under the noise line, and violet outlines that all fit the green readings and disagree behind the statue by metres (bar 146 cm; this draw errs by 13.9 cm). Choose smooth: bar 10.1 cm, error 7.9. Choose learned: bar 3.8, error 1.2; the statue is an easy object (0.7 to 2.1 cm over ten presses of another draw), and the typical error over 1000 objects is 3.9 cm with 95% of the truth inside the band. Add cameras: with the learned prior the typical error only falls to 2.8 cm at four, the prior doing the work; with no prior, four cameras on the walk still err by 27.5 cm and two all around give 4.3. Now the caution: learned prior, one camera, object out of family. The bar stays 3.8 cm, the typical error is 14.4 cm and the band holds only 44% of the truth. Last, back in family, raise the noise to 4.6 cm: the typical error moves only from 3.9 to 4.9 cm, three numbers being all the prior needs.

Road not taken · photograph the unseen side
The tempting fix assumes nothing about scenes: look at what you cannot see. In the toy it works if you can reach it: with no prior, two cameras on opposite sides give 4.3 cm, about what one camera gives with a prior learned from 20 scenes (4.5 cm). Count the cost in camera positions. A fully measured toy scene takes four cameras all around, so the 20 scenes cost 80 positions once, and the far-side camera costs one more position for every scene reconstructed afterwards: the prior is cheaper from the 81st scene. And cameras do not reach everywhere: four cameras walking past one side still err by 27.5 cm with no prior, and a car on a road or a photograph from the internet has no behind to stand in.

5 · A prior is only as good as its family

The out-of-family objects come from a world with a fourth trait that none of the 400 training objects had: an ordinary object squared off, its outline changed by −0.09 cos 2φ − 0.12 cos 4φ + 0.05 sin 4φ + 0.07 cos 5φ times a factor near 1. The learned prior is applied exactly as before, and it is wrong about this world. Three things follow.

First, the bar does not move: 3.8 cm, because the posterior covariance depends on A, σ and Λ and never on the readings. The error does: 14.4 cm, 3.8 times the bar, with 44% of the truth inside the band instead of 95%, and four cameras on the walk do not cure it (37% inside). The smooth prior is no better (17.2 cm) but less surprised: its bar is 10.1 cm and 79% of the truth is inside, because it never claimed to know much.

Second, the photographs hardly warn you. The prior supplies a typical back whatever the object, and the seen side is still fitted almost as well: the rms residual rises only from 0.65σ to 0.96σ, too little to notice on one object.

Third, an error bar is conditional. It says what the prior and the data together know if the prior is true, and nothing about whether it is. A learned prior is sharper than a hand-made one by exactly what it assumes about the world, and an object from outside is where that assumption is paid for.

6 · The same move on the ladder of today's systems

The recipe is complete in miniature: fully measured scenes supply a prior, and a prior turns an under-determined measurement into a guess with an error bar. The depth and reconstruction networks of the last decade are this recipe at scale, entering the pipeline at different rungs.

SystemPrior learned overGap it fillsWhere it enters
MVSNet (Yao et al., 2018)depth maps, given warped image featuresthe ambiguity of lesson 7's photo-consistency costinside geometry: homography warping builds the cost volume, a 3D network regularises it
MiDaS (Ranftl et al., 2022), Depth Anything (Yang et al., 2024)depth, given one imagea whole depth channel, up to scale and shiftinstead of a sensor; relative output
Metric3D (Yin et al., 2023), UniDepth (Piccinelli et al., 2024), Depth Pro (Bochkovskii et al., 2025)the same, plus sizes of things and camerasthe scale and the focal lengthinstead of a sensor; metric output
DUSt3R (Wang et al., 2024), MASt3R (Leroy et al., 2024)pointmaps of image pairs, with matching features (MASt3R)pose and structure with no calibration giveninstead of the matching and bundle adjustment of lesson 5
VGGT (Wang et al., 2025)cameras, depth, point maps and tracks for one to hundreds of viewspose, depth and points for many views at onceamortised: one forward pass instead of an optimisation

What one image cannot know. The pixel u = W/2 + fX/Z is the same for (X, Z) and (sX, sZ), so every single-image problem has a null space: lesson 1's crate, billboard and wall share the same 17.6 px. If the focal length is unknown too, (f, Z) → (cf, cZ) is a second. Eigen et al. (2014) called the global scale a fundamental ambiguity of depth prediction: telling their network the correct mean log depth, nothing else, lowered its log RMSE from 0.28 to 0.22. A relative-depth network is trained so that the null space is projected out of the loss: MiDaS predicts disparity up to scale and shift and fits both to the truth (least squares or a robust variant) before comparing. A metric network must fill it from learned sizes and cameras: Metric3D maps every image into a canonical camera, and UniDepth and Depth Pro work without intrinsics, Depth Pro by estimating the focal length itself.

What can the learned guess promise? Suppose experience says the thing is 1 m wide give or take a factor e0.5 = 1.65 (a Gaussian prior on ln w with σ = 0.5). The 17.6 px fix ln Z − ln w, so ln Z has the same Gaussian, shifted: a median of 2.5 m and, one standard deviation either way, 1.5 to 4.1 m, a factor 1.65 at every distance. The wall of lesson 1 (4 m wide, 10 m away) has the same 17.6 px and lies 2.8 standard deviations from the prior's median, at 0.02 of the crate's density: the network says "crate, 2.5 m", with confidence. That is lesson 1's road not taken with its promise made precise: a posterior over the missing number, calibrated for scenes like the experience and no other.

What the last rungs buy. DUSt3R learned its pointmaps from 8.5 million image pairs. VGGT predicts cameras, depth maps, point maps and tracks for one to hundreds of views in one pass: it was trained on 64 A100 GPUs for nine days (13,824 GPU-hours) and takes about 0.2 s for ten frames on one H100, against about 10 s for DUSt3R or MASt3R with global alignment, 50 times faster. The prior is paid for once, in training; each new scene costs a forward pass. Successors keep moving the null spaces around: π³ (Wang et al., 2025) drops the fixed reference view and predicts affine-invariant poses, and MapAnything (Keetha et al., 2025) takes optional intrinsics, poses or depth and returns metric geometry.

7 · The closed form does not scale: put the prior in a function

The toy's posterior costs one 10 × 10 inverse. Price it for a real scene. A voxel grid of 128³ has d = 2,097,152 unknowns (221), so its covariance Λ has d² = 4.4×1012 entries, 17.6 TB at four bytes each, and the posterior needs a d × d solve of about d³ = 9.2×1018 operations. Estimating Λ is no easier: a covariance estimated from M scenes has rank at most M − 1, so 400 scenes, plenty for our 100-entry matrix, are nothing for two million unknowns. The matrix can be neither stored nor estimated.

What can be kept is the map. The posterior mean is a function of the readings alone, m = m0 + K r with K = SAT/σ², and it can be fitted directly from examples without forming Λ or S: simulate N objects from the world with their readings and solve the least-squares problem readings → θ. With 10 examples the fit is no better than guessing the mean shape (13.0 cm against 13.3); with 100 it is within a few percent of the closed form (4.0 cm against 3.9), and with 2000 it is indistinguishable (3.85 against 3.87 cm). This is amortised inference: pay for the training scenes once, then answer each new scene in one pass. A network is the same with hidden layers, for when the posterior mean is not linear in the readings.

But look at what the fitted map consumed: nine numbers in a fixed order from a camera at a fixed place. The closed form consumed pairs (direction, reading). Test both on the same 1000 objects:

the input is…closed formfitted map
the nine readings as trained3.87 cm3.85 cm
the same readings, listed in a random order3.87 cm17.1 cm
two of the nine returns missing3.9 cm7.1 cm
the camera moved by 25°3.8 cm15.7 cm

The closed form barely notices: it adds one term per (direction, reading) pair, and a sum cares about neither order nor count. The fitted map reads a list, and a shuffled list is worse than not looking (17.1 cm, against 13.3 cm for guessing the mean shape). Real 3D data brings all of these at once: a point cloud has no order, a LiDAR sweep returns a different number of points every time, and a voxel grid is almost entirely empty. The function that carries the prior has to be built around that.

What this lesson did not do
It worked on ten numbers with a Gaussian. A real scene has millions of unknowns, and a real family can have kinds: when two different backs are both plausible the posterior has two hills and its mean is neither (lessons 12 and 13). It priced the covariance of a voxel grid but did not train a prior over fields or splats, and it measured calibration only in the toy: whether the confidences of depth and reconstruction networks hold out of family is the same question and was not tested. The fitted map reads a list in a fixed order; what a network must respect to read points, voxels and sweeps is lesson 12. And everything held still: lesson 14 moves it.

Common mistakes / failure modes

"a prior removes the ambiguity"
It ranks scenes that fit equally well. The twin of §1 still fits to 10−12; the doubt moves into the error bar, wide where no camera looked (§3).
"the error bar says how wrong the answer is"
It says how wrong, if the prior is right. Out of family the bar stays 3.8 cm while the error is 14.4 cm and the band holds 44% (§5).
"smoothness regularisation assumes nothing"
A penalty θTΛ−1θ is a Gaussian prior, and smoothness is the one with standard deviation 0.12/k, whose band holds 82% at four cameras (§3, §4). The posterior adds what a penalty throws away: the covariance, the error bar.
"monocular depth measures depth"
It estimates from experience; scale is not in the pixels. A crate at 2.5 m and a wall at 10 m give the same 17.6 px (§6).
"a map fitted on enough examples can read any input"
The map of §7 matched the closed form on its own layout and was 4.4 times worse on a shuffled list.

Checkpoint exercise

Try it
One radius deviation has the prior N(0, 0.10²). A camera reads it once, with noise σ = 0.05, and gets +0.15. (a) What are the posterior mean and standard deviation? (b) Repeat with a prior standard deviation of 0.01. (c) What is the posterior of a number the camera cannot see, in centimetres for the 1.15 m statue? Answer: (a) the precision is 1/0.10² + 1/0.05² = 500, so the standard deviation is 0.045 and the mean is (0.15/0.05²)/500 = 0.12: the reading gets weight 4/5. (b) The precision is 104 + 400, giving mean 0.0058 and standard deviation 0.0098: a very confident prior ignores the reading, which is §5 in one line. (c) Nothing is added to the precision, so the posterior is the prior: mean 0, standard deviation 0.10, that is 11.5 cm.

Where this points next

Other scenes supplied what the photographs could not. One camera and a prior learned from 400 scenes gave the unseen outline of a typical object to 3.9 cm, with an error bar of 3.8 cm that held 95% of the truth, against 38 cm with no prior. Two limits showed. Out of its family the bar stayed at 3.8 cm while the error became 14.4 cm. And at the size of a real scene the matrix needs 17.6 TB, so it must be replaced by a function fitted to many scenes; the function we fitted reproduced the closed form (3.85 cm) and broke exactly where 3D data are irregular: the same nine readings in a different order took its error to 17.1 cm and two missing returns to 7.1, while the closed form did not move. What must a network respect to learn from 3D data, and what can it then answer?

Takeaway
When the photographs leave directions of a scene unconstrained (at least 5824 for a 40 × 40 field), two scenes that differ there give the same photographs, and only information from outside the data can choose between them. That information exists in bulk as other scenes: a stack of fully measured scenes gives a mean μ and a covariance Λ, and with a linear measurement the posterior is Gaussian in closed form, as tight as the cameras where they looked and as wide as the prior where they did not. A learned prior was 2.5 times more accurate than a hand-made smoothness prior and calibrated (95% of the truth inside its 2σ band), but it needs enough scenes and is honest only inside its family: out of family its bar stayed at 3.8 cm while its error became 14.4 cm. Monocular depth and feed-forward reconstruction run this recipe at scale, with scale as the null space of one image. At real size the covariance cannot be stored (17.6 TB for 128³ voxels), so the prior is kept as a function fitted to many scenes and applied in one pass, which is brittle where 3D data are irregular.

Interview prompts

Companion reads: Computer Vision · 05 Multi-view, depth & SLAM (the classical pipeline whose gaps this lesson fills with learned priors) and Synthetic Vision · 01 Free labels, and the wall (training scenes a program draws, and what they cannot teach).