When the views run out: learning across scenes
Lesson 10 made a scene quick to fit and left it as hungry for views as ever: show a grid or a set of Gaussians three views and it explains those three photographs and guesses the rest. What is missing has a precise form, directions in the scene along which two different scenes give the same photographs, and more views help only if they look at the unseen side. The other source is other scenes: a stack of fully measured scenes becomes a prior, and the unseen side gets a best guess and an error bar in closed form, honest only inside the family the prior was learned from and impossible to store for a real scene.
New idea: information from outside the data enters as a prior learned over many scenes: with a linear measurement and a Gaussian prior the posterior is closed form, as tight as the photographs where they look and as wide as the prior where they do not. Its error bar is honest only for scenes like the ones it was learned from.
Forces next: The missing information comes from a prior learned over many scenes, and a prior has to live in a network. But a 3D input is not an image: a point cloud is an unordered set, a voxel grid is almost entirely empty, a LiDAR sweep has a different length every time. What must a network respect to learn from 3D data, and what can it then answer?
1 · What the photographs cannot say
Lesson 10 ended on a fit that is exact where the scene was shown and arbitrary where it was not. Count why. Lesson 9's grid field has 40 × 40 nodes of four numbers, 6400 unknowns; three photographs of 64 pixels and three channels are 576 numbers, so to first order at least 5824 directions of the field can change without changing any photograph, and the photographs do not choose among them. Plain gradient descent on a linear problem leaves the starting value along every such direction (§2). Adam, which lesson 9 used, divides each step by its parameter's own gradient scale and does not: behind a wall the gradient shrinks by the factor T (lesson 8) but is not zero, and in the three-camera fit all 1600 nodes moved, the node with the faintest gradient (1,180 times below the median node's) by 3.3 in raw density, the median node by 2.4. So the unseen part of a fitted field is its starting value plus whatever the rays that did reach it pushed it by, each push rescaled by Adam: the optimiser's doing, not the photographs'.
To see this structure with ten unknowns instead of 6400, shrink the object. A star-shaped object has an outline that is a radius R(φ) in every direction φ (seen from above, counter-clockwise from +x). In units of its mean radius (1.15 m for the statue):
R(φ) = 1 + Σk=1..5 ( ak cos kφ + bk sin kφ ), θ = (a1, b1, …, a5, b5) ∈ ℝ10
The ten numbers θ are the scene. (Five harmonics already forbid sharp corners, which is a prior built into the toy.) A depth camera at azimuth α reads R in nine directions across the arc it faces, |φ − α| ≤ 50°; beyond that the surface is too slanted to read. Take the 1 away from the nine readings and call the result r; it is linear in θ:
r = Aθ + n, row of A for direction φ: (cos φ, sin φ, cos 2φ, sin 2φ, …, cos 5φ, sin 5φ), n ~ N(0, σ²I), σ = 0.01 (1.2 cm)
One camera gives nine equations in ten unknowns, so the problem is ill-posed: A has a null space, a unit vector v with Av = 0. Add 0.2v to the statue's θ. The nine readings change by less than 10−12, and the outline behind the statue moves by up to 34 cm. These are two different statues and one photograph.
With noise, the null space blurs into a spectrum. The singular values si of A say how far the readings move when θ moves one unit along direction i; a direction with si < σ is lost in the noise: the readings fix it only to ±σ/si, more than a whole radius. One camera leaves 3 of the ten directions like that; each further camera, 30° along a path past the statue, brings the count down to 2, 1 and 0. But the back of the statue stays unread: 72% of the outline with one camera, 47% with four.
Now run the exam of this track in the toy: predict the radius in the directions no camera read, scoring the error in centimetres averaged over 1000 objects from the world of §4. Any answer that fits the readings is as good as any other as far as the data go. The plainest rule, the smallest θ that fits (in effect a unit Gaussian on each of the ten numbers, §3; the widget calls it none), is off by 38 cm, 2.8 times worse than guessing a plain circle (13 cm). The photographs fit and the exam fails, and no amount of fitting changes that, because what is missing is not in the data.
2 · Where else could information come from?
The photographs cannot choose, so something else must. Four candidates, priced in the toy (σ = 0.01; averages over 1000 objects from the world of §4).
| Source | What it costs | What the toy says |
|---|---|---|
| more photographs of this scene | a camera position each | Four cameras walking past one side, no prior: 27.5 cm, 47% of the outline unread. Two cameras on opposite sides do fix the toy (4.3 cm) if you can stand behind it. |
| a smoothness penalty chosen by hand | nothing | 9.6 cm with one camera, and an error bar that is right by luck (§4). |
| stop the optimiser early | nothing | Gradient descent from zero moves along direction i at a rate set by si2, and not at all along the null direction. With step 1/s12, 100 steps carry direction 7 (s7 = 0.035) 0.8% of the way: the starting value is the answer there, a prior chosen by accident. |
| photographs of other scenes | one fully measured scene each, paid once | the rest of the lesson |
Only the last asks the right question: what do scenes look like? Lessons 1 to 10 turn plentiful views of one scene into its parameters θ. Run that machine on 400 objects of the world and keep the 400 θ's. Each is a fully measured scene, and together they say which unseen backs go with which seen fronts.
3 · From a stack of scenes to a constraint
Summarise the stack by its mean μ and covariance Λ and take the Gaussian with that mean and covariance as the prior, θ ~ N(μ, Λ). It is the simplest distribution that remembers which numbers go together, and the one for which the algebra closes. Bayes' rule revises it by the readings:
log p(θ | r) = − ‖r − Aθ‖² / 2σ² − (θ − μ)T Λ−1 (θ − μ) / 2 + const
The right side is quadratic in θ, so the posterior is Gaussian, and completing the square gives its covariance S and mean m:
S = ( Λ−1 + ATA / σ² )−1, m = S ( Λ−1μ + ATr / σ² )
Read it twice. The mean m is the least-squares fit with a penalty (θ − μ)TΛ−1(θ − μ) (the MAP estimate), which makes it expensive to move along directions the stack rarely varied. And S says what is known: the data add ATA/σ² to the precision S−1, large along directions the cameras see and zero on the null space, where S−1v = Λ−1v and the answer is whatever the prior said. The error bar on the radius R = 1 + hTθ in direction φ, with h = (cos φ, sin φ, …), is √(hTSh): narrow where cameras looked, as wide as the prior elsewhere.
The twin of §1 shows the mechanism. The photographs cannot rank the statue and its twin; the penalty can.
| penalty (θ − μ)TΛ−1(θ − μ) | unit Gaussian | smooth (§4) | learned (§4) |
|---|---|---|---|
| the statue | 0.03 | 13.0 | 2.2 |
| its twin | 0.07 | 29.8 | 237.5 |
A prior does not remove the ambiguity (the twin still fits to 10−12). It ranks the scenes that fit equally well, and a stack of other scenes ranks them by how much they resemble those scenes.
4 · Hand-made against learned
Where do μ and Λ come from? A hand-made prior says rounded shapes are likelier: μ = 0, harmonic k with standard deviation 0.12/k. A learned prior takes the sample mean and covariance of the stack. The toy world makes its objects from three traits (elongated, three-lobed, lopsided) plus a little noise, and Λ has to find that out from examples. It does: its eigenvalues are 0.019, 0.012 and 0.006, then seven near 0.00015. Three directions carry 97.3% of the variance, and along the other seven the objects differ by only 1.3 to 1.5 cm. The ten numbers of an object have three degrees of freedom, which a smoothness prior cannot know.
Score the three priors on 1000 objects drawn from that world, one camera, σ = 0.01:
| prior | error on the unseen side | error bar it claims (1σ) | truth inside its 2σ band |
|---|---|---|---|
| unit Gaussian (none) | 37.8 cm | 146 cm | 99.8% |
| smooth, hand-made | 9.6 cm | 10.1 cm | 94.6% |
| learned from 400 scenes | 3.9 cm | 3.8 cm | 95.0% |
The learned prior is 2.5 times more accurate than the smooth one, and its bar equals its error with 95% of the truth inside the band, which is what a 2σ band of a Gaussian should hold: it is calibrated. The smooth prior is calibrated at one camera but not at four, where its bar has shrunk from 10.1 to 6.5 cm while its error has not moved (9.6 to 9.8 cm), so the band holds only 82%. A hand-made prior can be wrong about how much it knows, though it stays the right choice when there are no scenes to learn from.
How many scenes does the learned prior need? Averaged over twenty stacks, with one camera: 5 scenes err by 9.3 cm, close to the smooth prior, and their bars hold only 27% of the truth (a covariance estimated from 5 objects believes the world has four dimensions); 20 scenes give 4.5 cm and 86%, and 100 give 3.9 cm and 94%, as 400 do.
What to try. The page opens with one camera, no prior and the statue: 72% of the outline unseen, 3 of the ten bars under the noise line, and violet outlines that all fit the green readings and disagree behind the statue by metres (bar 146 cm; this draw errs by 13.9 cm). Choose smooth: bar 10.1 cm, error 7.9. Choose learned: bar 3.8, error 1.2; the statue is an easy object (0.7 to 2.1 cm over ten presses of another draw), and the typical error over 1000 objects is 3.9 cm with 95% of the truth inside the band. Add cameras: with the learned prior the typical error only falls to 2.8 cm at four, the prior doing the work; with no prior, four cameras on the walk still err by 27.5 cm and two all around give 4.3. Now the caution: learned prior, one camera, object out of family. The bar stays 3.8 cm, the typical error is 14.4 cm and the band holds only 44% of the truth. Last, back in family, raise the noise to 4.6 cm: the typical error moves only from 3.9 to 4.9 cm, three numbers being all the prior needs.
5 · A prior is only as good as its family
The out-of-family objects come from a world with a fourth trait that none of the 400 training objects had: an ordinary object squared off, its outline changed by −0.09 cos 2φ − 0.12 cos 4φ + 0.05 sin 4φ + 0.07 cos 5φ times a factor near 1. The learned prior is applied exactly as before, and it is wrong about this world. Three things follow.
First, the bar does not move: 3.8 cm, because the posterior covariance depends on A, σ and Λ and never on the readings. The error does: 14.4 cm, 3.8 times the bar, with 44% of the truth inside the band instead of 95%, and four cameras on the walk do not cure it (37% inside). The smooth prior is no better (17.2 cm) but less surprised: its bar is 10.1 cm and 79% of the truth is inside, because it never claimed to know much.
Second, the photographs hardly warn you. The prior supplies a typical back whatever the object, and the seen side is still fitted almost as well: the rms residual rises only from 0.65σ to 0.96σ, too little to notice on one object.
Third, an error bar is conditional. It says what the prior and the data together know if the prior is true, and nothing about whether it is. A learned prior is sharper than a hand-made one by exactly what it assumes about the world, and an object from outside is where that assumption is paid for.
6 · The same move on the ladder of today's systems
The recipe is complete in miniature: fully measured scenes supply a prior, and a prior turns an under-determined measurement into a guess with an error bar. The depth and reconstruction networks of the last decade are this recipe at scale, entering the pipeline at different rungs.
| System | Prior learned over | Gap it fills | Where it enters |
|---|---|---|---|
| MVSNet (Yao et al., 2018) | depth maps, given warped image features | the ambiguity of lesson 7's photo-consistency cost | inside geometry: homography warping builds the cost volume, a 3D network regularises it |
| MiDaS (Ranftl et al., 2022), Depth Anything (Yang et al., 2024) | depth, given one image | a whole depth channel, up to scale and shift | instead of a sensor; relative output |
| Metric3D (Yin et al., 2023), UniDepth (Piccinelli et al., 2024), Depth Pro (Bochkovskii et al., 2025) | the same, plus sizes of things and cameras | the scale and the focal length | instead of a sensor; metric output |
| DUSt3R (Wang et al., 2024), MASt3R (Leroy et al., 2024) | pointmaps of image pairs, with matching features (MASt3R) | pose and structure with no calibration given | instead of the matching and bundle adjustment of lesson 5 |
| VGGT (Wang et al., 2025) | cameras, depth, point maps and tracks for one to hundreds of views | pose, depth and points for many views at once | amortised: one forward pass instead of an optimisation |
What one image cannot know. The pixel u = W/2 + fX/Z is the same for (X, Z) and (sX, sZ), so every single-image problem has a null space: lesson 1's crate, billboard and wall share the same 17.6 px. If the focal length is unknown too, (f, Z) → (cf, cZ) is a second. Eigen et al. (2014) called the global scale a fundamental ambiguity of depth prediction: telling their network the correct mean log depth, nothing else, lowered its log RMSE from 0.28 to 0.22. A relative-depth network is trained so that the null space is projected out of the loss: MiDaS predicts disparity up to scale and shift and fits both to the truth (least squares or a robust variant) before comparing. A metric network must fill it from learned sizes and cameras: Metric3D maps every image into a canonical camera, and UniDepth and Depth Pro work without intrinsics, Depth Pro by estimating the focal length itself.
What can the learned guess promise? Suppose experience says the thing is 1 m wide give or take a factor e0.5 = 1.65 (a Gaussian prior on ln w with σ = 0.5). The 17.6 px fix ln Z − ln w, so ln Z has the same Gaussian, shifted: a median of 2.5 m and, one standard deviation either way, 1.5 to 4.1 m, a factor 1.65 at every distance. The wall of lesson 1 (4 m wide, 10 m away) has the same 17.6 px and lies 2.8 standard deviations from the prior's median, at 0.02 of the crate's density: the network says "crate, 2.5 m", with confidence. That is lesson 1's road not taken with its promise made precise: a posterior over the missing number, calibrated for scenes like the experience and no other.
What the last rungs buy. DUSt3R learned its pointmaps from 8.5 million image pairs. VGGT predicts cameras, depth maps, point maps and tracks for one to hundreds of views in one pass: it was trained on 64 A100 GPUs for nine days (13,824 GPU-hours) and takes about 0.2 s for ten frames on one H100, against about 10 s for DUSt3R or MASt3R with global alignment, 50 times faster. The prior is paid for once, in training; each new scene costs a forward pass. Successors keep moving the null spaces around: π³ (Wang et al., 2025) drops the fixed reference view and predicts affine-invariant poses, and MapAnything (Keetha et al., 2025) takes optional intrinsics, poses or depth and returns metric geometry.
7 · The closed form does not scale: put the prior in a function
The toy's posterior costs one 10 × 10 inverse. Price it for a real scene. A voxel grid of 128³ has d = 2,097,152 unknowns (221), so its covariance Λ has d² = 4.4×1012 entries, 17.6 TB at four bytes each, and the posterior needs a d × d solve of about d³ = 9.2×1018 operations. Estimating Λ is no easier: a covariance estimated from M scenes has rank at most M − 1, so 400 scenes, plenty for our 100-entry matrix, are nothing for two million unknowns. The matrix can be neither stored nor estimated.
What can be kept is the map. The posterior mean is a function of the readings alone, m = m0 + K r with K = SAT/σ², and it can be fitted directly from examples without forming Λ or S: simulate N objects from the world with their readings and solve the least-squares problem readings → θ. With 10 examples the fit is no better than guessing the mean shape (13.0 cm against 13.3); with 100 it is within a few percent of the closed form (4.0 cm against 3.9), and with 2000 it is indistinguishable (3.85 against 3.87 cm). This is amortised inference: pay for the training scenes once, then answer each new scene in one pass. A network is the same with hidden layers, for when the posterior mean is not linear in the readings.
But look at what the fitted map consumed: nine numbers in a fixed order from a camera at a fixed place. The closed form consumed pairs (direction, reading). Test both on the same 1000 objects:
| the input is… | closed form | fitted map |
|---|---|---|
| the nine readings as trained | 3.87 cm | 3.85 cm |
| the same readings, listed in a random order | 3.87 cm | 17.1 cm |
| two of the nine returns missing | 3.9 cm | 7.1 cm |
| the camera moved by 25° | 3.8 cm | 15.7 cm |
The closed form barely notices: it adds one term per (direction, reading) pair, and a sum cares about neither order nor count. The fitted map reads a list, and a shuffled list is worse than not looking (17.1 cm, against 13.3 cm for guessing the mean shape). Real 3D data brings all of these at once: a point cloud has no order, a LiDAR sweep returns a different number of points every time, and a voxel grid is almost entirely empty. The function that carries the prior has to be built around that.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
Other scenes supplied what the photographs could not. One camera and a prior learned from 400 scenes gave the unseen outline of a typical object to 3.9 cm, with an error bar of 3.8 cm that held 95% of the truth, against 38 cm with no prior. Two limits showed. Out of its family the bar stayed at 3.8 cm while the error became 14.4 cm. And at the size of a real scene the matrix needs 17.6 TB, so it must be replaced by a function fitted to many scenes; the function we fitted reproduced the closed form (3.85 cm) and broke exactly where 3D data are irregular: the same nine readings in a different order took its error to 17.1 cm and two missing returns to 7.1, while the closed form did not move. What must a network respect to learn from 3D data, and what can it then answer?
Interview prompts
- Why can a longer training run not repair a model fitted to three views? (§1 — the photographs leave directions of the scene unconstrained; more steps fit the readable ones better and leave the rest to the starting value and the optimiser.)
- Derive the posterior of a linear measurement with a Gaussian prior, and say what it gives along the null space. (§3 — the log posterior is quadratic in θ, and completing the square gives S and m; on the null space S−1v = Λ−1v, so the answer there is the prior.)
- How would you tell a good prior from a bad one using held-out data? (§4 — compare error and coverage: the claimed bar should equal the error and a 2σ band should hold about 95%.)
- Why is a learned prior over-confident out of family, and can the photographs warn you? (§5 — the bar depends on A, σ and Λ, not on the readings; the fit residual rises only from 0.65σ to 0.96σ.)
- Why do monocular depth networks come in relative and metric versions? (§6 — scale is a null space of the image; a relative network projects it out of the loss by fitting scale and shift, a metric one fills it from learned sizes and cameras.)
- Why can the prior over a voxel grid not be stored as a covariance, and what is kept instead? (§7 — it has d² = 4.4×1012 entries; the posterior-mean map is fitted from examples and applied in one pass, and it loses the closed form's indifference to order and count.)
Companion reads: Computer Vision · 05 Multi-view, depth & SLAM (the classical pipeline whose gaps this lesson fills with learned priors) and Synthetic Vision · 01 Free labels, and the wall (training scenes a program draws, and what they cannot teach).