3D Vision, from first principles
A camera turns a three-dimensional world into a flat array of numbers. This track runs that process backwards, and grades every step by one exam: show a model some photographs of a scene, then ask it for the photograph from a camera that took none of them. A pixel is a ray that forgot its distance; the exam is what makes the lost number matter; and each of the fourteen lessons is the move forced by the way the previous one failed the exam. From triangulation and pose, through surfaces, fog, radiance fields and splats, to priors, sampling and a scene that moves, the whole subject comes out of one question.
The exam
Hold one camera out. Give the model the photographs from the others and the cameras' positions, and ask for the photograph the held-out camera took. Score it by PSNR, the logarithm of the mean squared pixel error against the true photograph (higher is better). Lesson 1 sets the exam up with the simplest model there is, a single photograph reused as the answer, and shows exactly how it fails: it does not know how far away anything is. Everything after that is a better answer to a failure the previous answer left standing, and the failures arrive in a fixed order, which is why the track has the parts it has.
- Distance is missing. One pixel fixes a direction and nothing else, and a second view returns the distance by intersecting two rays (lessons 1–2).
- The cameras are not where we think. The intersection needs the relative placement of the cameras, which lives on a curved space and must be recovered from scans or from the images themselves (lessons 3–5).
- Points are not a scene. The depths form a cloud with holes; a surface answers questions between the points, and a forward model that renders it is what turns the exam from a slogan into a number (lessons 6–7).
- Visibility is a cliff. A surface is either seen or hidden, and a cliff has no gradient, so photographs cannot teach geometry until visibility is made soft; then a function of position holds the scene, and making it fast changes the representation again (lessons 8–10).
- The views run out. Where the photographs say nothing the answer comes from outside the data: a prior, a network that respects the symmetry of the data, and finally a sample instead of an average (lessons 11–13).
- The world will not hold still. Time breaks the assumption every earlier lesson made, and the exam changes from what does it look like from here? to what will it look like next? (lesson 14).
Part I · Recovering depth (lessons 01–02)
What a camera keeps and what it destroys, and the first way to get the lost number back.
Part II · Placing the cameras (lessons 03–05)
Rigid motion on a curved space, then pose by agreement: from scans, and from images alone.
Part III · From points to pictures (lessons 06–07)
A surface you can query, a forward model that predicts pixels, and the exam that grades both.
Part IV · Fitting a scene by gradient (lessons 08–10)
Make visibility soft so photographs can supervise geometry; store the scene as a field; make it fast.
Part V · Learning across scenes (lessons 11–13)
When views run out, priors; the networks that carry them; sampling instead of averaging.
Part VI · The world moves (lesson 14)
Add time, and hand the problem to world models.
The whole derivation on one page
Read the second column of any row and the fourth column of the row above it: they are the same sentence. That is the linear thinking made visible. Each lesson's problem it inherits is exactly the previous lesson's problem it creates, and the middle column is the single move that resolves the first and manufactures the second. The first row inherits its problem from the Computer Vision track; the last row hands its problem to World Models, from first principles.
| # | The problem it inherits | The one move | The problem it creates |
|---|---|---|---|
| Part I · Recovering depth | |||
| 01 | The Computer Vision track taught the camera as a function: a 3D point goes in, a pixel comes out. Vision has to run it backwards — the pixels are what we have and the scene is what we want. What does one pixel tell us about the point that made it, and what does it leave out? | A pixel is a ray: Read the camera as a function from rays to colors: a pixel fixes a direction and no distance, and parallax turns the missing distance into the shift between two views. | A pixel pins down a direction and nothing else: every point along its ray lands on the same pixel, so depth is the one number per pixel that the image destroys. It is also exactly the number a new viewpoint needs, because moving the camera sideways by b shifts a pixel by f·b/Z. So the shift between two views measures depth. How precisely can two rays pin down a point, and what limits that precision? |
| 02 | A pixel pins down a direction and nothing else: every point along its ray lands on the same pixel, so depth is the one number per pixel that the image destroys. It is also exactly the number a new viewpoint needs, because moving the camera sideways by b shifts a pixel by f·b/Z. So the shift between two views measures depth. How precisely can two rays pin down a point, and what limits that precision? | A second ray pins the point: Intersect two rays: disparity measures depth, the error grows as Z², the search is one-dimensional, and active sensors measure range directly. | Two rays pin a point only because we were told where the second camera is. In practice that is the unknown: a robot, a phone or a scanner moves between measurements, and every depth map lives in its own sensor frame. To compare or merge frames we must write down a rotation and a translation, and a rotation cannot be added, averaged or stepped along like a vector. How do we describe a rigid motion so that a computer can estimate it? |
| Part II · Placing the cameras | |||
| 03 | Two rays pin a point only because we were told where the second camera is. In practice that is the unknown: a robot, a phone or a scanner moves between measurements, and every depth map lives in its own sensor frame. To compare or merge frames we must write down a rotation and a translation, and a rotation cannot be added, averaged or stepped along like a vector. How do we describe a rigid motion so that a computer can estimate it? | Where is the camera? Rigid motion on a curved space: A pose is a rotation and a translation; rotations are a curved space, so take steps in the tangent space and map them back with the exponential. | A pose is now a rotation and a translation, and a small correction is a rotation vector applied through the exponential map, a step that never leaves the space of rotations. That says how to move, but not in which direction. Two scans of the same room, taken from unknown poses, must coincide where they overlap. How do we turn "the scans should coincide" into an equation we can solve for the pose? |
| 04 | A pose is now a rotation and a translation, and a small correction is a rotation vector applied through the exponential map, a step that never leaves the space of rotations. That says how to move, but not in which direction. Two scans of the same room, taken from unknown poses, must coincide where they overlap. How do we turn "the scans should coincide" into an equation we can solve for the pose? | Pose by agreement I: aligning scans: The right pose makes two scans coincide: closed form when matches are known (Procrustes), a match-then-solve loop when they are not (ICP). | Aligning scans has a closed-form answer when the matches are known, and a two-step loop when they are not: match each point to its nearest neighbour, solve, repeat. The loop converges to the nearest consistent answer if it starts close enough. All of it needs a sensor that measures 3D directly. A phone measures only pixels, so there is no scan to align. How can poses be recovered from images alone, and what plays the role of "the scans should coincide"? |
| 05 | Aligning scans has a closed-form answer when the matches are known, and a two-step loop when they are not: match each point to its nearest neighbour, solve, repeat. The loop converges to the nearest consistent answer if it starts close enough. All of it needs a sensor that measures 3D directly. A phone measures only pixels, so there is no scan to align. How can poses be recovered from images alone, and what plays the role of "the scans should coincide"? | Pose by agreement II: images alone: With no depth sensor the rays themselves must meet: the epipolar constraint gives relative pose up to scale, and bundle adjustment polishes poses and points together. | Images alone give camera poses and a sparse set of 3D points, up to one global scale, by adjusting both until the points reproject onto the pixels. That is the first model we fitted by comparing its prediction with the data. Now test it on the job it is for: predict a photograph from a camera we did not use. Points show through one another and nothing exists between them. What must a model of geometry provide to answer "what does this ray hit?", and how is it built from noisy depth? |
| Part III · From points to pictures | |||
| 06 | Images alone give camera poses and a sparse set of 3D points, up to one global scale, by adjusting both until the points reproject onto the pixels. That is the first model we fitted by comparing its prediction with the data. Now test it on the job it is for: predict a photograph from a camera we did not use. Points show through one another and nothing exists between them. What must a model of geometry provide to answer "what does this ray hit?", and how is it built from noisy depth? | From points to a surface: Fuse noisy depth maps by averaging in a signed distance field, then read the surface off its zero set (and a mesh off marching cubes). | Fusing noisy depth into a signed distance field gives a surface that answers "where is the first hit along this ray?", with a normal and an inside and an outside. But the held-out photograph is made of colors, not positions, and so far nothing has a color. What does a pixel actually measure, and how do we write down, and score, the image a model predicts? |
| 07 | Fusing noisy depth into a signed distance field gives a surface that answers "where is the first hit along this ray?", with a normal and an inside and an outside. But the held-out photograph is made of colors, not positions, and so far nothing has a color. What does a pixel actually measure, and how do we write down, and score, the image a model predicts? | What a pixel measures: the forward model and the exam: Write the forward model (first hit, then Lambert shading), score it on held-out cameras, and find the wall: visibility is a cliff in the geometry. | We can render a model and score it against a photograph, but we cannot yet improve it by gradient: where one surface hides another, a pixel is a cliff in the geometry, so the loss is flat almost everywhere and says nothing about which way to move. What would a renderer look like whose pixels change smoothly with the scene, and still reduce to the hard one when the scene is solid? |
| Part IV · Fitting a scene by gradient | |||
| 08 | We can render a model and score it against a photograph, but we cannot yet improve it by gradient: where one surface hides another, a pixel is a cliff in the geometry, so the loss is flat almost everywhere and says nothing about which way to move. What would a renderer look like whose pixels change smoothly with the scene, and still reduce to the hard one when the scene is solid? | Making visibility soft: volume rendering: Replace the first hit with a fog of absorbing, emitting particles: transmittance, alpha compositing, and a hard surface as the sharp limit. | Volume rendering makes every pixel a smooth function of the density and color along its ray, and a solid surface is simply its sharp limit. It still leaves the unknown: a density and a color at every point of space, which gradient descent on photographs must find, in a function class flexible enough for fine detail and smooth enough to train. What should that function be? |
| 09 | Volume rendering makes every pixel a smooth function of the density and color along its ray, and a solid surface is simply its sharp limit. It still leaves the unknown: a density and a color at every point of space, which gradient descent on photographs must find, in a function class flexible enough for fine detail and smooth enough to train. What should that function be? | Radiance fields: a scene as a function: Let a function of position (and direction) hold density and color, and train it by comparing rendered rays with photographs. | A radiance field trained on photographs can reproduce views it never saw, but only after a long training run, and each pixel costs hundreds of evaluations of a large network, almost all of them in empty space. Where is the waste, and what representation spends its effort only where the scene is? |
| 10 | A radiance field trained on photographs can reproduce views it never saw, but only after a long training run, and each pixel costs hundreds of evaluations of a large network, almost all of them in empty space. Where is the waste, and what representation spends its effort only where the scene is? | Speed: grids, hashes and Gaussian splats: Spend parameters only where the scene is: feature grids and hashes, then soft explicit primitives rendered by splatting. | Grids and Gaussians fit one scene in minutes and draw it in real time, but only by being shown that scene from many views. Give them three views, or one, and the unseen part is unconstrained: they score well on the photographs they were fitted to and fail the exam. Where can the missing information come from, if not from more photographs? |
| Part V · Learning across scenes | |||
| 11 | Grids and Gaussians fit one scene in minutes and draw it in real time, but only by being shown that scene from many views. Give them three views, or one, and the unseen part is unconstrained: they score well on the photographs they were fitted to and fail the exam. Where can the missing information come from, if not from more photographs? | When the views run out: learning across scenes: Add information from outside the data: a prior learned over many scenes, used inside geometry, instead of it, or amortized into one forward pass. | The missing information comes from a prior learned over many scenes, and a prior has to live in a network. But a 3D input is not an image: a point cloud is an unordered set, a voxel grid is almost entirely empty, a LiDAR sweep has a different length every time. What must a network respect to learn from 3D data, and what can it then answer? |
| 12 | The missing information comes from a prior learned over many scenes, and a prior has to live in a network. But a 3D input is not an image: a point cloud is an unordered set, a voxel grid is almost entirely empty, a LiDAR sweep has a different length every time. What must a network respect to learn from 3D data, and what can it then answer? | Networks that eat 3D: Respect the symmetry of the data: sets need order-free pooling, volumes need sparsity, driving scenes collapse to a bird's-eye view. | A trained network returns one answer per input, the average of everything compatible with what it saw. That is fine for "where is the car" and wrong for "what is behind the mug": the average of two plausible backs is a back that cannot exist. How do we produce a scene that agrees with the data and is also a valid sample of the kinds of scenes there are? |
| 13 | A trained network returns one answer per input, the average of everything compatible with what it saw. That is fine for "where is the car" and wrong for "what is behind the mug": the average of two plausible backs is a back that cannot exist. How do we produce a scene that agrees with the data and is also a valid sample of the kinds of scenes there are? | Sampling the unseen: generative 3D: Sample from the posterior instead of averaging it: distill a 2D prior, generate consistent views, or generate native 3D. | We can now reconstruct a scene, learn what scenes look like, and invent the parts nobody saw. But the scene holds still. In the world things move and so do we, and a model that explains the past cannot say what happens next or what would happen if we pushed. What does it take to represent a scene that changes, and then to predict it? |
| Part VI · The world moves | |||
| 14 | We can now reconstruct a scene, learn what scenes look like, and invent the parts nobody saw. But the scene holds still. In the world things move and so do we, and a model that explains the past cannot say what happens next or what would happen if we pushed. What does it take to represent a scene that changes, and then to predict it? | When the scene moves: time, SLAM, and the road to world models: Add time: a canonical scene plus motion, a camera that moves too, and the hand-off from explaining the past to predicting the future. | A scene that changes can be reconstructed from what has already happened. An agent that has to act needs the other direction: what will happen next, and what would happen if it did something else. Nothing in a reconstruction answers that, because it has no actions and no future. What must a model carry in its head to answer "if I do this, what happens next?", and why would an agent want one at all? |
python3 tools/chain/validate_chain.py --series 3d --widgets --oracles from the repository root: it re-derives each one with an independent script.
Where this track sits
Before it: Computer Vision · 04 Cameras and projection and 05 Multi-view geometry, depth and SLAM give the camera as a function and survey what this track derives; Computer Graphics, from first principles is the same pipeline run forward, from scene to image. Beside it: Computer Vision · 14 Self-supervised learning and 17 Generative vision cover the learned machinery that lessons 11–13 apply to geometry, and Synthetic Vision asks where the training scenes come from. After it: World Models, from first principles begins from the question this track ends on, how to predict what happens next when an agent acts.