all_lessons/3D Vision/index14 lessons · ~6h read

3D Vision, from first principles

A camera turns a three-dimensional world into a flat array of numbers. This track runs that process backwards, and grades every step by one exam: show a model some photographs of a scene, then ask it for the photograph from a camera that took none of them. A pixel is a ray that forgot its distance; the exam is what makes the lost number matter; and each of the fourteen lessons is the move forced by the way the previous one failed the exam. From triangulation and pose, through surfaces, fog, radiance fields and splats, to priors, sampling and a scene that moves, the whole subject comes out of one question.

The seed question
You have three photographs of a room. What must you know, and what must you assume, to say what a fourth camera, standing somewhere nobody photographed from, would see?
Who this is for
An engineer or student who is comfortable with linear algebra, calculus and basic probability and wants to understand reconstruction, NeRF and Gaussian splatting rather than call them. The track picks up where the camera-geometry lessons of Computer Vision stop (a 3D point goes in, a pixel comes out); nothing else is assumed, and every term is defined where it first appears. By the end you can say why a pixel cannot give depth and what does; how a camera's pose is recovered when no one tells you; why a point cloud fails the exam and a surface does not; where volume rendering comes from and what a radiance field is a solution to; why speed forces grids and splats; what a learned prior adds and what it costs; and why the whole subject ends by handing its problem to world models.
How this track is built
An original first-principles track, not a survey and not a translation of a textbook. Four things set it apart. (1) A relay, not a list. Every lesson opens with the problem the previous one left standing and closes by stating the next; that hand-off is one sentence, the baton, which appears word for word at the end of one lesson, at the start of the next and in the table below, and a script diffs them. (2) One exam. Every representation enters because the one before it failed the same test in a specific, measurable way, so "better" always means a number on the same ruler. (3) The mechanism runs. The lessons live in Flatland, a 2D world photographed by 1D cameras, where a widget can run the real algorithm (triangulation, ICP, bundle adjustment, volume rendering, a trained field) in a fraction of a second; where the extra dimension changes the answer, a lesson steps up to true 3D. (4) Every number is recomputed. Each quoted number is tagged, and an independent script re-derives it from scratch; facts about real systems come only from a verified list and are cited by name and year.

The exam

Hold one camera out. Give the model the photographs from the others and the cameras' positions, and ask for the photograph the held-out camera took. Score it by PSNR, the logarithm of the mean squared pixel error against the true photograph (higher is better). Lesson 1 sets the exam up with the simplest model there is, a single photograph reused as the answer, and shows exactly how it fails: it does not know how far away anything is. Everything after that is a better answer to a failure the previous answer left standing, and the failures arrive in a fixed order, which is why the track has the parts it has.

Part I · Recovering depth (lessons 01–02)

What a camera keeps and what it destroys, and the first way to get the lost number back.

01
A pixel is a ray
Projection is many-to-one, so one image cannot say how far; but moving the camera shifts each pixel by f·b/Z, so the exam "predict the view from a new camera" is exactly the depth problem.
02
A second ray pins the point
Triangulation and stereo, the Z² error law and the baseline trade-off, epipolar search, ToF and LiDAR error models, back-projection from a depth map to a point cloud.

Part II · Placing the cameras (lessons 03–05)

Rigid motion on a curved space, then pose by agreement: from scans, and from images alone.

03
Where is the camera? Rigid motion on a curved space
SO(3) and SE(3), why rotation matrices cannot be averaged and Euler angles lock, Rodrigues, exp and log, quaternions, and gradient steps that never leave the space.
04
Pose by agreement I: aligning scans
Kabsch/Procrustes with the reflection fix, ICP and its basin of attraction, point-to-plane, degeneracy, robust loss, and the first appearance of the explain-compare-update loop.
05
Pose by agreement II: images alone
The essential matrix and the 8-point algorithm, the scale ambiguity, cheirality, RANSAC, bundle adjustment as the loop with a pinhole renderer, gauge freedom, and how SLAM is the same loop online.

Part III · From points to pictures (lessons 06–07)

A surface you can query, a forward model that predicts pixels, and the exam that grades both.

06
From points to a surface
Why points fail the exam, what a surface must provide, TSDF fusion and its 1/sqrt(n) noise law, normals and sphere tracing, marching cubes, explicit versus implicit trade-offs.
07
What a pixel measures: the forward model and the exam
Radiance along a ray, Lambertian color and why photometric matching works, visibility by rasterizing or ray casting, PSNR as the exam, and the boundary term that autodiff misses.

Part IV · Fitting a scene by gradient (lessons 08–10)

Make visibility soft so photographs can supervise geometry; store the scene as a field; make it fast.

08
Making visibility soft: volume rendering
Derive T(t) = exp(-∫σ) and C = ∫Tσc, discretize to alpha compositing, take the hard limit, differentiate it, and see long-range gradients.
09
Radiance fields: a scene as a function
The field, view dependence, positional encoding and spectral bias, sampling, bundle adjustment with a dense renderer, and why few views still ghost.
10
Speed: grids, hashes and Gaussian splats
Where NeRF's time goes, feature grids and hash tables, 3D Gaussians and their projection, sorted alpha compositing, density control, and what splats cost.

Part V · Learning across scenes (lessons 11–13)

When views run out, priors; the networks that carry them; sampling instead of averaging.

11
When the views run out: learning across scenes
Null spaces and ill-posedness, MAP and posterior uncertainty, hand-made versus learned priors, monocular depth and its scale ambiguity, learned matching, feed-forward reconstruction.
12
Networks that eat 3D
Permutation invariance and PointNet, locality, sparse convolution and its cost, bird's-eye view, detection heads, and what a regression network cannot answer.
13
Sampling the unseen: generative 3D
Why the mean is not a valid object, mixture posteriors, score distillation and its failure modes, multiview generation, 3D latent diffusion, feed-forward generation.

Part VI · The world moves (lesson 14)

Add time, and hand the problem to world models.

14
When the scene moves: time, SLAM, and the road to world models
Dynamic scenes and their ambiguity, deformation fields, per-primitive trajectories, SLAM as the same loop online, and why reconstruction has no actions and no future.

The whole derivation on one page

Read the second column of any row and the fourth column of the row above it: they are the same sentence. That is the linear thinking made visible. Each lesson's problem it inherits is exactly the previous lesson's problem it creates, and the middle column is the single move that resolves the first and manufactures the second. The first row inherits its problem from the Computer Vision track; the last row hands its problem to World Models, from first principles.

#The problem it inheritsThe one moveThe problem it creates
Part I · Recovering depth
01The Computer Vision track taught the camera as a function: a 3D point goes in, a pixel comes out. Vision has to run it backwards — the pixels are what we have and the scene is what we want. What does one pixel tell us about the point that made it, and what does it leave out?A pixel is a ray: Read the camera as a function from rays to colors: a pixel fixes a direction and no distance, and parallax turns the missing distance into the shift between two views.A pixel pins down a direction and nothing else: every point along its ray lands on the same pixel, so depth is the one number per pixel that the image destroys. It is also exactly the number a new viewpoint needs, because moving the camera sideways by b shifts a pixel by f·b/Z. So the shift between two views measures depth. How precisely can two rays pin down a point, and what limits that precision?
02A pixel pins down a direction and nothing else: every point along its ray lands on the same pixel, so depth is the one number per pixel that the image destroys. It is also exactly the number a new viewpoint needs, because moving the camera sideways by b shifts a pixel by f·b/Z. So the shift between two views measures depth. How precisely can two rays pin down a point, and what limits that precision?A second ray pins the point: Intersect two rays: disparity measures depth, the error grows as Z², the search is one-dimensional, and active sensors measure range directly.Two rays pin a point only because we were told where the second camera is. In practice that is the unknown: a robot, a phone or a scanner moves between measurements, and every depth map lives in its own sensor frame. To compare or merge frames we must write down a rotation and a translation, and a rotation cannot be added, averaged or stepped along like a vector. How do we describe a rigid motion so that a computer can estimate it?
Part II · Placing the cameras
03Two rays pin a point only because we were told where the second camera is. In practice that is the unknown: a robot, a phone or a scanner moves between measurements, and every depth map lives in its own sensor frame. To compare or merge frames we must write down a rotation and a translation, and a rotation cannot be added, averaged or stepped along like a vector. How do we describe a rigid motion so that a computer can estimate it?Where is the camera? Rigid motion on a curved space: A pose is a rotation and a translation; rotations are a curved space, so take steps in the tangent space and map them back with the exponential.A pose is now a rotation and a translation, and a small correction is a rotation vector applied through the exponential map, a step that never leaves the space of rotations. That says how to move, but not in which direction. Two scans of the same room, taken from unknown poses, must coincide where they overlap. How do we turn "the scans should coincide" into an equation we can solve for the pose?
04A pose is now a rotation and a translation, and a small correction is a rotation vector applied through the exponential map, a step that never leaves the space of rotations. That says how to move, but not in which direction. Two scans of the same room, taken from unknown poses, must coincide where they overlap. How do we turn "the scans should coincide" into an equation we can solve for the pose?Pose by agreement I: aligning scans: The right pose makes two scans coincide: closed form when matches are known (Procrustes), a match-then-solve loop when they are not (ICP).Aligning scans has a closed-form answer when the matches are known, and a two-step loop when they are not: match each point to its nearest neighbour, solve, repeat. The loop converges to the nearest consistent answer if it starts close enough. All of it needs a sensor that measures 3D directly. A phone measures only pixels, so there is no scan to align. How can poses be recovered from images alone, and what plays the role of "the scans should coincide"?
05Aligning scans has a closed-form answer when the matches are known, and a two-step loop when they are not: match each point to its nearest neighbour, solve, repeat. The loop converges to the nearest consistent answer if it starts close enough. All of it needs a sensor that measures 3D directly. A phone measures only pixels, so there is no scan to align. How can poses be recovered from images alone, and what plays the role of "the scans should coincide"?Pose by agreement II: images alone: With no depth sensor the rays themselves must meet: the epipolar constraint gives relative pose up to scale, and bundle adjustment polishes poses and points together.Images alone give camera poses and a sparse set of 3D points, up to one global scale, by adjusting both until the points reproject onto the pixels. That is the first model we fitted by comparing its prediction with the data. Now test it on the job it is for: predict a photograph from a camera we did not use. Points show through one another and nothing exists between them. What must a model of geometry provide to answer "what does this ray hit?", and how is it built from noisy depth?
Part III · From points to pictures
06Images alone give camera poses and a sparse set of 3D points, up to one global scale, by adjusting both until the points reproject onto the pixels. That is the first model we fitted by comparing its prediction with the data. Now test it on the job it is for: predict a photograph from a camera we did not use. Points show through one another and nothing exists between them. What must a model of geometry provide to answer "what does this ray hit?", and how is it built from noisy depth?From points to a surface: Fuse noisy depth maps by averaging in a signed distance field, then read the surface off its zero set (and a mesh off marching cubes).Fusing noisy depth into a signed distance field gives a surface that answers "where is the first hit along this ray?", with a normal and an inside and an outside. But the held-out photograph is made of colors, not positions, and so far nothing has a color. What does a pixel actually measure, and how do we write down, and score, the image a model predicts?
07Fusing noisy depth into a signed distance field gives a surface that answers "where is the first hit along this ray?", with a normal and an inside and an outside. But the held-out photograph is made of colors, not positions, and so far nothing has a color. What does a pixel actually measure, and how do we write down, and score, the image a model predicts?What a pixel measures: the forward model and the exam: Write the forward model (first hit, then Lambert shading), score it on held-out cameras, and find the wall: visibility is a cliff in the geometry.We can render a model and score it against a photograph, but we cannot yet improve it by gradient: where one surface hides another, a pixel is a cliff in the geometry, so the loss is flat almost everywhere and says nothing about which way to move. What would a renderer look like whose pixels change smoothly with the scene, and still reduce to the hard one when the scene is solid?
Part IV · Fitting a scene by gradient
08We can render a model and score it against a photograph, but we cannot yet improve it by gradient: where one surface hides another, a pixel is a cliff in the geometry, so the loss is flat almost everywhere and says nothing about which way to move. What would a renderer look like whose pixels change smoothly with the scene, and still reduce to the hard one when the scene is solid?Making visibility soft: volume rendering: Replace the first hit with a fog of absorbing, emitting particles: transmittance, alpha compositing, and a hard surface as the sharp limit.Volume rendering makes every pixel a smooth function of the density and color along its ray, and a solid surface is simply its sharp limit. It still leaves the unknown: a density and a color at every point of space, which gradient descent on photographs must find, in a function class flexible enough for fine detail and smooth enough to train. What should that function be?
09Volume rendering makes every pixel a smooth function of the density and color along its ray, and a solid surface is simply its sharp limit. It still leaves the unknown: a density and a color at every point of space, which gradient descent on photographs must find, in a function class flexible enough for fine detail and smooth enough to train. What should that function be?Radiance fields: a scene as a function: Let a function of position (and direction) hold density and color, and train it by comparing rendered rays with photographs.A radiance field trained on photographs can reproduce views it never saw, but only after a long training run, and each pixel costs hundreds of evaluations of a large network, almost all of them in empty space. Where is the waste, and what representation spends its effort only where the scene is?
10A radiance field trained on photographs can reproduce views it never saw, but only after a long training run, and each pixel costs hundreds of evaluations of a large network, almost all of them in empty space. Where is the waste, and what representation spends its effort only where the scene is?Speed: grids, hashes and Gaussian splats: Spend parameters only where the scene is: feature grids and hashes, then soft explicit primitives rendered by splatting.Grids and Gaussians fit one scene in minutes and draw it in real time, but only by being shown that scene from many views. Give them three views, or one, and the unseen part is unconstrained: they score well on the photographs they were fitted to and fail the exam. Where can the missing information come from, if not from more photographs?
Part V · Learning across scenes
11Grids and Gaussians fit one scene in minutes and draw it in real time, but only by being shown that scene from many views. Give them three views, or one, and the unseen part is unconstrained: they score well on the photographs they were fitted to and fail the exam. Where can the missing information come from, if not from more photographs?When the views run out: learning across scenes: Add information from outside the data: a prior learned over many scenes, used inside geometry, instead of it, or amortized into one forward pass.The missing information comes from a prior learned over many scenes, and a prior has to live in a network. But a 3D input is not an image: a point cloud is an unordered set, a voxel grid is almost entirely empty, a LiDAR sweep has a different length every time. What must a network respect to learn from 3D data, and what can it then answer?
12The missing information comes from a prior learned over many scenes, and a prior has to live in a network. But a 3D input is not an image: a point cloud is an unordered set, a voxel grid is almost entirely empty, a LiDAR sweep has a different length every time. What must a network respect to learn from 3D data, and what can it then answer?Networks that eat 3D: Respect the symmetry of the data: sets need order-free pooling, volumes need sparsity, driving scenes collapse to a bird's-eye view.A trained network returns one answer per input, the average of everything compatible with what it saw. That is fine for "where is the car" and wrong for "what is behind the mug": the average of two plausible backs is a back that cannot exist. How do we produce a scene that agrees with the data and is also a valid sample of the kinds of scenes there are?
13A trained network returns one answer per input, the average of everything compatible with what it saw. That is fine for "where is the car" and wrong for "what is behind the mug": the average of two plausible backs is a back that cannot exist. How do we produce a scene that agrees with the data and is also a valid sample of the kinds of scenes there are?Sampling the unseen: generative 3D: Sample from the posterior instead of averaging it: distill a 2D prior, generate consistent views, or generate native 3D.We can now reconstruct a scene, learn what scenes look like, and invent the parts nobody saw. But the scene holds still. In the world things move and so do we, and a model that explains the past cannot say what happens next or what would happen if we pushed. What does it take to represent a scene that changes, and then to predict it?
Part VI · The world moves
14We can now reconstruct a scene, learn what scenes look like, and invent the parts nobody saw. But the scene holds still. In the world things move and so do we, and a model that explains the past cannot say what happens next or what would happen if we pushed. What does it take to represent a scene that changes, and then to predict it?When the scene moves: time, SLAM, and the road to world models: Add time: a canonical scene plus motion, a camera that moves too, and the hand-off from explaining the past to predicting the future.A scene that changes can be reconstructed from what has already happened. An agent that has to act needs the other direction: what will happen next, and what would happen if it did something else. Nothing in a reconstruction answers that, because it has no actions and no future. What must a model carry in its head to answer "if I do this, what happens next?", and why would an agent want one at all?
How to read this
Straight through, in order: the track is one argument, and each lesson's "Linear position" box names the problem it inherits, the single idea it adds and the problem it hands on. If you have time only for the spine, read 01 → 02 → 03 → 07 → 08 → 09 → 11 → 14. Each lesson has one widget that runs its mechanism, and a "What to try" paragraph that tells you which control to move first. To re-check every quoted number yourself, run python3 tools/chain/validate_chain.py --series 3d --widgets --oracles from the repository root: it re-derives each one with an independent script.

Where this track sits

Before it: Computer Vision · 04 Cameras and projection and 05 Multi-view geometry, depth and SLAM give the camera as a function and survey what this track derives; Computer Graphics, from first principles is the same pipeline run forward, from scene to image. Beside it: Computer Vision · 14 Self-supervised learning and 17 Generative vision cover the learned machinery that lessons 11–13 apply to geometry, and Synthetic Vision asks where the training scenes come from. After it: World Models, from first principles begins from the question this track ends on, how to predict what happens next when an agent acts.