When the scene moves: time, SLAM, and the road to world models
Thirteen lessons assumed that the scene holds still while the cameras come and go. A video breaks that: each instant is photographed once. This lesson runs lesson 10's splats on a video of a ball rolling past a crate and measures how a static model fails. It then finds that one camera cannot say how fast a moving thing goes, because it cannot say how far away it is: a whole family of worlds makes the same photographs. Sharing what does not change, one canonical scene with a small motion for each moment, shrinks that family from one open scale per frame to one, and a second camera closes it. The camera's own pose is one more such motion, so fitting both is bundle adjustment run along time. The result explains the recorded past well and says nothing reliable about what comes next.
New idea: a scene that changes is a canonical scene, shared by every frame, plus a small motion for each time, and the camera's pose is one more motion of the same kind. The same loop fits all of it; what the photographs cannot decide (how big a moving thing is) is settled by sharing across frames and by a second camera, not by more frames.
Forces next: A scene that changes can be reconstructed from what has already happened. An agent that has to act needs the other direction: what will happen next, and what would happen if it did something else. Nothing in a reconstruction answers that, because it has no actions and no future. What must a model carry in its head to answer "if I do this, what happens next?", and why would an agent want one at all?
1 · A static model, given a video
The scene is a yard seen from above (Flatland as before: x to the right, z up, metres). A crate and a tower stand still. A ball of radius 0.45 m rolls to the right at 1.2 m/s and passes the point (−0.5, 4.4) at the middle of the video. A camera slides to the right at 0.4 m/s, 0.04 m per frame, looking along +z. A frame is a 64-pixel photograph (f = 44 px, as in lesson 1), one every 0.1 s, and the video has 24 frames, 2.4 s. After the 24th the camera stands still and the world carries on; §6 uses those frames.
Stack the frames, one row each, time pointing down: a one-dimensional video becomes a two-dimensional picture, the space-time image in the widget below. A point that stands still draws a straight line whose slope is lesson 1's parallax, f·c/Z pixels per frame when the camera moves c = 0.04 m per frame: 0.45 for the crate's front face (Z = 3.9 m), and about 0.21 for the tower, whose near faces are 8.0 to 8.9 m away. The ball draws a line of its own. It moves relative to the camera at v − c = 0.08 m per frame (v = 0.12 m per frame is its own speed), so f·(v − c)/Z = 0.80 px per frame to the right, and in the window it crosses 19.2 px, more than twice its own width of 9.0 px.
Run lesson 10's loop on it without changing anything but the data. Seed 36 Gaussians at the points the camera saw in the middle frame (no densifying here), treat the 24 frames as 24 views from known cameras, and repeat: explain (render each frame), compare (squared error against its photograph), update (Adam on every Gaussian's nine numbers). The score is lesson 7's PSNR between render and photograph, here on the 24 frames the model was fitted to.
| ball speed | 0 | 0.4 | 0.8 | 1.2 | 1.6 | 2.0 m/s |
|---|---|---|---|---|---|---|
| static scene, score on the 24 frames | 27.3 dB | 21.4 | 17.4 | 14.8 | 13.1 | 12.3 |
With the ball at rest the static model is a good one. As it speeds up the score falls by 15.0 dB, by less and less for each step in speed, and the loss is not spread evenly. At 1.2 m/s the crate is as well fitted as at rest (27.1 dB against 27.2), the pixels that show the ball score 8.4 dB, and the empty background falls from 27.7 to 16.5 dB. A static scene has one ball. It puts it where the compromise over 24 frames lies, so every frame disagrees with the model twice, where the ball is and where the model's ball is.
So the exam has grown a time axis: predict the photograph of a camera at a time that was not among them. A static model has no time axis. Before giving it one, ask what the pixels of one camera can say about motion at all, because the answer decides what the model must look like.
2 · What one camera's pixels can say about motion
A photograph shows where things are relative to the camera at that instant and nothing else. So move the camera and every object by the same vector and the photograph does not change. Render the yard from the sliding camera, then from a camera that stays at its first pose in a yard that slides the other way, 0.04 m per frame: the two videos differ by 0.0000 in every pixel. "The camera moved right" and "the scene moved left" are one statement, and no number of frames separates them. It is lesson 5's gauge, now in time, and it is removed the way every motion estimate removes it: declare what stands still. The crate and the tower are the static reference, and the camera is what moves relative to them.
Fix the camera's path, which the static reference measures (up to the scale of lesson 5; take it as known). Is the ball's motion now determined? A disc of radius r with its centre at (X, Z) in camera coordinates is bounded by two tangent rays, at angles φ ± asin(r/d) from the optical axis, with φ = atan(X/Z) and d = √(X² + Z²). Multiply X, Z and r by the same s and neither φ nor r/d changes: the same two pixels bound the disc. So scale the ball about the camera, at every instant (p the ball's centre, C the camera's):
ps(t) = C(t) + s·(p(t) − C(t)), rs = s·r
whatever s, every photograph is unchanged. The ball's world velocity is not: differentiating, vs = c + s·(v − c), with c the camera's velocity and v the true one. Three members of the family, rendered (the two silhouette edges agree to within 10−12 px; the "largest pixel difference" is the renderer's own ray-marching tolerance, not a property of the family):
| s | ball depth | radius | world speed | largest pixel difference, camera A | edge shift seen by camera C, 0.5 m to the left | PSNR between the two C photographs |
|---|---|---|---|---|---|---|
| 0.5 | 2.2 m | 0.225 m | 0.80 m/s | 0.017 | 5.2 px | 11.9 dB |
| 1 (the truth) | 4.4 m | 0.45 m | 1.20 m/s | 0 | 0 | identical |
| 1.6 | 7.0 m | 0.72 m | 1.68 m/s | 0.016 | 1.9 px | 16.4 dB |
A ball half as far, half as big and rolling at 0.80 m/s is photographed exactly as the 1.2 m/s ball at 4.4 m. One camera cannot say how fast the ball moves because it cannot say how far away it is. Xian et al. (2021) put the general difficulty in words for monocular dynamic scenes: a video holds one observation of the scene at each instant, so motion and appearance can stand in for each other. When the ball rolls against the camera's motion, s = c/(c − v) makes vs zero: a static ball, nearer and smaller, explains the video. Motion has been traded for shape outright (the exercise at the end).
How much is left open? Count it. The ball's two silhouette edges give two numbers per frame. A model that gives the ball its own position and size in every frame has three unknowns against those two, so one scale per frame is free. Let the ball keep one size throughout and every frame's apparent size must agree with it: one scale is left. A velocity law cuts the unknowns from 49 to 5 and leaves that scale alone. The rank of the Jacobian of the 48 edge positions with respect to each model's unknowns, computed numerically:
| model of the ball | unknowns | measured by one camera | rank | left open | measured by two cameras | rank | left open |
|---|---|---|---|---|---|---|---|
| every frame on its own: (xt, zt, rt) | 72 | 48 | 48 | 24 | 96 | 72 | 0 |
| one size, free positions: (xt, zt; r) | 49 | 48 | 48 | 1 | 96 | 49 | 0 |
| one size, constant velocity: (x0, z0, vx, vz; r) | 5 | 48 | 4 | 1 | 96 | 5 | 0 |
Sharing the ball across frames takes the big step, from 24 open numbers to 1, because every frame must now agree about its size. A law of motion reduces what has to be estimated and leaves the scale where it was. The last number is not in one camera's photographs: a second camera at the same instant fixes it by triangulation (lesson 2), and in practice a prior does (lesson 11: the usual sizes of things).
3 · Share what does not change
The count says what to build. What is the same at every time should be stored once, and only what must differ should differ. The same at every time is the content of the scene, its shape, colour and size; what differs is where it is. Write the scene at time t as a canonical scene plus a motion, x(t) = x0 + D(x0, t) for each canonical point x0, or the same thing read backwards, a map from each point at time t to its canonical position. D-NeRF (Pumarola et al., 2021) learns the backward map with two networks, a deformation network and a canonical network that gives density and colour at the canonical position, from a single camera with one view per instant. Nerfies (Park et al., 2021) uses a deformation field in SE(3) with an elastic regulariser, and notes that optimising the radiance field and the field together is under-constrained, which §2 explains. Dynamic 3D Gaussians (Luiten et al., 2024) keep each Gaussian's colour, opacity and size fixed and let position and rotation change at every timestep, with priors of local rigidity and isometry, on 27 training cameras. 4D Gaussian splatting (Wu et al., 2024) keeps one set of Gaussians and a deformation field.
We take lesson 10's Gaussians. Their nine numbers (position, two log-scales, angle, opacity, colour) are the canonical scene, and the moving ones get a motion, one constant velocity V = (Vx, Vz) in metres per frame: xk(t) = xk + V·(t − tref), with tref the middle frame. All the moving Gaussians share it: the ball is one rigid group. The seeds that landed on the ball are marked as moving. The mask is read from the renderer here; in practice it comes from a segmentation model, or from the residual of the static fit, which §1 showed to be large on the ball. Everything else stands still: it is the static reference of §2.
Why a group, and not a velocity for every Gaussian? The same loop with a free velocity for each of the 36 Gaussians (72 numbers instead of 2) scores 26.5 dB, a point better than the group's 25.5, and sets 9 of the 20 Gaussians that belong to the crate and the tower moving faster than 0.1 m/s. The extra freedom is spent on the static scene's imperfections, and the family of §2 applies to every Gaussian. A group has two numbers and no room for that. What each choice costs, for K = 36 Gaussians and T = 24 frames:
| what is stored | numbers | here |
|---|---|---|
| a separate scene for every frame | 9·K·T | 7,776 |
| one scene, a free displacement of the moving group per frame | 9·K + 2·T | 372 |
| one scene, one constant velocity | 9·K + 2 | 326 |
| the video itself, one camera (64 pixels, 3 channels) | 3·64·T | 4,608 |
The first has more unknowns than the video has numbers: it can explain any video and learns nothing about this one. The other two are 12.4 and 14.1 times smaller than the video; they compress it. A displacement per frame has no value after frame 23. A velocity does, and §6 tests it.
The loop is lesson 5's. Explain: place the Gaussians at time t and render frame t. Compare: squared error against the photograph. Update: Adam on the nine numbers of every Gaussian and on V. The velocity's gradient is the chain rule through xk(t) = xk + Vτ, with τ = t − tref: ∂L/∂V = Σt τ Σk moving ∂Lt/∂xk, with Lt the error of frame t, which the per-frame gradients of lesson 10 already contain (checked against a finite difference: relative difference 0.14 %). One practical point: a velocity can only be found from frames in which the moving Gaussians still overlap the real ball, so the loop starts from the 8 frames nearest the middle and widens until it uses all 24, from step 160 of its 300.
4 · The camera moves too
Until now the camera's poses were given. In practice they are unknowns, and unknowns of the same kind. Moving the camera by δ along its slide in frame t gives the photograph of moving every Gaussian by −δ (the gauge of §2, as an equation), so ∂Lt/∂δt = −Σk ∂Lt/∂xk, with xk now the Gaussian's coordinate along the slide and the sum over every Gaussian, the moving group at its place at time t included: the sum of position gradients the loop computes anyway (checked by finite difference: relative difference 0.01 %). Poses join the unknowns at one number per frame, and the loop is bundle adjustment (lesson 5: explain, compare, update) with a photometric residual in place of a reprojection one. The gauge must be fixed, as in lesson 5: here the first and the last pose are known, which fixes where the world is and how big.
The widget can hand the loop poses with jitter in place of the true ones, as a hand-held camera would: 6 cm per frame, 5.5 cm RMS as drawn. With two cameras and the poses taken as given, the held-out camera falls from 25.0 to 21.6 dB: a pose error is a scene error, because the model has to explain photographs that disagree about where the camera stood. Let the loop estimate the poses as well and the score returns to 24.8 dB, with the pose error down to 1.2 cm. With one camera less is recovered: 2.0 cm of the 5.5 are left.
SLAM is this loop run while the frames arrive: the poses and the map of a recent window stay adjustable and older ones are frozen, because a step costs one render for every frame it adjusts, so adjusting everything gets dearer with every frame. Online has two prices. A pose is estimated from the frames so far, and later frames reach it only through the map; and its error is inherited by the map and by every pose estimated after it, so errors accumulate (drift): if each pose inherits an independent 1 cm error from the one before, n frames leave √n cm, 5 cm after 24 and 10 cm after 100, until a place seen before ties the present to the past (loop closure). MegaSaM (Li et al., 2024) builds camera and depth estimation for casual dynamic monocular video on a deep visual SLAM framework, with its own changes to training and inference.
5 · Fit it
The widget runs the loop of §3 and §4 on the yard. Fit takes 300 Adam steps and scores the model on the frames it was fitted to, on a camera it never saw (C, 0.5 m to the left of A), on the instants between the fitted frames, and on the future (§6).
What to try. Start as the page opens (1.2 m/s, one camera, canonical scene + motion) and press fit: the model scores 25.5 dB on the 24 frames and finds 1.14 m/s against the true 1.20. Switch the model to one static scene and fit again: 14.8 dB. Slide the speed to 0 and fit both models (27.3 and 27.5 dB), then to 2.0 m/s (12.3 and 24.3 dB): the static model's score falls by 15.0 dB, the moving model's by 3.3. Now the ambiguity. With one camera, set the depth guess to 0.6 and fit: the training score is still 25.5 dB, the speed settles at 0.82 m/s, and the held-out camera drops from 23.7 to 13.1 dB. At a guess of 1.5 the speed is 1.30 m/s, the training score 25.5 dB again and camera C scores 21.4 dB. The photographs the loop sees cannot tell these fits apart. The camera it never saw tells the 0.6 guess from the right one, and the 1.5 guess only weakly: six other draws of the 5 cm of noise on the starting Gaussians move the right guess's own score anywhere from 18.5 to 23.3 dB. Switch to two cameras and repeat: for guesses of 0.6, 1 and 1.5 the speed is 1.19, 1.20 and 1.19 m/s and camera C scores 24.3, 25.0 and 24.1 dB. Last, the poses: with two cameras and a guess of 1, take 6 cm of jitter as given (21.6 dB on C, pose error 5.5 cm), then estimated (24.8 dB, 1.2 cm). On every fit, read the right-hand half of the plot: that is §6.
6 · Run it forward
A fit is judged on frames it did not see, and a video offers three kinds: a held-out camera at a fitted time (§5); held-out times inside the window, the 23 instants half-way between fitted frames; and the future, frames 24–47, photographed with the camera standing at its last pose. Take the two-camera fit, so that the scale is not in question. Between the fitted frames the model scores 25.6 dB, as on the fitted frames (25.6): interpolating in time is free. Beyond them it is not. Over the next 0.8 s (frames 24–31) the score is 21.6 dB, and 3.0–3.7 s ahead (frames 40–47) it is 10.2 dB.
Two things lose the points, and they differ in kind. The first is ordinary, and more frames cure it: fitted on frames 0–31 instead, the same loop scores 23.8 dB on frames 24–31, not 21.6. The loss sits at the ball. On its pixels and the 2 px around them the score is 16.1 dB in these frames against 21.9 in the window; everywhere else it is 25.3 against 26.5. The second is the drop after frame 32.75. The ball met the crate (at 3.3 s) and rolled back at 0.96 m/s, and the model's ball went on: by frame 47 it has moved 4.24 m from where it was at the middle frame and the real one 1.18 m, 3.06 m apart, through a crate that is in the model.
The model has the crate (Gaussians, at the right place), the ball and its velocity. What it lacks is the rule that turned the ball round: solid things do not pass through one another. The constant velocity was our choice, a description of the window. Any other motion that agrees with 24 frames, a displacement per frame for one, fits as well and has nothing to say beyond frame 23. Nothing in the fit learned a law.
Nor can it be asked what if. Its inputs are a camera and a time, and a push is neither. Suppose the ball is struck at frame 24 and goes twice as fast: it hits the crate at frame 28.4 and bounces back at 1.92 m/s. The model, fitted on the same 24 frames, returns the frames it returned before: they match the world where nothing happens at 21.6 dB over frames 24–31, and the world where the push happens at 14.5 dB. A reconstruction can be re-fitted to the world in which the push happened, once it has happened and been photographed. It cannot say beforehand what the push will do, and an agent that has to choose what to push needs exactly that.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
A scene that moves can now be represented and fitted: one canonical scene, a motion for what moves, a pose for every frame, found by the loop of lesson 5. With one camera it scores 25.5 dB on the 24 frames where a static scene reaches 14.8, and the scale of the moving thing stays outside the photographs until a second camera supplies it. What the fit cannot do is predict. With two cameras it scores 25.6 dB on the frames it was fitted to and on the instants between them, 21.6 dB over the next 0.8 s and 10.2 dB at 3.0–3.7 s ahead, where the ball has met a crate that the model contains and knows nothing about; and when the ball is struck at frame 24 it returns the same frames as if nothing had happened, 14.5 dB from the world that does. It carries a state for every instant of the past, where everything was, and no rule that takes one state to the next, and nowhere to put what an agent does. Genie 3 (DeepMind, 2025) is described as generating worlds one can navigate in real time at 24 frames per second, consistent for a few minutes at 720p, and Marble (World Labs, 2025) as creating 3D worlds from text, images, video or coarse layouts, exportable as Gaussian splats, meshes or videos. They stand where reconstruction and generation meet, and a world that can be navigated has to answer the navigator, an input no reconstruction in this track has. What must a model carry in its head to answer "if I do this, what happens next?", and why would an agent want one at all?
Interview prompts
- Why can a single camera, however many frames it takes, not give the speed of a moving object? (§2 — scaling the object about the camera by s leaves every photograph unchanged and changes its world speed to c + s(v − c); the depth is the missing number.)
- What is the gauge of a dynamic scene, and how is it fixed? (§2, §4 — a photograph holds only relative position, so camera motion and scene motion are one statement; declare the static majority fixed, and fix a frame and a scale with known poses.)
- What does sharing a canonical scene across frames buy, and what does it not? (§2, §3 — it cuts the open scales from one per frame to one in all, and the unknowns from 9KT to 9K plus a few; the last scale stays open.)
- Count the numbers stored by a per-frame reconstruction, a canonical scene with a displacement per frame, and one with a velocity. (§3 — 9KT = 7,776; 9K + 2T = 372; 9K + 2 = 326, against 4,608 numbers in the video.)
- How are camera poses estimated together with a scene, and what makes SLAM different from batch bundle adjustment? (§4 — moving the camera equals moving the scene the other way, so the pose gradient is minus the sum of position gradients; SLAM runs the loop over a window as frames arrive, with drift and loop closure.)
- Where does SLAM drift come from, and what removes it? (§4 — each pose is estimated from the frames so far and its error is inherited by the map and by every later pose, so independent errors add as √n; a loop closure ties the present to a place seen before.)
- A model fits a video at 25 dB and scores 10 dB at three seconds ahead. What follows? (§6 — nothing about its fitted frames; it has a state and a description of the window, no law: here it passes the ball through a crate.)
- Why can a reconstruction not answer "what if I push it"? (§6 — its inputs are a camera and a time; a push is neither, so it returns the same frames with or without it.)
Companion reads: Computer Vision · 05 Multi-view, depth and SLAM (structure from motion and SLAM, which this lesson turns into one loop), Computer Vision · 16 Video and temporal models (motion learned from data), Computer Vision · 11 Keypoints, pose and tracking (the tracking stage of track-then-reconstruct), and Computer Graphics · 13 Animation and motion (how authored motion is parameterised).