Places: a world you can leave and return to
A video model's state is the last few hundred frames it is allowed to look at, so a place it has stopped looking at is gone: walk away, come back, and what it draws is a guess. A camera walks a loop through a yard of eight coloured posts and must say, from memory alone, what it will see a fifth of a lap ahead; a window of 64 frames does no better than a model that never saw the yard. The move is to index memory by where instead of when. The state becomes a pose and a map, written at the place a thing was seen and read at the place the agent is about to look. It costs a pose that drifts, and it cannot say which thing is which.
New idea: the state of a world you can leave and return to is a pose and a map: memory indexed by where, not when, written at the estimated pose and read at the current one. It does not age, it costs a pose that drifts, and it stores what is at a place, not which thing it is.
Forces next: A map remembers where things are, not what they are: it cannot say that this cup is the cup that was on the table a minute ago, or what happens when one thing pushes another. A world made of things needs a state made of things. How should a model represent entities and their interactions?
1 · The revisit test
The yard is a 12 m square with a wall, and eight posts of radius 0.25 m, each its own colour, stand on a ring outside the path. The agent walks a circle of radius 3.4 m counter-clockwise and looks straight ahead (x to the right, y up, angles counter-clockwise from +x). A lap is always the same 21.4 m, taken in n steps, so a larger n is a slower walk and a longer absence from every place. At n = 240 a step is 8.9 cm long and turns 1.5°, and at 0.1 s per step a lap takes 24 s. What the agent sees is a strip of 64 columns across a field of view of 64°, one degree each. A column holds the colour of the first thing its ray hits (grey for the wall) and the distance to it. A frame is those 64 colours, 64 distances and the pose (3 numbers): 131 numbers.
The obvious exam, predict the next frame, is passed by every memory, because consecutive strips differ by a shift of a column or two. An exam that separates memories has to ask about places that have left the view. So we use a revisit test: during lap 2, at 12 moments, each memory gets a look-ahead question, if I walk the next K steps, what will I see? with K = n/5, a turn of 72° and 4.27 m of path at n = 240. The memory answers alone, at the pose the commanded motion leads to. The answer is scored against the strip the agent really sees there by the IoU of the columns that show a post: the columns where answer and truth show the same colour, divided by the columns where either shows a post. Only posts count: the wall is easy. A column the memory has nothing for is filled by a prior, a yard of the same kind drawn afresh with the colours shuffled. That is what a model with no memory would answer, and it scores 0.014: chance.
Give a window of M = 64 frames (lesson 11's W; here a hand-written store of the last M frames, nothing learned) the exact pose, so that nothing drifts yet, and ask on the 240-step lap. For K = 1, 8, 24 and 48 it scores 0.90, 0.74, 0.43 and 0.015, against chance 0.014. It is nearly perfect while the answer lies in what it holds and gone when it does not: at K = 48 the columns it is asked about were last seen 177 steps ago (median; 10th to 90th percentile 157 to 189), far outside 64 frames. Nothing in the window knows the post is there. This is lesson 11's failure measured for places: what has left the last few frames has left the state.
The repair is not better prediction but a memory that does not age. Three constraints follow. What is at a place must survive any absence. It must be found from what the question contains, and a question names a place, not a time. And its size should not grow with the time spent away.
2 · Two reasons a picture changes
A returning view differs from the stored one for two reasons that a window cannot tell apart: the agent has moved, or the world has. Measure how much each explains. Over lap 2, compare the strip at step t with the strip 48 steps later: 29.5 % of the 64 columns have changed colour. Now suppose the world held still. Store everything lap 1 saw as points where the rays landed (§3 says how) and draw that store from the pose the agent has reached. It disagrees with the real strip in only 1.98 % of the columns. The pose alone explains 93.3 % of the change.
So a view is, to a good approximation, a function of two things that change at different rates:
state = (pose, map), viewt = render(map, poset), poset+1 = poset ⊕ motiont
The pose is the camera's position and heading in a fixed world frame, three numbers (x, y, θ); ⊕ composes one step of motion into it, and it changes every step. The map is what the world looks like at each place, and it changes only where the world does. The 1.98 % that the pose does not explain is the floor of this reader. What the world itself does comes on top, and that is what is left for a model to predict: the innovation of lesson 2, the part of the observation the state did not already know. If a hidden hand moves one post by 1.2 m while the agent is away (§6), the unexplained columns rise to 6.96 % and 15.7 % of the columns that show a post are wrong.
Lesson 11's window holds pictures, and a picture does not say where it was taken, so a returning view looks like any new one and nothing links it to the one seen 177 steps before. The windows scored here are kinder, since every frame carries its exact pose, and they still forget, because they are keyed by when. The state has to contain the pose, and memory has to be keyed by it.
3 · Remember by place
What should the memory be made of? The candidates differ in what they are indexed by, and the index decides what happens to an old place.
| Memory | Keeps | Indexed by | Costs, and what goes wrong |
|---|---|---|---|
| window | the last M frames | when | 131 numbers a frame; an absence longer than M is chance |
| leaky state | one vector, faded by ρ each step | when | fixed size; a place fades as ρage |
| keyframes | a frame, if its pose is at least 0.43 (metres plus radians) from every kept one; at most M | where the camera was | 131 numbers each; refuses new places when full; overlapping views stored twice |
| map cells | 5 cm cells, each with one running-mean point and its colour; a new colour writes over the old | where the thing is | 3 numbers a cell; needs the pose; holds a colour, not a thing |
To compare them fairly, one reader answers every query. A stored frame becomes points: each of its 64 rays lands at a world position (x, y) with a colour, using the pose at which the frame was written (the estimated pose; exact for now). To answer a query at pose q, the reader projects the stored points into the camera at q: a point lands in the column its bearing falls in, the nearest point wins the column (a z-buffer), a hole of one or two columns between equal colours is closed, and a column no point reaches stays unknown. The memories differ only in which points they keep. Nothing here is learned: the memories and the reader are written by hand, the depth of every column is supplied, and every score is the look-ahead exam of §1.
At exact poses on the 240-step lap, with M = 64, the window scores 0.015 and holds 8,384 numbers (64 × 131). The keyframes score 0.908 with 60 kept, 7,860 numbers. The map scores 0.894 with 983 cells, 2,949 numbers. The two place-indexed memories sit near the ceiling of this reader (a point lands in one column, so post edges can be a column off). The map does not care how far ahead it is asked: 0.895, 0.894, 0.915 and 0.894 at K = 1, 8, 24 and 48. It does not grow with time either: after lap 2 it holds 983 cells, as after lap 1, because seeing a place again writes over the same cells. The keyframe store needs room for the 60 places it keeps; at M = 32 it is full and refuses the rest, and scores 0.550.
How long must a window be to match? A question looks K = n/5 steps ahead. The oldest thing it can ask about is a place that comes into view then and was last seen on the previous lap: n − K = 0.8 n steps ago (the oldest asked about at n = 240 is 192), so a window must reach back that far. The smallest M (a multiple of 8) at which it reaches 80 % of the map's score is 56, 96, 192, 376 and 752 frames for laps of 60, 120, 240, 480 and 960 steps: about 0.8 n. At n = 960 that is 752 × 131 = 98,512 numbers. A cell, in contrast, is written once however many rays land in it, so the map grows with the surface seen (a cell is 5 cm of wall or post), not with the time spent: 917, 963, 983, 995 and 1,007 cells at the same five laps, against 960 for the 48 m of wall alone. At n = 960 that is 3,021 numbers, and the window needs 32.6 times as many. A window's size grows with the absence it covers; a map's with the surface there is to remember.
4 · The price of where: the pose drifts
Everything so far used the exact pose. The agent does not have it. It has its own motion: at each step an odometer reports the distance moved and the angle turned, with small independent errors (at noise scale 1, 1 cm and 0.001 rad = 0.057° per step), and the pose estimate is the sum of the reports. This is dead reckoning, and it is the error recursion of lesson 6 with nothing in it to pull the error back: eN+1 = eN + δN for the new step's error δ. Independent errors add in variance, so the heading error after N steps has standard deviation σθ√N: 0.888° at N = 240 by the formula, 0.888° measured over 3,000 simulated walks. The position error grows about the same way, 1.10 cm·√N to within 5 %: 8.1, 12.1, 17.0, 24.3 and 34.6 cm after 60, 120, 240, 480 and 960 steps. A lap of 240 steps ends 17 cm off, 0.80 % of the 21.4 m walked, and a longer absence is more steps and more drift.
A pose that is wrong writes the same post at slightly different places on each pass, and the map blurs; the question, too, is asked at the believed pose. With a window of 256 frames (long enough to reach back a lap), noise scale 1 gives the map 3,518 cells where the exact walk needs 983, and its score falls from 0.894 to 0.689; at scales 2, 4 and 8 it is 0.456, 0.165 and 0.038. The window with the same pose pays the same price: 0.697 at scale 1 against the map's 0.689. Neither is at fault. Both place what they hold by the believed pose, so both inherit its error. That is what the idea costs.
Dead reckoning cannot notice that it has come back; a place can. A post of a known colour should stand where it stood. Loop closure uses this. The agent keeps a registry of the posts it has seen, each with a colour and an estimated position, frozen once the post leaves the view. When a post comes fully into view it is matched to the nearest frozen entry of the same colour within 1 m. The matched pairs (where I now put the post, where the registry says it is) give the rigid motion that best lines them up, a rotation and a translation by least squares (the alignment of 3D lesson 4), and the pose is moved by it. The rotation is fitted only if two matched posts are at least 1 m apart, because one point cannot fix an angle.
At noise 1 closure cuts the pose error from 16.0 to 6.7 cm and lifts the map from 0.689 to 0.814. Over 40 worlds (posts drawn afresh) it lifts the mean from 0.698 to 0.784 and wins in 35 of them. It repairs the pose from then on, not the lap-1 map, which was written with a drifting pose. On twelve worlds, with an exact first lap closure scores 0.873 against a ceiling of 0.886; with a drifting first lap, 0.797. And it degrades as the drift grows: 0.814, 0.725, 0.580 and 0.390 at scales 1, 2, 4 and 8, where the open-loop pose error (123 cm at scale 8) is past its 1 m matching gate.
What to try. Start at the defaults: M = 64, n = 240, exact poses. The window scores 0.015 (chance 0.014), the keyframes 0.908 and the map 0.894. Move M to 256: the window rises to 0.887, at 33,536 numbers, and the other two do not move. Now set n = 480: the window falls to 0.080; the keyframes (0.908) and the map (0.894) do not. Set the noise to ×1 (n = 240, M = 256): the map blurs to 0.689 with 3,518 cells and the pose is off by 16.0 cm; the window falls with it (0.697). Turn loop closure on at ×1: 0.814, pose error 6.7 cm; at ×4, the map's 0.165 becomes 0.580 (pose 63.0 to 26.5 cm); at ×8 it reaches only 0.390.
5 · What real systems report
| System | What its authors report about memory |
|---|---|
| Genie (Bruce et al., 2024) | 16 frames of memory, about 1 frame per second |
| Genie 2 (DeepMind, 2024) | consistent worlds for up to a minute, most examples 10–20 s; remembers regions that left the view |
| Genie 3 (DeepMind, 2025) | 24 frames per second at 720p; consistent for a few minutes, with visual memory extending as far back as one minute |
| RTFM (World Labs, 2025) | posed frames act as a spatial memory: "context juggling" retrieves the stored frames near the pose asked about, so memory is not bounded by an ever-growing context; no explicit 3D representation |
| Marble (World Labs, 2025) | worlds can be exported as Gaussian splats, meshes or videos |
Read against §1 and §3. Genie's memory is 16 frames, and lesson 11 listed windows of seconds (WHAM, Oasis, Dreamer 4) and the limited memory of DIAMOND on CS:GO. Genie 2 and 3 state their horizons in seconds and minutes, the quantity §1 measured in steps: a minute at 24 frames per second is 1,440 frames. RTFM resembles the keyframe store of §3 inside a learned renderer: posed frames kept as memory, the ones near the pose asked about retrieved. Marble exports an explicit 3D world, the road above. These are the authors' statements about their systems; the revisit test of §1 is our own measurement, on hand-written memories.
6 · What a map cannot say
A cell holds a place and a colour. Two things follow that a map cannot express, and each runs in the widget.
Which is which. Put two posts of one colour d = 1.2 m apart (posts: two look-alikes). With exact poses nothing is lost: the map scores 0.896 against 0.894 with distinct colours, because each post is where it is. Trouble starts when the pose is wrong. A returning post must be matched to a stored one, colour cannot decide, so position must. Let each sighting land at its stored place plus an independent error of standard deviation s along the line between the posts: A′ = A + eA, B′ = B + eB, B = A + d. The best position can do is keep the order, right exactly when eA − eB < d. That difference has standard deviation s√2, so, with Φ the standard normal distribution function,
P(right) = Φ( d / (s√2) )
| s / d | 0.25 | 0.5 | 1 | 2 | 4 | 8 |
|---|---|---|---|---|---|---|
| P(right), % | 99.8 | 92.1 | 76.0 | 63.8 | 57.0 | 53.5 |
It falls from certainty to a coin toss once the error passes the gap (a simulation in the plane agrees within 1.5 points: 74.7 % at s = d). No amount of appearance memory changes this: the two look the same, and what could separate them is their history, which a map does not keep. In the walk the registry of §4 does the matching. Over 16 worlds on the 240-step lap it pairs the right twin 100 % of the time at noise ×1 (1,720 pairings) and 86.9 % at ×4 (1,719); on the 480-step lap, with more drift, 75.0 % at ×4 (3,089). A wrong pairing is not a wrong label: it moves the pose. At ×2 on the 480-step lap the mean pose error is 39.0 cm with twins and 28.6 cm with distinct posts. One yard shows how bad it can get, the widget's own, with look-alikes, ×1, closure on and n = 480: the registry takes the second twin for the first, pairs right 29.6 % of 213 times and ends with a pose error of 76.3 cm, against 21.7 cm with no closure at all.
What happened. Set while the agent is away to the push: after lap 1 a hidden hand moves post 4 by 1.2 m. With exact poses and M = 256 the window, keyframes and map score 0.642, 0.655 and 0.646, against 0.887, 0.908 and 0.894 unpushed. The arena shows why; the red ring marks where the post was. After lap 2 the map holds 19 cells of that post's colour at the old place and 18 at the new one: the post was written again where it now stands and never erased where it was. A map that carves free space along every ray would erase the ghost, and that is a better map. It would still hold a post at the new place with no link to the old one. Nothing in it says that the two are one post, that it moved, what moved it, or what it would do on touching another.
Both failures have one cause. The map's unit is a cell at a place. A world with things in it has units that persist when they move, carry a pose of their own, and meet. A state for such a world has an entry per thing, not per place.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
A pose and a map let a world model stay consistent about places it has stopped looking at: on the look-ahead test a window of 64 frames scores 0.015 and the map 0.894, with 983 cells at n = 240 and 1,007 at 960. The price, a pose that drifts, takes the map to 0.689 at noise ×1, and loop closure brings it back to 0.814. But the map holds appearance at places and nothing else. Two posts of one colour 1.2 m apart are told apart by position alone, right with probability 76.0 % when the pose error equals the gap; closure pairs them right 86.9 % of the time at noise ×4 on the 240-step lap and 75.0 % on the 480-step lap. When a hidden hand pushed one post, the map kept 19 cells at the old place, gained 18 at the new one, and scored 0.646 against 0.894. It cannot say that the post at the new place is the post that stood at the old one, or what moved it. How should a model represent entities and their interactions?
Interview prompts
- Why does a finite window forget a place, and how long must it be to remember a lap? (§1, §3 — whatever is older than M steps is gone; a lap needs about 0.8 n frames, so its size grows with the absence.)
- Why is "predict the next frame" the wrong exam for memory? (§1 — consecutive views overlap, so every memory passes; the exam must ask about a place that has left the view.)
- What changes between two views of one place, and what does the pose explain? (§2 — the agent's motion and the world's change; the pose explains most of the changed columns, and the rest is the innovation.)
- What is the difference between indexing memory by when and by where? (§3 — by when it ages out and grows with the absence; by where its size is set by the surface seen and an old place reads as well as a new one.)
- How fast does a dead-reckoned pose drift, and why? (§4 — independent step errors add in variance, so the error grows as √N; nothing pulls it back.)
- What does loop closure repair, and what can it not? (§4 — it re-anchors the pose on known posts; it cannot straighten a map written with a drifting pose, and it degrades as drift outgrows its match gate.)
- Why can a map not tell two look-alike objects apart? (§6 — it stores colour at a place, so only position can pair them: right with probability Φ(d/(s√2)).)
- What does a place-indexed memory do when something is pushed? (§6 — it keeps a ghost at the old place and a new entry at the new one, with nothing linking them.)
Companion reads: Lesson 25 · Horizon, memory, and the quadratic bill (what a long window costs), 3D Vision · 05 Pose by agreement II: images alone (the pose from pictures instead of an odometer), 3D Vision · 14 When the scene moves (SLAM, and what breaks when the scene is not static) and Computer Vision · 05 Multi-view, depth and SLAM (the survey of mapping and loop closure).