all_lessons/World Models/12 · Placeslesson 12 / 31

Places: a world you can leave and return to

A video model's state is the last few hundred frames it is allowed to look at, so a place it has stopped looking at is gone: walk away, come back, and what it draws is a guess. A camera walks a loop through a yard of eight coloured posts and must say, from memory alone, what it will see a fifth of a lap ahead; a window of 64 frames does no better than a model that never saw the yard. The move is to index memory by where instead of when. The state becomes a pose and a map, written at the place a thing was seen and read at the place the agent is about to look. It costs a pose that drifts, and it cannot say which thing is which.

The thesis, here
A window forgets by age: whatever was seen more than M steps ago is gone, however much it still matters. Memory indexed by place does not age, and, with a true pose, its size is set by how much of the world has been seen, not by the time spent. What it asks in return is a pose, and a pose built from motion alone drifts. What it holds is appearance at a place, with no notion of which thing is which.
Linear position
Forced by: A video model predicts the next frames convincingly, but its state is the last few hundred frames it is allowed to look at. Walk away from the sofa and come back and the sofa has changed. A finite window is a finite memory. What would let a world model stay consistent about places it has stopped looking at?
New idea: the state of a world you can leave and return to is a pose and a map: memory indexed by where, not when, written at the estimated pose and read at the current one. It does not age, it costs a pose that drifts, and it stores what is at a place, not which thing it is.
Forces next: A map remembers where things are, not what they are: it cannot say that this cup is the cup that was on the table a minute ago, or what happens when one thing pushes another. A world made of things needs a state made of things. How should a model represent entities and their interactions?
The plan
Six moves. (1) Put a number on the entry failure with a revisit test. (2) Split what changes between two views of one place into the agent's own motion and the world's change. (3) Index memory by place and compare three ways to keep it. (4) Pay for the pose: dead reckoning drifts and blurs the map, and recognising a revisit repairs it. (5) Set the result beside what real systems report. (6) Ask what a map cannot say.

1 · The revisit test

The yard is a 12 m square with a wall, and eight posts of radius 0.25 m, each its own colour, stand on a ring outside the path. The agent walks a circle of radius 3.4 m counter-clockwise and looks straight ahead (x to the right, y up, angles counter-clockwise from +x). A lap is always the same 21.4 m, taken in n steps, so a larger n is a slower walk and a longer absence from every place. At n = 240 a step is 8.9 cm long and turns 1.5°, and at 0.1 s per step a lap takes 24 s. What the agent sees is a strip of 64 columns across a field of view of 64°, one degree each. A column holds the colour of the first thing its ray hits (grey for the wall) and the distance to it. A frame is those 64 colours, 64 distances and the pose (3 numbers): 131 numbers.

The obvious exam, predict the next frame, is passed by every memory, because consecutive strips differ by a shift of a column or two. An exam that separates memories has to ask about places that have left the view. So we use a revisit test: during lap 2, at 12 moments, each memory gets a look-ahead question, if I walk the next K steps, what will I see? with K = n/5, a turn of 72° and 4.27 m of path at n = 240. The memory answers alone, at the pose the commanded motion leads to. The answer is scored against the strip the agent really sees there by the IoU of the columns that show a post: the columns where answer and truth show the same colour, divided by the columns where either shows a post. Only posts count: the wall is easy. A column the memory has nothing for is filled by a prior, a yard of the same kind drawn afresh with the colours shuffled. That is what a model with no memory would answer, and it scores 0.014: chance.

Give a window of M = 64 frames (lesson 11's W; here a hand-written store of the last M frames, nothing learned) the exact pose, so that nothing drifts yet, and ask on the 240-step lap. For K = 1, 8, 24 and 48 it scores 0.90, 0.74, 0.43 and 0.015, against chance 0.014. It is nearly perfect while the answer lies in what it holds and gone when it does not: at K = 48 the columns it is asked about were last seen 177 steps ago (median; 10th to 90th percentile 157 to 189), far outside 64 frames. Nothing in the window knows the post is there. This is lesson 11's failure measured for places: what has left the last few frames has left the state.

The repair is not better prediction but a memory that does not age. Three constraints follow. What is at a place must survive any absence. It must be found from what the question contains, and a question names a place, not a time. And its size should not grow with the time spent away.

2 · Two reasons a picture changes

A returning view differs from the stored one for two reasons that a window cannot tell apart: the agent has moved, or the world has. Measure how much each explains. Over lap 2, compare the strip at step t with the strip 48 steps later: 29.5 % of the 64 columns have changed colour. Now suppose the world held still. Store everything lap 1 saw as points where the rays landed (§3 says how) and draw that store from the pose the agent has reached. It disagrees with the real strip in only 1.98 % of the columns. The pose alone explains 93.3 % of the change.

So a view is, to a good approximation, a function of two things that change at different rates:

state = (pose, map), viewt = render(map, poset), poset+1 = poset ⊕ motiont

The pose is the camera's position and heading in a fixed world frame, three numbers (x, y, θ); ⊕ composes one step of motion into it, and it changes every step. The map is what the world looks like at each place, and it changes only where the world does. The 1.98 % that the pose does not explain is the floor of this reader. What the world itself does comes on top, and that is what is left for a model to predict: the innovation of lesson 2, the part of the observation the state did not already know. If a hidden hand moves one post by 1.2 m while the agent is away (§6), the unexplained columns rise to 6.96 % and 15.7 % of the columns that show a post are wrong.

Lesson 11's window holds pictures, and a picture does not say where it was taken, so a returning view looks like any new one and nothing links it to the one seen 177 steps before. The windows scored here are kinder, since every frame carries its exact pose, and they still forget, because they are keyed by when. The state has to contain the pose, and memory has to be keyed by it.

3 · Remember by place

What should the memory be made of? The candidates differ in what they are indexed by, and the index decides what happens to an old place.

MemoryKeepsIndexed byCosts, and what goes wrong
windowthe last M frameswhen131 numbers a frame; an absence longer than M is chance
leaky stateone vector, faded by ρ each stepwhenfixed size; a place fades as ρage
keyframesa frame, if its pose is at least 0.43 (metres plus radians) from every kept one; at most Mwhere the camera was131 numbers each; refuses new places when full; overlapping views stored twice
map cells5 cm cells, each with one running-mean point and its colour; a new colour writes over the oldwhere the thing is3 numbers a cell; needs the pose; holds a colour, not a thing

To compare them fairly, one reader answers every query. A stored frame becomes points: each of its 64 rays lands at a world position (x, y) with a colour, using the pose at which the frame was written (the estimated pose; exact for now). To answer a query at pose q, the reader projects the stored points into the camera at q: a point lands in the column its bearing falls in, the nearest point wins the column (a z-buffer), a hole of one or two columns between equal colours is closed, and a column no point reaches stays unknown. The memories differ only in which points they keep. Nothing here is learned: the memories and the reader are written by hand, the depth of every column is supplied, and every score is the look-ahead exam of §1.

At exact poses on the 240-step lap, with M = 64, the window scores 0.015 and holds 8,384 numbers (64 × 131). The keyframes score 0.908 with 60 kept, 7,860 numbers. The map scores 0.894 with 983 cells, 2,949 numbers. The two place-indexed memories sit near the ceiling of this reader (a point lands in one column, so post edges can be a column off). The map does not care how far ahead it is asked: 0.895, 0.894, 0.915 and 0.894 at K = 1, 8, 24 and 48. It does not grow with time either: after lap 2 it holds 983 cells, as after lap 1, because seeing a place again writes over the same cells. The keyframe store needs room for the 60 places it keeps; at M = 32 it is full and refuses the rest, and scores 0.550.

How long must a window be to match? A question looks K = n/5 steps ahead. The oldest thing it can ask about is a place that comes into view then and was last seen on the previous lap: n − K = 0.8 n steps ago (the oldest asked about at n = 240 is 192), so a window must reach back that far. The smallest M (a multiple of 8) at which it reaches 80 % of the map's score is 56, 96, 192, 376 and 752 frames for laps of 60, 120, 240, 480 and 960 steps: about 0.8 n. At n = 960 that is 752 × 131 = 98,512 numbers. A cell, in contrast, is written once however many rays land in it, so the map grows with the surface seen (a cell is 5 cm of wall or post), not with the time spent: 917, 963, 983, 995 and 1,007 cells at the same five laps, against 960 for the 48 m of wall alone. At n = 960 that is 3,021 numbers, and the window needs 32.6 times as many. A window's size grows with the absence it covers; a map's with the surface there is to remember.

Road not taken · a longer window
It works exactly as far as it reaches, and the bill grows with n: the 98,512 numbers above and, for a transformer, the attention over every frame in it (lesson 11 counted it). Real systems state horizons of this kind: Genie 3 (DeepMind, 2025) reports visual memory extending as far back as one minute, which at 24 frames per second is 1,440 frames. For a window of that length, a place unseen for longer is at chance (§1).
Road not taken · a fixed-size recurrent state
A leaky state does not grow, but it blends a place with everything that came after it. If it keeps a fraction ρ of what it knew each step, a place last seen 177 steps ago has left ρ177: 0.0046 for ρ = 0.97, 0.169 for 0.99 and 0.492 for 0.996. Half survives only if ρ ≥ 0.99609, and then the state forgets nothing quickly, including what is no longer true.

4 · The price of where: the pose drifts

Everything so far used the exact pose. The agent does not have it. It has its own motion: at each step an odometer reports the distance moved and the angle turned, with small independent errors (at noise scale 1, 1 cm and 0.001 rad = 0.057° per step), and the pose estimate is the sum of the reports. This is dead reckoning, and it is the error recursion of lesson 6 with nothing in it to pull the error back: eN+1 = eN + δN for the new step's error δ. Independent errors add in variance, so the heading error after N steps has standard deviation σθ√N: 0.888° at N = 240 by the formula, 0.888° measured over 3,000 simulated walks. The position error grows about the same way, 1.10 cm·√N to within 5 %: 8.1, 12.1, 17.0, 24.3 and 34.6 cm after 60, 120, 240, 480 and 960 steps. A lap of 240 steps ends 17 cm off, 0.80 % of the 21.4 m walked, and a longer absence is more steps and more drift.

A pose that is wrong writes the same post at slightly different places on each pass, and the map blurs; the question, too, is asked at the believed pose. With a window of 256 frames (long enough to reach back a lap), noise scale 1 gives the map 3,518 cells where the exact walk needs 983, and its score falls from 0.894 to 0.689; at scales 2, 4 and 8 it is 0.456, 0.165 and 0.038. The window with the same pose pays the same price: 0.697 at scale 1 against the map's 0.689. Neither is at fault. Both place what they hold by the believed pose, so both inherit its error. That is what the idea costs.

Dead reckoning cannot notice that it has come back; a place can. A post of a known colour should stand where it stood. Loop closure uses this. The agent keeps a registry of the posts it has seen, each with a colour and an estimated position, frozen once the post leaves the view. When a post comes fully into view it is matched to the nearest frozen entry of the same colour within 1 m. The matched pairs (where I now put the post, where the registry says it is) give the rigid motion that best lines them up, a rotation and a translation by least squares (the alignment of 3D lesson 4), and the pose is moved by it. The rotation is fitted only if two matched posts are at least 1 m apart, because one point cannot fix an angle.

At noise 1 closure cuts the pose error from 16.0 to 6.7 cm and lifts the map from 0.689 to 0.814. Over 40 worlds (posts drawn afresh) it lifts the mean from 0.698 to 0.784 and wins in 35 of them. It repairs the pose from then on, not the lap-1 map, which was written with a drifting pose. On twelve worlds, with an exact first lap closure scores 0.873 against a ceiling of 0.886; with a drifting first lap, 0.797. And it degrades as the drift grows: 0.814, 0.725, 0.580 and 0.390 at scales 1, 2, 4 and 8, where the open-loop pose error (123 cm at scale 8) is past its 1 m matching gate.

Leave, return, and ask what is there
Left: the yard from above, with a key under it. Right, top: the strip really there at the moment asked about (truth) and what each memory says it will be, a red bar under every column the memory had nothing for and the prior had to guess, and the score on this question. Right, bottom: each memory's score over the 12 questions against the lap length n; the dashed line is chance. The first slider sets the memory size M; the last two selects belong to §6.
window score
—
keyframes score
—
map score
—
chance (no memory)
—
window, numbers
—
keyframes, numbers
—
map, numbers
—
map cells
—
keyframes kept
—
pose error, lap 2 (rms)
—
look-alike pairings right
—
Show the core JS
PL.splat = function (pts, q, out) {
  var lab = out.lab, zb = out.z, c = Math.cos(q[2]), s = Math.sin(q[2]), tn = Math.tan(FOV / 2), i, j;
  for (i = 0; i < pts.length; i += 3) {
    var dx = pts[i] - q[0], dy = pts[i + 1] - q[1], fw = dx * c + dy * s, lf = -dx * s + dy * c;       // forward and leftward offsets in the camera frame
    if (fw < 0.05 || Math.abs(lf) > tn * fw) continue;
    j = Math.floor((FOV / 2 - Math.atan2(lf, fw)) / DB);
    var z = Math.sqrt(dx * dx + dy * dy);
    if (j >= 0 && j < NPX && z < zb[j]) { zb[j] = z; lab[j] = pts[i + 2]; }
  }
};

PL.Cells.prototype.write = function (s, p) {
  var P = PL.points(s, p), j, k, e;
  for (j = 0; j < NPX; j++) {
    k = (Math.round(P[3 * j] / PL.CELL) + 4096) * 8192 + Math.round(P[3 * j + 1] / PL.CELL) + 4096; e = this.c.get(k);
    if (e && e[3] === P[3 * j + 2]) { e[0] += P[3 * j]; e[1] += P[3 * j + 1]; e[2]++; }
    else this.c.set(k, [P[3 * j], P[3 * j + 1], 1, P[3 * j + 2]]);                               // a new place, or the thing there changed: write over it
  }
};

What to try. Start at the defaults: M = 64, n = 240, exact poses. The window scores 0.015 (chance 0.014), the keyframes 0.908 and the map 0.894. Move M to 256: the window rises to 0.887, at 33,536 numbers, and the other two do not move. Now set n = 480: the window falls to 0.080; the keyframes (0.908) and the map (0.894) do not. Set the noise to ×1 (n = 240, M = 256): the map blurs to 0.689 with 3,518 cells and the pose is off by 16.0 cm; the window falls with it (0.697). Turn loop closure on at ×1: 0.814, pose error 6.7 cm; at ×4, the map's 0.165 becomes 0.580 (pose 63.0 to 26.5 cm); at ×8 it reaches only 0.390.

Road not taken · make the state a full 3D reconstruction
The 3D Vision track builds the full thing: poses from images by bundle adjustment (3D lesson 5) and surfaces; Marble (World Labs, 2025) exports its worlds as splats or meshes. It is the same idea at higher fidelity and it inherits both limits: it is static, so a pushed post is a stale surface (§6; 3D lesson 14 asks what to do when the scene moves), and it holds surfaces, not things.

5 · What real systems report

SystemWhat its authors report about memory
Genie (Bruce et al., 2024)16 frames of memory, about 1 frame per second
Genie 2 (DeepMind, 2024)consistent worlds for up to a minute, most examples 10–20 s; remembers regions that left the view
Genie 3 (DeepMind, 2025)24 frames per second at 720p; consistent for a few minutes, with visual memory extending as far back as one minute
RTFM (World Labs, 2025)posed frames act as a spatial memory: "context juggling" retrieves the stored frames near the pose asked about, so memory is not bounded by an ever-growing context; no explicit 3D representation
Marble (World Labs, 2025)worlds can be exported as Gaussian splats, meshes or videos

Read against §1 and §3. Genie's memory is 16 frames, and lesson 11 listed windows of seconds (WHAM, Oasis, Dreamer 4) and the limited memory of DIAMOND on CS:GO. Genie 2 and 3 state their horizons in seconds and minutes, the quantity §1 measured in steps: a minute at 24 frames per second is 1,440 frames. RTFM resembles the keyframe store of §3 inside a learned renderer: posed frames kept as memory, the ones near the pose asked about retrieved. Marble exports an explicit 3D world, the road above. These are the authors' statements about their systems; the revisit test of §1 is our own measurement, on hand-written memories.

6 · What a map cannot say

A cell holds a place and a colour. Two things follow that a map cannot express, and each runs in the widget.

Which is which. Put two posts of one colour d = 1.2 m apart (posts: two look-alikes). With exact poses nothing is lost: the map scores 0.896 against 0.894 with distinct colours, because each post is where it is. Trouble starts when the pose is wrong. A returning post must be matched to a stored one, colour cannot decide, so position must. Let each sighting land at its stored place plus an independent error of standard deviation s along the line between the posts: A′ = A + eA, B′ = B + eB, B = A + d. The best position can do is keep the order, right exactly when eA − eB < d. That difference has standard deviation s√2, so, with Φ the standard normal distribution function,

P(right) = Φ( d / (s√2) )

s / d0.250.51248
P(right), %99.892.176.063.857.053.5

It falls from certainty to a coin toss once the error passes the gap (a simulation in the plane agrees within 1.5 points: 74.7 % at s = d). No amount of appearance memory changes this: the two look the same, and what could separate them is their history, which a map does not keep. In the walk the registry of §4 does the matching. Over 16 worlds on the 240-step lap it pairs the right twin 100 % of the time at noise ×1 (1,720 pairings) and 86.9 % at ×4 (1,719); on the 480-step lap, with more drift, 75.0 % at ×4 (3,089). A wrong pairing is not a wrong label: it moves the pose. At ×2 on the 480-step lap the mean pose error is 39.0 cm with twins and 28.6 cm with distinct posts. One yard shows how bad it can get, the widget's own, with look-alikes, ×1, closure on and n = 480: the registry takes the second twin for the first, pairs right 29.6 % of 213 times and ends with a pose error of 76.3 cm, against 21.7 cm with no closure at all.

What happened. Set while the agent is away to the push: after lap 1 a hidden hand moves post 4 by 1.2 m. With exact poses and M = 256 the window, keyframes and map score 0.642, 0.655 and 0.646, against 0.887, 0.908 and 0.894 unpushed. The arena shows why; the red ring marks where the post was. After lap 2 the map holds 19 cells of that post's colour at the old place and 18 at the new one: the post was written again where it now stands and never erased where it was. A map that carves free space along every ray would erase the ghost, and that is a better map. It would still hold a post at the new place with no link to the old one. Nothing in it says that the two are one post, that it moved, what moved it, or what it would do on touching another.

Both failures have one cause. The map's unit is a cell at a place. A world with things in it has units that persist when they move, carry a pose of their own, and meet. A state for such a world has an entry per thing, not per place.

What this lesson did not do
It did not learn anything, where real systems learn what to keep and how to read it. The yard is flat, the strip comes with a depth channel, the scene is static except for one push, and the pose comes from an odometer, not from images (3D lessons 4, 5 and 14 give the image side). It did not carve free space, keep a pose graph or re-optimise the map after a closure. How a state should hold things is lesson 13; when a model can be trusted about places it cannot check is lesson 16.

Common mistakes / failure modes

"a longer window fixes memory"
It reaches exactly as far as it is long: about 0.8 n frames for a lap, 98,512 numbers at n = 960, 32.6 times the map (§3).
"if it predicts the next frame well, it remembers"
A window of 64 frames scores 0.90 one step ahead and 0.015 48 steps ahead (§1).
"a map does not need a pose"
It is written and read by pose: at noise ×1 the map scores 0.689, and a window using the same pose 0.697 (§4).
"loop closure fixes the map"
It fixes the pose from then on, not a map drawn with a drifting one: 0.873 with an exact first lap, 0.797 without (§4).
"a thing is identified by how it looks"
Look-alikes are told apart by position, right with probability Φ(d/(s√2)): 76.0 % when the error equals the gap (§6).
"a map follows the world as it changes"
After a push it keeps 19 cells at the old place and scores 0.646 against 0.894 (§6).

Checkpoint exercise

Try it
A lap has n = 480 steps. (a) About how many frames must a window hold to reach 80 % of the map's score, and how many numbers is that at 131 per frame? (b) How many numbers do the keyframes need to cover the same loop, and the map? (c) Two look-alike posts stand 1.2 m apart and the position error is s = 1.2 m: how often does matching by position pair them right? Answer: (a) about 0.8 × 480 ≈ 384; the sweep finds 376 frames, and 376 × 131 = 49,256 numbers. (b) The same 60 keyframes as at n = 240, 7,860 numbers, and the map 2,985 (995 cells): they depend on the yard, not the lap. (c) Φ(1/√2) = 76.0 %.

Where this points next

A pose and a map let a world model stay consistent about places it has stopped looking at: on the look-ahead test a window of 64 frames scores 0.015 and the map 0.894, with 983 cells at n = 240 and 1,007 at 960. The price, a pose that drifts, takes the map to 0.689 at noise ×1, and loop closure brings it back to 0.814. But the map holds appearance at places and nothing else. Two posts of one colour 1.2 m apart are told apart by position alone, right with probability 76.0 % when the pose error equals the gap; closure pairs them right 86.9 % of the time at noise ×4 on the 240-step lap and 75.0 % on the 480-step lap. When a hidden hand pushed one post, the map kept 19 cells at the old place, gained 18 at the new one, and scored 0.646 against 0.894. It cannot say that the post at the new place is the post that stood at the old one, or what moved it. How should a model represent entities and their interactions?

Takeaway
A finite window forgets by age. On the revisit test a window of 64 frames scores 0.015, chance 0.014, and holding a lap takes about 0.8 n frames: 376 at n = 480. Memory indexed by place does not age: keyframes (0.908) and a map (0.894) keep their score however long the lap. The state of such a world is a pose and a map, view = render(map, pose), and the pose explains 93.3 % of what changes between views. The price is the pose: dead reckoning drifts as √N, the map blurs to 0.689, and recognising a revisit repairs the pose (0.814) but not a map already bent. What a map holds is what is at a place: look-alike posts are paired right with probability Φ(d/(s√2)) for independent errors of spread s, and a pushed post leaves a ghost, 0.894 falling to 0.646. A world made of things needs a state made of things.

Interview prompts

Companion reads: Lesson 25 · Horizon, memory, and the quadratic bill (what a long window costs), 3D Vision · 05 Pose by agreement II: images alone (the pose from pictures instead of an odometer), 3D Vision · 14 When the scene moves (SLAM, and what breaks when the scene is not static) and Computer Vision · 05 Multi-view, depth and SLAM (the survey of mapping and loop closure).