all_lessons/World Models/13 · Thingslesson 13 / 31

Things: a world made of objects

Lesson 12's map knows where things are, not which is which, and not what happens when one pushes another. Send two identical balls behind the curtain, one from each side: whether the left one to come out is the one that went in on the left depends on whether they bounced, and no map holds that. This lesson builds a Courtyard of several balls and derives the state such a world needs. Relabelling the balls must relabel the prediction, the weights cannot depend on how many balls there are, and pushes add; together these leave one rule, a pairwise effect shared by every pair and summed over neighbours. Trained on three balls it runs on ten, and it keeps names through the curtain. It fixes what the model can represent, not what it can learn from.

The thesis, here
A world of several things needs a state that is a set of things and dynamics that treat them alike. If the balls are identical, relabelling them must relabel the prediction; if the model must run with any number of them, its weights cannot depend on how many there are; and if pushes add, so must the effects it learns. One form is left: each ball's next state is the one-ball model plus the sum, over the other balls, of one shared pairwise effect. It carries a network trained on 3 balls to 10 and lets a name survive an encounter nobody saw; it says nothing about where the data come from.
Linear position
Forced by: A map remembers where things are, not what they are: it cannot say that this cup is the cup that was on the table a minute ago, or what happens when one thing pushes another. A world made of things needs a state made of things. How should a model represent entities and their interactions?
New idea: a state made of things: a set of slots, one shared self-model and one shared pairwise rule summed over the neighbours. Relabelling then relabels the answer, the parameters do not depend on how many things there are, and a name is a slot whose prior includes the rule.
Forces next: Structure inside the model fixes what it can represent, not what it can learn from. The video on the internet shows what happened and not what anyone did to make it happen, and an interactive world must answer a button press now, not after a minute of sampling. Where do actions come from when the data has none, and how fast can a model run?
The plan
Six moves. (1) Send two balls behind the curtain: a map fails. (2) Write the balls side by side and count what that costs. (3) Derive the form that relabelling, a variable number of balls and additive pushes leave. (4) Bind names to slots by a prior and a competition. (5) Run one trained rule on 2 to 10 balls. (6) Ask what the structure cannot supply: the pushes, and the speed.

1 · Two balls behind a curtain

The Soft Courtyard keeps the Courtyard's court (8 m by 5 m, x to the right, y up), its friction (3.4 % of a ball's velocity lost per step of 0.1 s) and its curtain (x from 2.6 m to 3.4 m, where the detector, which reads a position with 0.1 m of noise, sees nothing). It holds N identical unit-mass balls. Each carries a soft shell of radius 0.25 m; where two shells overlap by δ metres each ball is pushed away from the other along the line of centres with acceleration kδ², k = 300 m−1s−2, and a wall pushes a shell the same way (k = 600). The push comes from a potential, so a collision loses nothing (friction off, a head-on encounter at ±1 m/s ends within 0.0002 % of its starting energy, kinetic plus contact, with the two velocities exchanged). Contact is soft so that a small network can learn it; the hard bounce is lesson 15's.

Which is which. Fire ball A from the left (x = 1.4 m) and ball B from the right (x = 4.6 m) at similar speeds (1.0 to 1.8 m/s, B's within 15 % of A's), their paths a height b apart, so that they meet under the curtain. The detector sees both, goes blind, and sees two points again, one on each side. In 2000 such trials with b drawn from ±0.5 m (1672 in which both balls come out again within six seconds) the balls bounced in 41 % and passed in the rest. A map, "a ball is where one was last seen", names the points right in 42 %, about the share that bounced. A model that lets each ball go on alone is right in 59 %, about the share that passed. After the curtain the picture is a point on each side either way: which ball is which is decided by how the balls interact, and neither picture holds that.

What happens next. Drop 3 balls into a room of 2.4 m by 1.6 m in the middle of the court, speeds 0.3 to 1.8 m/s, and ask where they are after one second. The one-ball model of lessons 1 to 6, run on each ball, is off by 20.5 cm on average, because in 64 % of those seconds two shells touch; with 10 balls it is off by 56.7 cm. The state must hold several things, let them push one another and say which is which.

2 · The obvious state: the balls side by side

Write the balls after one another and let one network read the row. Let g be the one-ball model (lessons 4 to 6; it is exact here), so g(si) is where ball i would go alone, si being its position and velocity. For receiver i the row holds each other ball's position and velocity relative to i (in units of 0.5 m and 1.5 m/s: the most favourable form to give a flat net), in index order: 4(N − 1) numbers. The network writes a correction to g(si), the only thing left to learn (§3's rule gets g too). Train it on 800 scenes of 3 balls followed for 5 steps, 6000 Adam steps, 16 hidden units twice: 484 parameters. At N = 3, ten steps ahead, it is off by 5.7 cm against 20.5 for the one-ball model. That is the best of eight initialisations (the other seven end up to 17.9 cm off); the widget trains that one, so the comparison favours the flat net. Three costs follow.

The count is in the weights. A row has 8 numbers, so there is no input for a third neighbour; to run the net on other N each receiver keeps its first two neighbours in index order. Ten steps ahead the error is 25.7 cm at N = 5 against 34.0 for the one-ball model, and 53.1 cm at N = 10 against 56.7: it no longer sees most neighbours. A net built for N balls needs 64(N − 1) + 356 parameters.

The order is in the weights. The balls are identical and the numbering is ours, yet the answer depends on it: relabel the balls at random, predict, relabel the answer back, and in states where two balls are within 0.6 m the velocities differ by 3.9 cm/s (rms). The same pair effect has to be learned again for every place in the row.

Learning it again is no answer. A receiver's senders can come in (N − 1)! orders: 2, 24, 5040, 362,880 for N = 3, 5, 8, 10. Retraining at each N does not pay either (Road, below).

3 · What relabelling forces

Three constraints, two on the model and one on the world. R1, relabelling relabels the answer (permutation equivariance): rename the balls by any permutation and the predicted next states are renamed the same way. R2, the weights do not depend on N. R3, pair effects add: the change a ball gets from its neighbours is the sum of the changes each would give it alone. R3 is a fact about the world, and it is checked: a ball's velocity change beyond g, in a 3-ball step, differs from the sum over the 2-ball steps by 0.20 cm/s rms, 1.5 % of the change itself (12.7 cm/s); at N = 8 the gap is 5.7 %. (A ball touching two others moves during the step, which changes the second contact: the sum is near, not exact.)

form of the answer for ball iR1R2R3what goes wrong
the others side by side (§2)nono–the count and the order are in the weights
the nearest two, nearest first (Road, below)yesnonoa third neighbour has no input
the mean of the neighbours' effectsyesyesnoa push is divided by the number of neighbours, N − 1
the largest of the neighbours' effectsyesyesnotwo neighbours push no harder than one
the sum of the neighbours' effectsyesyesyesnothing: this is the form

That leaves

ŝi′ = g(si) + Σj≠i φ(si, sj)

where ŝi′ is the predicted next state of ball i and φ is one network that every ordered pair shares. It reads the sender as seen from the receiver (position and velocity differences: contact does not care where in the court it happens, since walls live in g, nor how fast the pair moves together). R1 holds because a sum does not care in what order it is added, R2 because φ has 4 inputs and 4 outputs whatever N is: φ (4 → 16 → 16 → 4) has 420 parameters for any N. The price is one φ per ordered pair, N(N − 1) per step: 90 at N = 10. This is the shape of the Interaction Network (Battaglia et al., 2016), a graph network over objects and relations that simulated n-body, rigid-body collision and non-rigid dynamics: one function of each pair, summed over a receiver's senders (listing under the widget).

Train it on the same data and steps. At N = 3, ten steps ahead, it is off by 4.6 cm (rank 3 of its eight initialisations, which end between 4.3 and 5.9), the flat net by 5.7 (its best) and the one-ball model by 20.5. Relabel the balls: the answer changes by 0 cm/s, and at 5 and 8 balls by nothing beyond rounding.

4 · Who is who: binding points to slots

A detector hands over a bag of points with no names. The model keeps one slot per thing, a vector of its position and velocity: the belief of lesson 2, a Kalman filter per ball and axis, fed by every detection while the ball is visible (before the curtain the points are taken as named). Behind the curtain the dynamics advance the slot; when points reappear each must be bound to a slot. Two choices decide how well: what each slot expects (its prior) and how points are given to slots (the rule).

Priors: map, the slot stays where it was last seen; free flight, g advances it with no neighbours; interaction-aware, g + Σφ advances all the slots together, once with the true physics and once with the trained network. Rules: nearest, each point takes the slot whose expectation is nearest; competition, each point splits its attention among the slots by a softmax of the squared distance (temperature 0.25 m), each slot moves to the attention-weighted mean of the points, and after four rounds each point goes to the slot with most of its attention; optimal, the one-to-one matching with the least total squared distance. Competition is the step of Slot Attention (Locatello et al., 2020), whose slots bind to objects through a competitive procedure over several attention rounds, here on detected points with distances in place of learned projections.

prior of each slotnearest (%)competition (%)right when
map4243they bounce
free flight5959they pass
interaction-aware, true physics94.796.7both
interaction-aware, trained rule94.796.9both

Names kept through the curtain, in the 1672 trials of §1. Only a prior that contains the interaction is right in both cases, and the trained network keeps names about as well as the true physics. It is not perfect: started from the true state instead of the filter's belief, the true-physics prior is right in 100 % of the trials, so the missing points are the belief's noise, which a collision amplifies. With two points the rules differ by two; with two balls on each side at random heights, giving each slot one point matters (667 trials with four balls, trained rule): nearest 55 %, competition 70 %, optimal 75 %. Names are not guaranteed: with six balls even the true physics and the optimal matching keep them in 35 % of trials, since the future of a crowd out of sight is a distribution (lesson 5). SlotFormer (Wu et al., 2023) learns the dynamics over slots with a Transformer and C-SWM (Kipf et al., 2020) learns object embeddings and relations without reconstructing pixels: learned versions of this section.

5 · Run one rule on 2 to 10 balls

Both nets were trained on 3 balls only. The widget trains them in front of you and runs them on any N. The exam is 100 scenes per N (scene q with N + 1 balls is scene q with one ball more), followed for 10 steps; errors are mean distances between predicted and true positions.

Many balls, one rule
Left: ten steps of one scene of N balls in the dashed room; black is what happens, grey dashes the one-ball model, cyan the flat net, purple the pairwise rule. Right: ten-step error against N; N = 3, the training size, is dotted. The view who is who sends two balls through the curtain, offset by b, and plots the share of trials in which each prior keeps the names (competition rule).
ten steps, cm: one-ball / flat / pairwise
—
one step, cm: one-ball / flat / pairwise
—
parameters: pairwise / flat net for this N
—
relabelling changes the answer by, cm/s: pairwise / flat
—
training steps: pairwise / flat
—
orders of the senders (N − 1)! · φ evaluations N(N − 1)
—
names kept, %: map (nearest / competition)
—
free flight
—
true physics
—
trained rule
—
trials counted (share that bounced)
—
Show the core JS
Pair.prototype.messages = function (X, B, N) {           // phi on every sender alone, then the sum over each receiver's N-1 senders: B*N x 4
  var M = this.phi.forward(X, B * N * (N - 1)), E = new Float64Array(B * N * 4), k, r, c, w = this.agg === 'mean' ? 1 / (N - 1) : 1;
  for (k = 0; k < B * N; k++) for (r = 0; r < N - 1; r++) for (c = 0; c < 4; c++) E[4 * k + c] += w * M[4 * (k * (N - 1) + r) + c];
  return E;
};
Pair.prototype.predictB = function (SB, B, N) { return addResid(L13.stepAloneB(SB, B, N), this.messages(L13.pairRows(SB, B, N), B, N), new Float64Array(SB.length)); };
// who is who: Z the points, mu the slots; A[d][k] is the share of point d that slot k wins (softmax over slots), then every slot moves to its weighted mean
A = Z.map(function (z) { var w = mu.map(function (m) { return Math.exp(-d2(m, z) / tau2); }), s = w.reduce(function (a, c) { return a + c; }, 0) || 1e-300; return w.map(function (v) { return v / s; }); });
mu = mu.map(function (m, k2) { var sw = 0, sx = 0, sy = 0; Z.forEach(function (z, d3) { sw += A[d3][k2]; sx += A[d3][k2] * z[0]; sy += A[d3][k2] * z[1]; }); return sw > 1e-9 ? [sx / sw, sy / sw] : m; });

What to try. Press train the pairwise rule six times and train the flat net six times (the best of eight initialisations). At N = 3, ten steps ahead, the one-ball model is off by 20.5 cm, the flat net by 5.7 and the pairwise rule by 4.6. Slide to N = 5: 34.0, 25.7 and 7.0 cm. At N = 10: 56.7, 53.1 and 12.6 cm: with weights trained on 3 balls the pairwise rule removes 78 % of the error. Its error grows slowly with N (3.0 cm at 2, 10.4 at 8); the flat net climbs towards the one-ball model. The parameter readout stays at 420 for the pairwise rule, and relabelling changes its answer by 0 cm/s and the flat net's, at N = 3, by 3.9. Now switch the view to who is who. Head on (b = 0, 60 trials) the map names 100 % of the points right and free flight 0 %, the true physics 83 % and the trained rule 80 % (nearest; competition 88 and 87). At b = 0.5 m the map is right 0 %, free flight 100 %, the trained rule 100 %. Between b = 0.15 and 0.25 m a bounce turns into a pass: the balls glance off sideways along the curtain, most trials never finish (11 of 60 count at 0.2 m) and the two curves cross.

Road not taken · a bigger flat net, or one per N
A wider net still has a fixed number of places in a row, and a net retrained at each N on 800 scenes of N balls, 6000 steps, is hardly better than no model: 31.4 cm ten steps ahead at N = 5 against the one-ball model's 34.0, and at N = 8 50.0 against 50.0, with an answer that still depends on the order of the balls (10.4 cm/s).
Road not taken · sort the senders by distance
The flat net's order has a cheaper cure than a sum: give each ball its two nearest senders, nearest first. Trained as in §2 (eight initialisations), that net is exactly equivariant (relabelling changes its answer by 0 cm/s) and at the training size it beats the pairwise rule, 2.8 to 3.4 cm at N = 3 against 4.3 to 5.9. But two is in its weights as a cap: a third neighbour has no input, however close, and at N = 10 it is off by 15.1 to 16.1 cm against 11.3 to 14.2. Sorting repairs the order and keeps the cap; the sum has neither.
Road not taken · average the messages
A mean looks safer than a sum, since it does not grow with the number of neighbours. But a push does. Trained at N = 3 (same initialisation) the mean does at least as well as the sum (2.6 cm ten steps ahead against 4.6); at N = 8 it is off by 28.5 cm where the sum is off by 10.4, because it divides by N − 1: a contact learned at N = 3 (÷ 2) is felt at 8 (÷ 7) as two sevenths of its strength.
Road not taken · objects for everything
Objects are the wrong abstraction where there is nothing to count: wind is a field with a value at every point, a fluid or a cloth has no ten items whose identity matters, and evaluating every pair costs N(N − 1) messages a step, 999,000 for a thousand particles. Such worlds need a field or a grid beside or instead of the slots.

6 · What the structure cannot supply

The form says what the model can represent, any number of balls pushing in pairs, and nothing about what it is given to learn from. Everything so far came from transitions whose causes were all in the state. Add an agent that nudges: at 29 % of the steps it presses one of four buttons on a random ball, an impulse of 0.8 m/s in a direction the model does not know. Record the video, the states before and after, not the presses. The trained pairwise rule predicts a video without presses with an error of 2.9 cm/s in velocity and the nudged video with 24.0 cm/s, and 97.6 % of the squared error sits on the balls that were pressed: the nudges are what the structure leaves unexplained. Ask the rule what a press of one button would do and it answers that nothing happens, wrong by 76.4 cm/s, nearly the whole push. Give it one labelled press per button and the error falls to 3.9 cm/s: in the toy each button is a fixed impulse, whatever the state.

Real systems are short of labels. Dreamer 4 (Hafner et al., 2025) learned action conditioning from about 100 hours of video paired with actions inside about 2500 hours, 4 % of the data. VPT (Baker et al., 2022) trained an inverse-dynamics model on 1,962 hours of labelled contractor data, as few as about 100 hours giving a fairly accurate one, and used it to label about 70,000 hours of filtered web video, 36 times what was labelled. Genie (Bruce et al., 2024) learned 8 latent actions from 30,000 hours of platformer video with no labels at all.

The second bill is time. A 256 × 256 frame is 196,608 numbers and, in 16 × 16 patches, 256 tokens (lesson 11). At 24 frames a second a frame has 41.7 ms. Sampled token by token that is 256 sequential steps of 0.16 ms; sampled by a denoiser in K passes the budget is 41.7/K ms a pass: 13.9 ms for K = 3, the denoising steps per frame of DIAMOND (Alonso et al., 2024), and 10.4 ms for K = 4, those of GameNGen (Valevski et al., 2025). Dreamer 4 (Hafner et al., 2025) runs a 2-billion-parameter model on one H100 at 21 frames a second with 4 passes: 47.6 ms a frame and 11.9 ms a pass, 14 % more than the 24 fps budget.

What this lesson did not do
It gave the tracker detected points, not pictures: finding slots in pixels needs a learned encoder. The one-ball model was exact and the contact soft; learning the whole step, and hard contact, are lesson 15. The additive rule is near, not exact (5.7 % off at 8 balls), and names stay uncertain in crowds. Unlabelled video leaves the pushes out; inferring the actions, and answering a press within a frame time, are lesson 14.

Common mistakes / failure modes

"a bigger network on the concatenated state will handle more balls"
It has no input for them: retrained at 8 balls on 800 scenes of 8 it is no better than the one-ball model (§2, Road).
"permutation symmetry means the model ignores the order"
Relabelling relabels the answer, exactly: the pairwise rule changes by 0 cm/s, the flat net by 3.9 (§3).
"a mean over the neighbours is safer than a sum"
Forces add. The mean is as good at 3 balls and 28.5 cm off at 8, against 10.4 for the sum (§5, Road).
"a model with the right structure learns whatever the data contain"
It learns what is in the data. On video of nudged balls 97.6 % of its error is the nudges, and what-if is wrong by 76.4 cm/s (§6).

Checkpoint exercise

Try it
A flat net with 16 hidden units twice reads, for each ball, the 4(N − 1) numbers of its senders. (a) How many parameters does it have for N = 6, and how many does φ (4 → 16 → 16 → 4) have? (b) How many orders of the senders can one receiver see at N = 6, and how many evaluations of φ does one step cost? (c) A ball has 7 neighbours: 5 touch it and each alone would change its velocity by f, 2 are far away. What do the sum and the mean of the messages predict? Answer: (a) 64(N − 1) + 356 = 676; φ has 4·16 + 16 + 16·16 + 16 + 16·4 + 4 = 420 whatever N is. (b) (N − 1)! = 120 orders; N(N − 1) = 30 evaluations. (c) The sum predicts 5f, which is the physics; the mean predicts 5f/7, because it divides by the number of neighbours, touching or not.

Where this points next

One rule shared by every pair, trained on 3 balls, runs on 10 with a ten-step error of 12.6 cm where the one-ball model has 56.7, and with the rule in its prior a slot keeps its name through the curtain in 97 % of the trials where a map keeps it in 42. But on video of balls that someone nudged, 97.6 % of the error is the nudges and what a press would do is wrong by 76.4 cm/s; in Dreamer 4 the hours paired with actions are 4 % of the video; and a frame drawn in 41.7 ms leaves 0.16 ms a token, or a few passes of about 10 ms. Structure inside the model fixes what it can represent, not what it can learn from. Where do actions come from when the data has none, and how fast can a model run?

Takeaway
A world of identical things needs a state made of things. Relabelling the balls must relabel the prediction, the weights must not depend on how many there are, and where pushes add that leaves one form: the one-ball model plus the sum, over the other balls, of one shared pairwise effect. It has 420 parameters for any number of balls and relabelling changes its answer by 0; trained on 3 balls it cuts the ten-step error at 10 balls from 56.7 to 12.6 cm, where a flat net sees two neighbours and stays at 53.1, and one that sorts its two nearest senders fixes the order but keeps the cap. A name is a slot with a prior that contains the rule and a competition among slots: through the curtain it is kept 97 % of the time, against 42 for a map. What the form cannot supply is what the data lack: on video of nudged balls the nudges are 97.6 % of the error, and neither the labelled hours nor the time per frame come from structure.

Interview prompts

Companion reads: Computer Vision · 11 Keypoints, pose and tracking (giving detections to tracks), Lesson 23 · Inferring the actions (labelling video by inverse dynamics) and Robot Model Training · 17 What an hour costs: prices and the ledger (what a labelled hour costs).