Things: a world made of objects
Lesson 12's map knows where things are, not which is which, and not what happens when one pushes another. Send two identical balls behind the curtain, one from each side: whether the left one to come out is the one that went in on the left depends on whether they bounced, and no map holds that. This lesson builds a Courtyard of several balls and derives the state such a world needs. Relabelling the balls must relabel the prediction, the weights cannot depend on how many balls there are, and pushes add; together these leave one rule, a pairwise effect shared by every pair and summed over neighbours. Trained on three balls it runs on ten, and it keeps names through the curtain. It fixes what the model can represent, not what it can learn from.
New idea: a state made of things: a set of slots, one shared self-model and one shared pairwise rule summed over the neighbours. Relabelling then relabels the answer, the parameters do not depend on how many things there are, and a name is a slot whose prior includes the rule.
Forces next: Structure inside the model fixes what it can represent, not what it can learn from. The video on the internet shows what happened and not what anyone did to make it happen, and an interactive world must answer a button press now, not after a minute of sampling. Where do actions come from when the data has none, and how fast can a model run?
1 · Two balls behind a curtain
The Soft Courtyard keeps the Courtyard's court (8 m by 5 m, x to the right, y up), its friction (3.4 % of a ball's velocity lost per step of 0.1 s) and its curtain (x from 2.6 m to 3.4 m, where the detector, which reads a position with 0.1 m of noise, sees nothing). It holds N identical unit-mass balls. Each carries a soft shell of radius 0.25 m; where two shells overlap by δ metres each ball is pushed away from the other along the line of centres with acceleration kδ², k = 300 m−1s−2, and a wall pushes a shell the same way (k = 600). The push comes from a potential, so a collision loses nothing (friction off, a head-on encounter at ±1 m/s ends within 0.0002 % of its starting energy, kinetic plus contact, with the two velocities exchanged). Contact is soft so that a small network can learn it; the hard bounce is lesson 15's.
Which is which. Fire ball A from the left (x = 1.4 m) and ball B from the right (x = 4.6 m) at similar speeds (1.0 to 1.8 m/s, B's within 15 % of A's), their paths a height b apart, so that they meet under the curtain. The detector sees both, goes blind, and sees two points again, one on each side. In 2000 such trials with b drawn from ±0.5 m (1672 in which both balls come out again within six seconds) the balls bounced in 41 % and passed in the rest. A map, "a ball is where one was last seen", names the points right in 42 %, about the share that bounced. A model that lets each ball go on alone is right in 59 %, about the share that passed. After the curtain the picture is a point on each side either way: which ball is which is decided by how the balls interact, and neither picture holds that.
What happens next. Drop 3 balls into a room of 2.4 m by 1.6 m in the middle of the court, speeds 0.3 to 1.8 m/s, and ask where they are after one second. The one-ball model of lessons 1 to 6, run on each ball, is off by 20.5 cm on average, because in 64 % of those seconds two shells touch; with 10 balls it is off by 56.7 cm. The state must hold several things, let them push one another and say which is which.
2 · The obvious state: the balls side by side
Write the balls after one another and let one network read the row. Let g be the one-ball model (lessons 4 to 6; it is exact here), so g(si) is where ball i would go alone, si being its position and velocity. For receiver i the row holds each other ball's position and velocity relative to i (in units of 0.5 m and 1.5 m/s: the most favourable form to give a flat net), in index order: 4(N − 1) numbers. The network writes a correction to g(si), the only thing left to learn (§3's rule gets g too). Train it on 800 scenes of 3 balls followed for 5 steps, 6000 Adam steps, 16 hidden units twice: 484 parameters. At N = 3, ten steps ahead, it is off by 5.7 cm against 20.5 for the one-ball model. That is the best of eight initialisations (the other seven end up to 17.9 cm off); the widget trains that one, so the comparison favours the flat net. Three costs follow.
The count is in the weights. A row has 8 numbers, so there is no input for a third neighbour; to run the net on other N each receiver keeps its first two neighbours in index order. Ten steps ahead the error is 25.7 cm at N = 5 against 34.0 for the one-ball model, and 53.1 cm at N = 10 against 56.7: it no longer sees most neighbours. A net built for N balls needs 64(N − 1) + 356 parameters.
The order is in the weights. The balls are identical and the numbering is ours, yet the answer depends on it: relabel the balls at random, predict, relabel the answer back, and in states where two balls are within 0.6 m the velocities differ by 3.9 cm/s (rms). The same pair effect has to be learned again for every place in the row.
Learning it again is no answer. A receiver's senders can come in (N − 1)! orders: 2, 24, 5040, 362,880 for N = 3, 5, 8, 10. Retraining at each N does not pay either (Road, below).
3 · What relabelling forces
Three constraints, two on the model and one on the world. R1, relabelling relabels the answer (permutation equivariance): rename the balls by any permutation and the predicted next states are renamed the same way. R2, the weights do not depend on N. R3, pair effects add: the change a ball gets from its neighbours is the sum of the changes each would give it alone. R3 is a fact about the world, and it is checked: a ball's velocity change beyond g, in a 3-ball step, differs from the sum over the 2-ball steps by 0.20 cm/s rms, 1.5 % of the change itself (12.7 cm/s); at N = 8 the gap is 5.7 %. (A ball touching two others moves during the step, which changes the second contact: the sum is near, not exact.)
| form of the answer for ball i | R1 | R2 | R3 | what goes wrong |
|---|---|---|---|---|
| the others side by side (§2) | no | no | – | the count and the order are in the weights |
| the nearest two, nearest first (Road, below) | yes | no | no | a third neighbour has no input |
| the mean of the neighbours' effects | yes | yes | no | a push is divided by the number of neighbours, N − 1 |
| the largest of the neighbours' effects | yes | yes | no | two neighbours push no harder than one |
| the sum of the neighbours' effects | yes | yes | yes | nothing: this is the form |
That leaves
ŝi′ = g(si) + Σj≠i φ(si, sj)
where ŝi′ is the predicted next state of ball i and φ is one network that every ordered pair shares. It reads the sender as seen from the receiver (position and velocity differences: contact does not care where in the court it happens, since walls live in g, nor how fast the pair moves together). R1 holds because a sum does not care in what order it is added, R2 because φ has 4 inputs and 4 outputs whatever N is: φ (4 → 16 → 16 → 4) has 420 parameters for any N. The price is one φ per ordered pair, N(N − 1) per step: 90 at N = 10. This is the shape of the Interaction Network (Battaglia et al., 2016), a graph network over objects and relations that simulated n-body, rigid-body collision and non-rigid dynamics: one function of each pair, summed over a receiver's senders (listing under the widget).
Train it on the same data and steps. At N = 3, ten steps ahead, it is off by 4.6 cm (rank 3 of its eight initialisations, which end between 4.3 and 5.9), the flat net by 5.7 (its best) and the one-ball model by 20.5. Relabel the balls: the answer changes by 0 cm/s, and at 5 and 8 balls by nothing beyond rounding.
4 · Who is who: binding points to slots
A detector hands over a bag of points with no names. The model keeps one slot per thing, a vector of its position and velocity: the belief of lesson 2, a Kalman filter per ball and axis, fed by every detection while the ball is visible (before the curtain the points are taken as named). Behind the curtain the dynamics advance the slot; when points reappear each must be bound to a slot. Two choices decide how well: what each slot expects (its prior) and how points are given to slots (the rule).
Priors: map, the slot stays where it was last seen; free flight, g advances it with no neighbours; interaction-aware, g + Σφ advances all the slots together, once with the true physics and once with the trained network. Rules: nearest, each point takes the slot whose expectation is nearest; competition, each point splits its attention among the slots by a softmax of the squared distance (temperature 0.25 m), each slot moves to the attention-weighted mean of the points, and after four rounds each point goes to the slot with most of its attention; optimal, the one-to-one matching with the least total squared distance. Competition is the step of Slot Attention (Locatello et al., 2020), whose slots bind to objects through a competitive procedure over several attention rounds, here on detected points with distances in place of learned projections.
| prior of each slot | nearest (%) | competition (%) | right when |
|---|---|---|---|
| map | 42 | 43 | they bounce |
| free flight | 59 | 59 | they pass |
| interaction-aware, true physics | 94.7 | 96.7 | both |
| interaction-aware, trained rule | 94.7 | 96.9 | both |
Names kept through the curtain, in the 1672 trials of §1. Only a prior that contains the interaction is right in both cases, and the trained network keeps names about as well as the true physics. It is not perfect: started from the true state instead of the filter's belief, the true-physics prior is right in 100 % of the trials, so the missing points are the belief's noise, which a collision amplifies. With two points the rules differ by two; with two balls on each side at random heights, giving each slot one point matters (667 trials with four balls, trained rule): nearest 55 %, competition 70 %, optimal 75 %. Names are not guaranteed: with six balls even the true physics and the optimal matching keep them in 35 % of trials, since the future of a crowd out of sight is a distribution (lesson 5). SlotFormer (Wu et al., 2023) learns the dynamics over slots with a Transformer and C-SWM (Kipf et al., 2020) learns object embeddings and relations without reconstructing pixels: learned versions of this section.
5 · Run one rule on 2 to 10 balls
Both nets were trained on 3 balls only. The widget trains them in front of you and runs them on any N. The exam is 100 scenes per N (scene q with N + 1 balls is scene q with one ball more), followed for 10 steps; errors are mean distances between predicted and true positions.
What to try. Press train the pairwise rule six times and train the flat net six times (the best of eight initialisations). At N = 3, ten steps ahead, the one-ball model is off by 20.5 cm, the flat net by 5.7 and the pairwise rule by 4.6. Slide to N = 5: 34.0, 25.7 and 7.0 cm. At N = 10: 56.7, 53.1 and 12.6 cm: with weights trained on 3 balls the pairwise rule removes 78 % of the error. Its error grows slowly with N (3.0 cm at 2, 10.4 at 8); the flat net climbs towards the one-ball model. The parameter readout stays at 420 for the pairwise rule, and relabelling changes its answer by 0 cm/s and the flat net's, at N = 3, by 3.9. Now switch the view to who is who. Head on (b = 0, 60 trials) the map names 100 % of the points right and free flight 0 %, the true physics 83 % and the trained rule 80 % (nearest; competition 88 and 87). At b = 0.5 m the map is right 0 %, free flight 100 %, the trained rule 100 %. Between b = 0.15 and 0.25 m a bounce turns into a pass: the balls glance off sideways along the curtain, most trials never finish (11 of 60 count at 0.2 m) and the two curves cross.
6 · What the structure cannot supply
The form says what the model can represent, any number of balls pushing in pairs, and nothing about what it is given to learn from. Everything so far came from transitions whose causes were all in the state. Add an agent that nudges: at 29 % of the steps it presses one of four buttons on a random ball, an impulse of 0.8 m/s in a direction the model does not know. Record the video, the states before and after, not the presses. The trained pairwise rule predicts a video without presses with an error of 2.9 cm/s in velocity and the nudged video with 24.0 cm/s, and 97.6 % of the squared error sits on the balls that were pressed: the nudges are what the structure leaves unexplained. Ask the rule what a press of one button would do and it answers that nothing happens, wrong by 76.4 cm/s, nearly the whole push. Give it one labelled press per button and the error falls to 3.9 cm/s: in the toy each button is a fixed impulse, whatever the state.
Real systems are short of labels. Dreamer 4 (Hafner et al., 2025) learned action conditioning from about 100 hours of video paired with actions inside about 2500 hours, 4 % of the data. VPT (Baker et al., 2022) trained an inverse-dynamics model on 1,962 hours of labelled contractor data, as few as about 100 hours giving a fairly accurate one, and used it to label about 70,000 hours of filtered web video, 36 times what was labelled. Genie (Bruce et al., 2024) learned 8 latent actions from 30,000 hours of platformer video with no labels at all.
The second bill is time. A 256 × 256 frame is 196,608 numbers and, in 16 × 16 patches, 256 tokens (lesson 11). At 24 frames a second a frame has 41.7 ms. Sampled token by token that is 256 sequential steps of 0.16 ms; sampled by a denoiser in K passes the budget is 41.7/K ms a pass: 13.9 ms for K = 3, the denoising steps per frame of DIAMOND (Alonso et al., 2024), and 10.4 ms for K = 4, those of GameNGen (Valevski et al., 2025). Dreamer 4 (Hafner et al., 2025) runs a 2-billion-parameter model on one H100 at 21 frames a second with 4 passes: 47.6 ms a frame and 11.9 ms a pass, 14 % more than the 24 fps budget.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
One rule shared by every pair, trained on 3 balls, runs on 10 with a ten-step error of 12.6 cm where the one-ball model has 56.7, and with the rule in its prior a slot keeps its name through the curtain in 97 % of the trials where a map keeps it in 42. But on video of balls that someone nudged, 97.6 % of the error is the nudges and what a press would do is wrong by 76.4 cm/s; in Dreamer 4 the hours paired with actions are 4 % of the video; and a frame drawn in 41.7 ms leaves 0.16 ms a token, or a few passes of about 10 ms. Structure inside the model fixes what it can represent, not what it can learn from. Where do actions come from when the data has none, and how fast can a model run?
Interview prompts
- Why can't a map say which of two identical balls is which? (§1, §4 — after the curtain the picture is a point on each side whether they bounced or passed; only the interaction decides.)
- Why does an MLP on the concatenated states of N objects generalise badly to other N? (§2 — the count and order of the objects are in its weights: no inputs for new ones, and one pair effect learned once per place.)
- What does permutation equivariance require, and what form does it leave? (§3 — relabelling relabels the output, weights independent of N, effects that add: the one-ball model plus a sum over senders of a shared pairwise function.)
- Why a sum and not a mean over neighbours? (§3, §5 — forces add; a mean divides the push by the number of neighbours.)
- What does a slot add to a detection, and what does competition add to nearest-neighbour matching? (§4 — a slot carries a belief advanced through occlusion, with a prior that includes the interaction; competition gives each point to one slot.)
- If the model has the right structure, why can it still fail to answer "what if I press"? (§6 — presses are not in the video; the nudges are the unexplained 97.6 % of the error, and the labels are the scarce part.)
Companion reads: Computer Vision · 11 Keypoints, pose and tracking (giving detections to tracks), Lesson 23 · Inferring the actions (labelling video by inverse dynamics) and Robot Model Training · 17 What an hour costs: prices and the ledger (what a labelled hour costs).