What should a state keep?
The belief of lesson 2 solved the curtain, but we built it by hand, choosing the state variables and writing down the sensor. A camera hands the agent 360 numbers per picture and does not say which matter, so what should a learned state keep? The tempting answer, whatever rebuilds the picture, spends the code on the loudest variation, here a bright band unrelated to the ball. Asking the code to predict the next code lets it drop what does not persist, but its easiest solution is to keep nothing, so it needs a guard; what persists without mattering only the task retires. Every code here reads one picture, so it has no velocity and no memory of a ball behind the curtain.
New idea: a state is a code that is sufficient for the future that matters and no larger, and it is found by asking the code to predict the next code, with a guard against collapse, not by asking it to rebuild the picture. Where predictability alone cannot say what matters, the task decides.
Forces next: A good state keeps what predicts the future that matters and drops the rest; reconstructing the observation is the wrong way to ask for it, and predicting in a latent space needs protection against collapse. We still lack the machine itself: something that turns a stream of observations and actions into that state and moves it forward. How is the filter learned from data?
1 · What a state is for
The Courtyard now reaches the agent as pictures: 24 × 15 = 360 brightness numbers (x right, y up), a ball drawn as a Gaussian blob of width σ = 1.5 pixels (lesson 2's has 0.8; see the limits), the goal and the curtain. A clip is three consecutive pictures of a ball in free flight, started anywhere in view at 0.5 to 3 m/s in any direction, so that, apart from the walls and the curtain, position and velocity are independent. Two nuisances, both unrelated to the ball, are added: a vertical bright band 0.4 m wide, whose centre is redrawn in every picture and whose peak is 0.8 η (the ball's is 0.8), and a global gain 1 + 0.1 η N(0, 1), also redrawn. η is the nuisance level: 0 is none, 1 is a band as bright as the ball.
Nobody says which numbers matter. Four do, the state s = (x, y, vx, vy), and the only teacher is the stream. The total pixel variance (each pixel's variance across clips, summed) is 4.05 with no nuisance, all of it the ball. At η = 1 it is 21.0: the ball is 19 % of it, and the single loudest direction carries 4.33, more than the ball does in all its pixels together.
What a state is for. A code zt of everything seen so far is sufficient if p(future | zt, actions) = p(future | whole history, actions), so that nothing in the history improves a prediction once zt is known, and minimal if no smaller code is sufficient. In free flight that is s, and it is not a function of one picture: with friction γ = 0.35 s⁻¹ the velocity shrinks by a factor d = 0.966 per picture (0.1 s) and a step moves x to x + g vx, g = (1 − d)/γ = 0.098 s, so vx = (xt+1 − xt)/g: velocity needs two positions. With the stream as the only teacher there are three things to ask of a code: that it rebuild this picture (§2), that it predict the next code (§3, §4), and, if the agent is paid for something, that it predict the reward (§5).
How a code is judged. A probe reads a target out of the code. Ours uses the 10 nearest neighbours: for a test picture, average the true targets of the 10 training pictures whose codes are closest, and score the squared correlation r² with the truth over 200 test clips. It needs only that nearby balls have nearby codes; a linear probe (ridge regression, scored by the share of variance it explains) needs position to be linear in the code, and for a linear code of a blob it is not: with no nuisance, a linear probe reads the ball from the 12 loudest components at 0.88 where neighbours read 0.98. Each code number is read with noise 0.05, the finite precision of any reader: a code whose numbers are much smaller than that says nothing, which matters in §4.
2 · Keep what rebuilds the picture
The first answer is the autoencoder's: keep k numbers from which the picture can be rebuilt as well as possible in squared error. For a linear code the rebuild lies in a k-dimensional subspace, and its average error is tr C − tr(PC), with C the covariance of the centred pictures and P the projector onto the subspace. Write C = Σi μiuiui′, μ1 ≥ μ2 ≥ …: then tr(PC) = Σ μipi with pi = ui′Pui in [0, 1] and Σpi = k, largest when pi = 1 for the k largest μi. The best code is the top k eigenvectors of the pixel covariance (PCA), and the error left is the sum of the discarded eigenvalues.
What it keeps is decided by pixel variance, and pixel variance does not know what a pixel means. At η = 1 the first eight components carry 82 % of the variance and none can be predicted from the previous picture: the correlation between a component now and one picture later is at most 0.14 in size. They are the band and the gain, redrawn every picture. The ball enters from the ninth on (correlations 0.47, 0.52, 0.86 for the ninth to eleventh). The first twelve keep 88 % of the training variance (by the eigenvalues) and 86 % of a held-out picture's, hold the band's centre almost exactly (r² = 1.00) and localise the ball at 0.13. A bigger code does not help this probe: k = 24 gives 0.11, because neighbours are found by distance in the code's own numbers, where the loud band outweighs the ball, whose variance is spread over many quiet components (with no nuisance the loudest carries 10 % of it). Rescaled to unit variance per number, the code shows the ball at 0.20 for k = 12 and 0.67 for k = 24; the predictive code of §3 reads 0.97 and 0.95. With no nuisance the same twelve read the ball at 0.98. PCA is exactly the optimum of its objective, which asks which numbers explain the most pixel variance; that is not our question.
3 · Predict the next code
What does the ball have that a redrawn band does not? It persists: the next picture shows the ball almost where this one does, and the band elsewhere. So ask for a code that can be predicted. Let wt be the picture's 32 leading principal components (98 % of the variance), whitened (rotated and rescaled to unit variance and no correlation over the pictures that enter the loss), which keeps the algebra small and chooses a direction for how well it persists, not how loud it is. The code is zt = E wt, E being k × 32, and a predictor A (k × k) must guess the next picture's code:
ℒ(E, A) = mean over pairs of |zt+1 − A zt|² / k
The target is the next code, not the next picture, so a code that has not kept the band owes nothing for where the band goes next. Every pair is used in both orders: a ball launched backwards in every direction is as likely as one launched forwards, so the stream looks the same run in reverse, up to walls and friction.
The best code, in six lines. (i) For fixed E the best predictor is least squares, A = H G⁻¹, with G = E[ztzt′] and H = E[zt+1zt′] (the next code has the same G, by the reversal); the loss left is (tr G − tr H G⁻¹ H′)/k. (ii) Replace E by cE: G and H scale by c², so the loss does, and E = 0 reaches 0 (§4). (iii) Fix the scale: ask for k numbers of unit variance and no correlation, G = I. Then A = H and ℒ = 1 − |H|F²/k, with |H|F² the sum of the squared entries of H. (iv) In whitened coordinates that constraint is E E′ = I and H = E S E′, S = E[wt+1wt′] being symmetric by the reversal; write S = Σ ρivivi′. (v) The singular values of a compression E S E′ cannot exceed those of S, so |E S E′|F² ≤ Σρi² over the k largest |ρi|, reached when the rows of E are those eigenvectors. (vi) So ℒ = 1 − (1/k) Σi≤k ρi²: a code number costs 1 − ρi², the share of its variance the previous picture cannot predict.
Each ρi is the correlation of a code number with itself one picture later. At η = 1, |ρ| = 0.99 for the best direction, 0.95 for the twelfth and 0.02 for the last of the 32; the best 12-number code loses 0.061 per number, the best of 300 random orthonormal codes 0.51. In the widget (predictive, closed form, η = 1, k = 12) it reads the ball at r² = 0.97, where reconstruction read 0.13; the band at 0.01, where reconstruction read 1.00; and rebuilds 5 % of the pixel variance, where reconstruction rebuilds 86 %. The two objectives disagree about what to keep, and the one that asks about the future answers our question.
| η | 0 | 0.5 | 1 | 1.5 | 2 | 3 |
|---|---|---|---|---|---|---|
| ball's share of pixel variance (%) | 100 | 49 | 19 | 9 | 5 | 2 |
| reconstruction, ball r² at k = 12 | 0.98 | 0.56 | 0.13 | 0.04 | 0.01 | 0.00 |
| predictive, ball r² at k = 12 | 0.97 | 0.96 | 0.97 | 0.96 | 0.95 | 0.83 |
| smallest k with r² ≥ 0.9, on the grid 2, 4, 6, 8, 12, 16, 24: reconstruction | predictive | 4 | 4 | none | 6 | none | 6 | none | 6 | none | 8 | none | 16 |
Reconstruction fails as soon as the nuisance is on (at η = 0.5 the ball is still 49 % of the variance); the predictive code keeps reading the ball up to η = 3. It is not magic: with no nuisance the two tie, and prediction pays a few more numbers as the nuisance grows. The smallest k for the predictive code also depends on the clips drawn: over four draws of the clips it is 4 to 6 at η = 1.
4 · The constraint was doing the work
§3 removed a trivial solution by fiat. A deep network gets no such constraint, only the loss and a gradient. Do that: the same E (12 × 32) and A from a random start, 200 Adam steps on the loss alone (the loss and its gradient are in “Show the core JS”). The spread of a code is the mean standard deviation of its numbers; its effective rank is (Σλ)²/Σλ² for the eigenvalues λ of its covariance, 12 if the numbers vary independently with equal variance and 1 if they all copy one direction. The loss falls from 0.81 to 5.9e-7, and not because the code predicts better: its spread falls from 0.95 to 0.0030 and its rank from 9.0 to 1.8, and the probe reads the ball at r² = 0.00. This is collapse, and the loss calls it a success. By (ii) of §3 it had to: ℒ(0.1 E, A) = 0.01 ℒ(E, A), and nothing in ℒ prefers twelve persistent directions to a constant, which is perfectly predictable. (It is a collapse of scale and rank, not of every trace: with unlimited precision the one or two surviving directions still read r² = 0.78, 0.74 to 0.78 over six random starts.)
| cure | what it does | what it costs |
|---|---|---|
| a reconstruction term | the code must rebuild the picture, so it cannot be constant | the loudest nuisance returns (§2) |
| contrastive negatives | codes of different pictures must stay apart | many negatives: LeCun (2022) argues that in the worst case their number may grow exponentially with the dimension of the representation |
| a held-fixed target (BYOL, Grill et al.; SimSiam, Chen and He) | the target code comes from a slowly moving copy (BYOL) or passes no gradient (SimSiam) | the loss keeps its collapsed minimum; training only avoids it (SimSiam reports that collapsing solutions exist) |
| variance and covariance terms (VICReg, 2022; Barlow Twins, 2021) | penalise a code number whose standard deviation is below 1, and correlation between numbers (Barlow Twins: a cross-correlation target) | two more terms with weights; "1" is a choice of units |
| an isotropic Gaussian regulariser (LeJEPA, 2025) | pushes a batch of codes toward N(0, I) | a statistic over a batch; the authors report no stop-gradient or teacher is needed |
We implement the fourth: for the codes of both pictures of a pair, a variance term Σj max(0, 1 − stdj)²/2k and a covariance term Σi≠j covij²/2k, each with weight 1. Same start, each term alone and both:
| guard | final loss | spread | effective rank | ball r² |
|---|---|---|---|---|
| none | 5.9e-7 | 0.0030 | 1.8 | 0.00 |
| covariance only | 3.7e-6 | 0.0048 | 1.0 | 0.01 |
| variance only | 0.024 | 0.97 | 1.5 | 0.83 |
| both | 0.055 | 0.95 | 12.0 | 0.97 |
Decorrelation alone cannot stop the shrinking: a constant has no correlation to penalise. The variance term alone stops the shrinking and not the copying: twelve numbers that mostly repeat one or two persistent directions are predicted almost perfectly (loss 0.024, below the constrained optimum 0.061) at an effective rank of 1.5. With both, the rank is 12.0 and the ball reads 0.97 (0.96 to 0.97 over six starts). Its loss 0.055 is the optimum 0.061 times the squared spread, 0.90: the hinge lets the code shrink a little. This recipe, a prediction term plus a regulariser that keeps the code spread out, is also the shape of LeWorldModel (Maes et al., 2026), a world model trained end to end from pixels with two loss terms, next-embedding prediction and a Gaussian regulariser. A constant code is perfect for any encoder and predictor; with a linear encoder output and a linear predictor, as here, the descent toward it follows the same c² law.
The widget
What to try. Start at η = 1, k = 12, reconstruction: the rebuilt picture has the band and almost no ball; ball 0.13, band 1.00, pixels rebuilt 86 %. Drag η to 0: the ball climbs to 0.98; at 0.5 it is 0.56. Raise k to 24: it stays at 0.11 (§2 says why). Switch to the closed-form predictive code: ball 0.97, band 0.01, pixels rebuilt 5 %, loss 0.061; at η = 3 it still reads 0.83. Choose the gradient code without the guard and press Train: spread 0.0030, rank 1.8, ball 0.00; with the guard, spread 0.95, rank 12.0, ball 0.97, loss 0.055. Switch the band to drifts (§5): the predictive code reads it at 0.69; the reward-predictive code, after Train, at 0.02, with ball 0.89 and reward 0.94. The velocity readout hovers near zero for every code (at most 0.05 for the first two).
5 · What persists and does not matter
Make the band predictable. In the drifting setting its centre starts anywhere and moves 0.15 m per picture. The next band is a function of this one, persistence has no reason to discard it, and it does not: at η = 1 and k = 12 the predictive code reads the band's centre at r² = 0.69 (hopping band: 0.01) and the ball at 0.93 (hopping: 0.97), and reaching 0.9 takes k = 12 instead of 6. A predictive objective cannot call a predictable thing irrelevant.
Only the task can. If the agent is paid for ending near a goal at (6.6, 2.5), a state is sufficient when it predicts that reward. MuZero (Schrittwieser et al., 2020) trains its latent state to predict reward, value and policy, with no requirement that it rebuild the observation, and justifies planning in the learned model by value equivalence: from the same real state, the cumulative reward along a path through the abstract model matches the real one. EfficientZero (Ye et al., 2021) adds a SimSiam-style temporal consistency loss: both ideas of this lesson in one model. Here: the same 32 coordinates, a bottleneck of k tanh units and a small head, trained on the 600 labelled pictures of the training clips to output minus the distance to the goal; the bottleneck is the code. With the drifting band, η = 1 and k = 12 it reads the band at 0.02 (predictive code: 0.69), the ball at 0.89 and the reward at 0.94; the head alone explains 0.96 of the held-out reward variance. The band is retired because it does not change the reward.
It is not free. It is built from labels: the same head explains 0.37 of the held-out reward variance when trained on 150 labelled pictures and 0.96 on 600; the predictive code needed none to be built. And it is committed to its question: with the hopping band and k = 2, the reward code reads the reward it was trained for at 0.95 and the reward of a goal moved to (4.5, 4.2) at 0.60, where the closed-form predictive code of that size reads 0.76 and 0.78 (widget: goals A and B).
6 · What no code of one picture can hold
Every code so far reads one picture, and a picture is drawn from the ball's position alone (plus the nuisance). So the velocity a code can read is at most what a flexible function of the exact position reads, and the stream draws velocity independently of position except through the walls and the curtain: fitted on 200,000 clips, such a function reaches r² = 0.04. The codes agree. Over the reconstruction and predictive codes, every k from 2 to 24, every nuisance level and both kinds of band, the largest velocity reading is 0.05, and shuffling the velocities, so that there is nothing to read, gives up to 0.03 on 200 test clips. Two pictures do hold it: ridge regression from the difference of two clean consecutive pictures explains 0.92 of the velocity variance, since the difference is the displacement g v. With the nuisance (η = 1) the same regression explains −0.87, worse than guessing the mean: the difference of two bands is far louder than a ball's displacement.
And no code of a fixed number of pictures can hold what is not in them. A ball behind the curtain is not drawn, so two pictures with the ball at different hidden positions are identical to the last digit. A ball launched at 2 to 3 m/s (lesson 1's launcher) stays hidden for 5.4 pictures in a row on average, and for 81 % of the pictures taken behind the curtain the picture before was blind too: even a code of the last two pictures has nothing to read. The state has to be carried: zt+1 must be computed from zt, the new picture and the action, so that what was seen before the curtain is still there behind it.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
A state can now be asked for in the right way: keep what predicts, guard against keeping nothing, and let the task retire what persists without mattering. On the Courtyard's pictures that gives a 12-number code that reads the ball at r² = 0.97 through a band as bright as the ball, where rebuilding the picture reads 0.13. But every such code is a function of one picture: it reads the ball's velocity at 0.05 at best, and for 81 % of the pictures taken behind the curtain the picture before was blind too. A state has to be carried from picture to picture, and no machine of ours does it: lesson 2's filter was written by hand for a model we knew. How is the filter learned from data?
Interview prompts
- What do sufficiency and minimality mean for a state? (§1 — the future is independent of the history given the code, and nothing smaller works; in the Courtyard, position and velocity, which needs two pictures.)
- Why does a linear autoencoder keep the nuisance? (§2 — its optimum is the top eigenvectors of the pixel covariance, ranked by variance, and the band carries more variance than the ball.)
- Derive the best code for predicting the next code. (§3 — least-squares predictor, unit-variance decorrelated constraint, top eigenvectors of the symmetrised lag-1 covariance by |ρ|, loss 1 − mean ρ².)
- Why does the unconstrained latent-prediction loss collapse, and how is it prevented? (§4 — scaling the encoder by c scales the loss by c², so a constant code reaches zero; add a reconstruction term, negatives, a slow target, variance and covariance terms, or a Gaussian regulariser, each with a price.)
- A nuisance drifts predictably. What does a predictive code do, and what retires it? (§5 — it keeps it, because it is predictable; a code trained to predict the reward drops it because it does not change the reward.)
- Why can a code of one picture have no velocity and no memory? (§6 — the picture depends on position only and a hidden ball is not drawn; the state must be carried across pictures.)
Companion reads: Computer Vision · 14 Self-supervised learning (the guards of §4), Lesson 19 · The prediction space is a bit budget (what the loss spends capacity on) and Lesson 20 · The tokenizer is the ceiling (what an encoder drops is gone).