all_lessons/World Models/03 · What a state keepslesson 3 / 31

What should a state keep?

The belief of lesson 2 solved the curtain, but we built it by hand, choosing the state variables and writing down the sensor. A camera hands the agent 360 numbers per picture and does not say which matter, so what should a learned state keep? The tempting answer, whatever rebuilds the picture, spends the code on the loudest variation, here a bright band unrelated to the ball. Asking the code to predict the next code lets it drop what does not persist, but its easiest solution is to keep nothing, so it needs a guard; what persists without mattering only the task retires. Every code here reads one picture, so it has no velocity and no memory of a ball behind the curtain.

The thesis, here
A state is judged by what it lets you predict, not by what it lets you rebuild. Rebuilding pays in pixel variance, so the loudest nuisance wins the code. Predicting the next code pays for persistence, which finds the ball, but the loss has a cheaper solution, keeping nothing, so it needs an explicit guard against collapse. What persists without mattering can only be retired by asking what the state is for.
Linear position
Forced by: The belief, a distribution over the hidden state updated by predict-then-correct, solves the curtain problem. But we built it from a model of the ball and a model of the sensor that we wrote down by hand, using state variables we chose. A real agent receives pixels, and nobody tells it which of a million numbers matter. What should an agent keep from what it sees, when the only teacher is the stream itself?
New idea: a state is a code that is sufficient for the future that matters and no larger, and it is found by asking the code to predict the next code, with a guard against collapse, not by asking it to rebuild the picture. Where predictability alone cannot say what matters, the task decides.
Forces next: A good state keeps what predicts the future that matters and drops the rest; reconstructing the observation is the wrong way to ask for it, and predicting in a latent space needs protection against collapse. We still lack the machine itself: something that turns a stream of observations and actions into that state and moves it forward. How is the filter learned from data?
The plan
Six moves. (1) Say what a state is for, and put a probe on the pictures. (2) Rebuild the picture and measure what that keeps. (3) Predict the next code instead and derive the best code. (4) Drop the constraint of that derivation, watch the code collapse, and compare the guards. (5) Make the nuisance predictable and let the task retire it. (6) Show what no code of one picture can hold.

1 · What a state is for

The Courtyard now reaches the agent as pictures: 24 × 15 = 360 brightness numbers (x right, y up), a ball drawn as a Gaussian blob of width σ = 1.5 pixels (lesson 2's has 0.8; see the limits), the goal and the curtain. A clip is three consecutive pictures of a ball in free flight, started anywhere in view at 0.5 to 3 m/s in any direction, so that, apart from the walls and the curtain, position and velocity are independent. Two nuisances, both unrelated to the ball, are added: a vertical bright band 0.4 m wide, whose centre is redrawn in every picture and whose peak is 0.8 η (the ball's is 0.8), and a global gain 1 + 0.1 η N(0, 1), also redrawn. η is the nuisance level: 0 is none, 1 is a band as bright as the ball.

Nobody says which numbers matter. Four do, the state s = (x, y, vx, vy), and the only teacher is the stream. The total pixel variance (each pixel's variance across clips, summed) is 4.05 with no nuisance, all of it the ball. At η = 1 it is 21.0: the ball is 19 % of it, and the single loudest direction carries 4.33, more than the ball does in all its pixels together.

What a state is for. A code zt of everything seen so far is sufficient if p(future | zt, actions) = p(future | whole history, actions), so that nothing in the history improves a prediction once zt is known, and minimal if no smaller code is sufficient. In free flight that is s, and it is not a function of one picture: with friction γ = 0.35 s⁻¹ the velocity shrinks by a factor d = 0.966 per picture (0.1 s) and a step moves x to x + g vx, g = (1 − d)/γ = 0.098 s, so vx = (xt+1 − xt)/g: velocity needs two positions. With the stream as the only teacher there are three things to ask of a code: that it rebuild this picture (§2), that it predict the next code (§3, §4), and, if the agent is paid for something, that it predict the reward (§5).

How a code is judged. A probe reads a target out of the code. Ours uses the 10 nearest neighbours: for a test picture, average the true targets of the 10 training pictures whose codes are closest, and score the squared correlation r² with the truth over 200 test clips. It needs only that nearby balls have nearby codes; a linear probe (ridge regression, scored by the share of variance it explains) needs position to be linear in the code, and for a linear code of a blob it is not: with no nuisance, a linear probe reads the ball from the 12 loudest components at 0.88 where neighbours read 0.98. Each code number is read with noise 0.05, the finite precision of any reader: a code whose numbers are much smaller than that says nothing, which matters in §4.

2 · Keep what rebuilds the picture

The first answer is the autoencoder's: keep k numbers from which the picture can be rebuilt as well as possible in squared error. For a linear code the rebuild lies in a k-dimensional subspace, and its average error is tr C − tr(PC), with C the covariance of the centred pictures and P the projector onto the subspace. Write C = Σi μiuiui′, μ1 ≥ μ2 ≥ …: then tr(PC) = Σ μipi with pi = ui′Pui in [0, 1] and Σpi = k, largest when pi = 1 for the k largest μi. The best code is the top k eigenvectors of the pixel covariance (PCA), and the error left is the sum of the discarded eigenvalues.

What it keeps is decided by pixel variance, and pixel variance does not know what a pixel means. At η = 1 the first eight components carry 82 % of the variance and none can be predicted from the previous picture: the correlation between a component now and one picture later is at most 0.14 in size. They are the band and the gain, redrawn every picture. The ball enters from the ninth on (correlations 0.47, 0.52, 0.86 for the ninth to eleventh). The first twelve keep 88 % of the training variance (by the eigenvalues) and 86 % of a held-out picture's, hold the band's centre almost exactly (r² = 1.00) and localise the ball at 0.13. A bigger code does not help this probe: k = 24 gives 0.11, because neighbours are found by distance in the code's own numbers, where the loud band outweighs the ball, whose variance is spread over many quiet components (with no nuisance the loudest carries 10 % of it). Rescaled to unit variance per number, the code shows the ball at 0.20 for k = 12 and 0.67 for k = 24; the predictive code of §3 reads 0.97 and 0.95. With no nuisance the same twelve read the ball at 0.98. PCA is exactly the optimum of its objective, which asks which numbers explain the most pixel variance; that is not our question.

3 · Predict the next code

What does the ball have that a redrawn band does not? It persists: the next picture shows the ball almost where this one does, and the band elsewhere. So ask for a code that can be predicted. Let wt be the picture's 32 leading principal components (98 % of the variance), whitened (rotated and rescaled to unit variance and no correlation over the pictures that enter the loss), which keeps the algebra small and chooses a direction for how well it persists, not how loud it is. The code is zt = E wt, E being k × 32, and a predictor A (k × k) must guess the next picture's code:

ℒ(E, A) = mean over pairs of |zt+1 − A zt|² / k

The target is the next code, not the next picture, so a code that has not kept the band owes nothing for where the band goes next. Every pair is used in both orders: a ball launched backwards in every direction is as likely as one launched forwards, so the stream looks the same run in reverse, up to walls and friction.

The best code, in six lines. (i) For fixed E the best predictor is least squares, A = H G⁻¹, with G = E[ztzt′] and H = E[zt+1zt′] (the next code has the same G, by the reversal); the loss left is (tr G − tr H G⁻¹ H′)/k. (ii) Replace E by cE: G and H scale by c², so the loss does, and E = 0 reaches 0 (§4). (iii) Fix the scale: ask for k numbers of unit variance and no correlation, G = I. Then A = H and ℒ = 1 − |H|F²/k, with |H|F² the sum of the squared entries of H. (iv) In whitened coordinates that constraint is E E′ = I and H = E S E′, S = E[wt+1wt′] being symmetric by the reversal; write S = Σ ρivivi′. (v) The singular values of a compression E S E′ cannot exceed those of S, so |E S E′|F² ≤ Σρi² over the k largest |ρi|, reached when the rows of E are those eigenvectors. (vi) So ℒ = 1 − (1/k) Σi≤k ρi²: a code number costs 1 − ρi², the share of its variance the previous picture cannot predict.

Each ρi is the correlation of a code number with itself one picture later. At η = 1, |ρ| = 0.99 for the best direction, 0.95 for the twelfth and 0.02 for the last of the 32; the best 12-number code loses 0.061 per number, the best of 300 random orthonormal codes 0.51. In the widget (predictive, closed form, η = 1, k = 12) it reads the ball at r² = 0.97, where reconstruction read 0.13; the band at 0.01, where reconstruction read 1.00; and rebuilds 5 % of the pixel variance, where reconstruction rebuilds 86 %. The two objectives disagree about what to keep, and the one that asks about the future answers our question.

η00.511.523
ball's share of pixel variance (%)1004919952
reconstruction, ball r² at k = 120.980.560.130.040.010.00
predictive, ball r² at k = 120.970.960.970.960.950.83
smallest k with r² ≥ 0.9, on the grid 2, 4, 6, 8, 12, 16, 24: reconstruction | predictive4 | 4none | 6none | 6none | 6none | 8none | 16

Reconstruction fails as soon as the nuisance is on (at η = 0.5 the ball is still 49 % of the variance); the predictive code keeps reading the ball up to η = 3. It is not magic: with no nuisance the two tie, and prediction pays a few more numbers as the nuisance grows. The smallest k for the predictive code also depends on the clips drawn: over four draws of the clips it is 4 to 6 at η = 1.

4 · The constraint was doing the work

§3 removed a trivial solution by fiat. A deep network gets no such constraint, only the loss and a gradient. Do that: the same E (12 × 32) and A from a random start, 200 Adam steps on the loss alone (the loss and its gradient are in “Show the core JS”). The spread of a code is the mean standard deviation of its numbers; its effective rank is (Σλ)²/Σλ² for the eigenvalues λ of its covariance, 12 if the numbers vary independently with equal variance and 1 if they all copy one direction. The loss falls from 0.81 to 5.9e-7, and not because the code predicts better: its spread falls from 0.95 to 0.0030 and its rank from 9.0 to 1.8, and the probe reads the ball at r² = 0.00. This is collapse, and the loss calls it a success. By (ii) of §3 it had to: ℒ(0.1 E, A) = 0.01 ℒ(E, A), and nothing in ℒ prefers twelve persistent directions to a constant, which is perfectly predictable. (It is a collapse of scale and rank, not of every trace: with unlimited precision the one or two surviving directions still read r² = 0.78, 0.74 to 0.78 over six random starts.)

curewhat it doeswhat it costs
a reconstruction termthe code must rebuild the picture, so it cannot be constantthe loudest nuisance returns (§2)
contrastive negativescodes of different pictures must stay apartmany negatives: LeCun (2022) argues that in the worst case their number may grow exponentially with the dimension of the representation
a held-fixed target (BYOL, Grill et al.; SimSiam, Chen and He)the target code comes from a slowly moving copy (BYOL) or passes no gradient (SimSiam)the loss keeps its collapsed minimum; training only avoids it (SimSiam reports that collapsing solutions exist)
variance and covariance terms (VICReg, 2022; Barlow Twins, 2021)penalise a code number whose standard deviation is below 1, and correlation between numbers (Barlow Twins: a cross-correlation target)two more terms with weights; "1" is a choice of units
an isotropic Gaussian regulariser (LeJEPA, 2025)pushes a batch of codes toward N(0, I)a statistic over a batch; the authors report no stop-gradient or teacher is needed

We implement the fourth: for the codes of both pictures of a pair, a variance term Σj max(0, 1 − stdj)²/2k and a covariance term Σi≠j covij²/2k, each with weight 1. Same start, each term alone and both:

guardfinal lossspreadeffective rankball r²
none5.9e-70.00301.80.00
covariance only3.7e-60.00481.00.01
variance only0.0240.971.50.83
both0.0550.9512.00.97

Decorrelation alone cannot stop the shrinking: a constant has no correlation to penalise. The variance term alone stops the shrinking and not the copying: twelve numbers that mostly repeat one or two persistent directions are predicted almost perfectly (loss 0.024, below the constrained optimum 0.061) at an effective rank of 1.5. With both, the rank is 12.0 and the ball reads 0.97 (0.96 to 0.97 over six starts). Its loss 0.055 is the optimum 0.061 times the squared spread, 0.90: the hinge lets the code shrink a little. This recipe, a prediction term plus a regulariser that keeps the code spread out, is also the shape of LeWorldModel (Maes et al., 2026), a world model trained end to end from pixels with two loss terms, next-embedding prediction and a Gaussian regulariser. A constant code is perfect for any encoder and predictor; with a linear encoder output and a linear predictor, as here, the descent toward it follows the same c² law.

The widget

What a code keeps of a picture
Panels, in reading order: three test pictures (ball ringed, band's centre ticked) with what the code rebuilds of the middle one, what it leaves out, the ball alone and the code; component variances, coloured by how well one picture predicts the next; probe readings (r²); ball position against code size; the code's spread in a gradient run. Trained codes need Train.
ball's share of pixel variance
—
ball position r²
—
ball velocity r²
—
band position r²
—
reward r², goal A
—
reward r², goal B (moved)
—
pixel variance rebuilt
—
spread of the code
—
effective rank
—
next-code loss
—
smallest k with r² ≥ 0.9 (reconstruction · predictive)
—
Show the core JS
// the next-code loss and its gradient, in covariance form (E: k x 32 encoder, A: k x k predictor, H = E[z1 z0'])
var EC0 = mul(E, k, M, mx.C00, M), G0 = mul(EC0, k, M, Et, k), EC1 = mul(E, k, M, mx.C11, M), G1 = mul(EC1, k, M, Et, k);
var H = mul(mul(E, k, M, mx.M10, M), k, M, Et, k);
var AG0 = mul(A, k, k, G0, k), AH = mul(A, k, k, T(H, k, k), k), AG0At = mul(AG0, k, k, At, k), pred = 0;
for (i = 0; i < k; i++) pred += G1[i * k + i] - 2 * AH[i * k + i] + AG0At[i * k + i];
...
// the guard: a hinge on each code number's standard deviation, and the squared off-diagonal covariances
for (j = 0; j < k; j++) { var v = Math.sqrt(G[j * k + j] + 1e-4), h = Math.max(0, 1 - v); V += h * h / k / 2; if (h > 0) Gam[j * k + j] += guard.v * (-h / (k * v)) / 2; }
for (i = 0; i < k; i++) for (j = 0; j < k; j++) if (i !== j) { Cv += G[i * k + j] * G[i * k + j] / k / 2; Gam[i * k + j] += guard.c * (G[i * k + j] / k); }
...
// the closed form: eigenvectors of the symmetrised lag-1 covariance, ranked by |rho|
ord.sort(function (p, q) { return Math.abs(ec.vals[q]) - Math.abs(ec.vals[p]); });

What to try. Start at η = 1, k = 12, reconstruction: the rebuilt picture has the band and almost no ball; ball 0.13, band 1.00, pixels rebuilt 86 %. Drag η to 0: the ball climbs to 0.98; at 0.5 it is 0.56. Raise k to 24: it stays at 0.11 (§2 says why). Switch to the closed-form predictive code: ball 0.97, band 0.01, pixels rebuilt 5 %, loss 0.061; at η = 3 it still reads 0.83. Choose the gradient code without the guard and press Train: spread 0.0030, rank 1.8, ball 0.00; with the guard, spread 0.95, rank 12.0, ball 0.97, loss 0.055. Switch the band to drifts (§5): the predictive code reads it at 0.69; the reward-predictive code, after Train, at 0.02, with ball 0.89 and reward 0.94. The velocity readout hovers near zero for every code (at most 0.05 for the first two).

Road not taken · engineer the features
Write a detector instead: the brightest pixel is the ball. It needs no training, and with no nuisance it finds the ball (within 0.5 m) in 100 % of the test pictures; at η = 1 in 50 %, at η = 2 in 17 %, since the band is then as bright or brighter. Subtract each column's median first and it recovers to 100 % at both, because the band is a vertical stripe: knowledge about the nuisance that the agent was not given. Make the band horizontal and the column fix finds the ball in 28 %; only a row fix (100 %) repairs it. Every nuisance needs a new engineer; the learned code must find what to ignore from the stream alone, and persistence is its clue.

5 · What persists and does not matter

Make the band predictable. In the drifting setting its centre starts anywhere and moves 0.15 m per picture. The next band is a function of this one, persistence has no reason to discard it, and it does not: at η = 1 and k = 12 the predictive code reads the band's centre at r² = 0.69 (hopping band: 0.01) and the ball at 0.93 (hopping: 0.97), and reaching 0.9 takes k = 12 instead of 6. A predictive objective cannot call a predictable thing irrelevant.

Only the task can. If the agent is paid for ending near a goal at (6.6, 2.5), a state is sufficient when it predicts that reward. MuZero (Schrittwieser et al., 2020) trains its latent state to predict reward, value and policy, with no requirement that it rebuild the observation, and justifies planning in the learned model by value equivalence: from the same real state, the cumulative reward along a path through the abstract model matches the real one. EfficientZero (Ye et al., 2021) adds a SimSiam-style temporal consistency loss: both ideas of this lesson in one model. Here: the same 32 coordinates, a bottleneck of k tanh units and a small head, trained on the 600 labelled pictures of the training clips to output minus the distance to the goal; the bottleneck is the code. With the drifting band, η = 1 and k = 12 it reads the band at 0.02 (predictive code: 0.69), the ball at 0.89 and the reward at 0.94; the head alone explains 0.96 of the held-out reward variance. The band is retired because it does not change the reward.

It is not free. It is built from labels: the same head explains 0.37 of the held-out reward variance when trained on 150 labelled pictures and 0.96 on 600; the predictive code needed none to be built. And it is committed to its question: with the hopping band and k = 2, the reward code reads the reward it was trained for at 0.95 and the reward of a goal moved to (4.5, 4.2) at 0.60, where the closed-form predictive code of that size reads 0.76 and 0.78 (widget: goals A and B).

6 · What no code of one picture can hold

Every code so far reads one picture, and a picture is drawn from the ball's position alone (plus the nuisance). So the velocity a code can read is at most what a flexible function of the exact position reads, and the stream draws velocity independently of position except through the walls and the curtain: fitted on 200,000 clips, such a function reaches r² = 0.04. The codes agree. Over the reconstruction and predictive codes, every k from 2 to 24, every nuisance level and both kinds of band, the largest velocity reading is 0.05, and shuffling the velocities, so that there is nothing to read, gives up to 0.03 on 200 test clips. Two pictures do hold it: ridge regression from the difference of two clean consecutive pictures explains 0.92 of the velocity variance, since the difference is the displacement g v. With the nuisance (η = 1) the same regression explains −0.87, worse than guessing the mean: the difference of two bands is far louder than a ball's displacement.

And no code of a fixed number of pictures can hold what is not in them. A ball behind the curtain is not drawn, so two pictures with the ball at different hidden positions are identical to the last digit. A ball launched at 2 to 3 m/s (lesson 1's launcher) stays hidden for 5.4 pictures in a row on average, and for 81 % of the pictures taken behind the curtain the picture before was blind too: even a code of the last two pictures has nothing to read. The state has to be carried: zt+1 must be computed from zt, the new picture and the action, so that what was seen before the curtain is still there behind it.

What this lesson did not do
The codes were linear maps of 32 principal components of one picture, trained on 200 clips. A nonlinear encoder can keep the band's centre and the ball in a handful of numbers, so reconstruction fails less harshly there; the lesson derived what each objective pays for, not how large the failure is. A ridge probe reads the same twelve-number codes at 0.33 (reconstruction) and 0.47 (predictive) at η = 1. With lesson 2's sharper ball (σ = 0.8 pixel) the order of the codes is the same (0.01 against 0.82 at k = 12) but 200 clips are too few for any code to reach 0.9. Predicting only the next code is not sufficient in general: Littman, Sutton and Singh (2001) give a case where one-step predictions cannot tell two states apart and a two-step test can; here the pictures hold the persistent quantity directly. How a state is used to choose actions is lessons 7 and 9; a future that is not one point is lesson 5.

Common mistakes / failure modes

"a code that rebuilds the picture well must know where the ball is"
Twelve components rebuild 86 % of a held-out picture and read the ball at 0.13: the band is louder (§2).
"predicting the next code always beats rebuilding"
With no nuisance both reach 0.9 with 4 numbers. Prediction wins when the loud variation does not persist (§3).
"a loss near zero means the dynamics are learned"
The collapsed run reached 5.9e-7 with a code of spread 0.0030: zero is the signature of keeping nothing (§4).
"a variance penalty is enough to stop collapse"
It stops the shrinking (spread 0.97) and leaves a rank of 1.5: twelve numbers, about two things (§4).
"what is predictable is signal"
A drifting band is predictable and irrelevant: the predictive code reads it at 0.69, the reward code at 0.02 (§5).

Checkpoint exercise

Try it
A code number has variance 1 and correlates 0.9 with itself one picture later. (a) With the best predictor, what is its loss, and what share of its variance is unpredictable? (b) Multiply the number by 0.1: what are the loss and the share now? (c) What does that say about comparing losses between codes? Answer: (a) the best predictor is A = ρ = 0.9, the loss is 1 − ρ² = 0.19, and 19 % of the variance is unpredictable. (b) The loss becomes 0.1² × 0.19 = 0.0019, a hundred times smaller, and the unpredictable share is still 19 %. (c) A bare loss rewards shrinking the code; only a scale-free quantity such as that share can be compared between codes, which is why §3 fixes the scale and §4 guards it.

Where this points next

A state can now be asked for in the right way: keep what predicts, guard against keeping nothing, and let the task retire what persists without mattering. On the Courtyard's pictures that gives a 12-number code that reads the ball at r² = 0.97 through a band as bright as the ball, where rebuilding the picture reads 0.13. But every such code is a function of one picture: it reads the ball's velocity at 0.05 at best, and for 81 % of the pictures taken behind the curtain the picture before was blind too. A state has to be carried from picture to picture, and no machine of ours does it: lesson 2's filter was written by hand for a model we knew. How is the filter learned from data?

Takeaway
A state is sufficient if it makes the future independent of the history, and minimal if nothing smaller does. Asking a code to rebuild the picture pays in pixel variance and keeps the loudest nuisance: at η = 1 twelve components rebuild 86 % of the picture and read the ball at 0.13. Asking it to predict the next code pays for persistence; under unit variance the best code is the top eigenvectors of the lag-1 covariance, ranked by |ρ| (ball 0.97). The bare loss has a trivial minimum, since ℒ(cE) = c²ℒ(E), and gradient descent finds it (collapse: spread 0.0030); a variance term and a covariance term hold the spread at 0.95 and the rank at 12.0. A predictable nuisance survives any predictive objective (0.69 for a drifting band); only the task retires it, with a value-equivalent code (0.02), at the price of labels and of commitment to one question. Every such code is a function of one picture: no velocity (r² 0.05 at best), nothing behind the curtain.

Interview prompts

Companion reads: Computer Vision · 14 Self-supervised learning (the guards of §4), Lesson 19 · The prediction space is a bit budget (what the loss spends capacity on) and Lesson 20 · The tokenizer is the ceiling (what an encoder drops is gone).