The prediction space is a bit budget
"Should I predict pixels, tokens, or latents?" sounds like an architecture question. It is an accounting question. Your model has a finite number of bits to spend describing the future, and the loss function — not you — decides where they go. Here is the calculation that shows where they actually go.
1 · A model has a bit rate, whether or not you chose one
Start with the picture. A camera is pointed at a bench: a wall behind, a connector in a gripper, and a socket two millimetres wider than the pin. Now imagine you are allowed to describe the next frame in a fixed number of words — say two hundred. How many do you spend on the wall? Because the wall is most of the image, and the two millimetres are the whole task, and a model that is scored on how well its description matches the picture will spend almost all two hundred words on the wall. This lesson is that intuition made exact.
Start from something uncontroversial. Your model, at some layer, carries a finite description of the future — a fixed number of activations at finite precision, or a fixed number of discrete tokens from a fixed codebook. Call the total B bits per frame. It is finite, it is not very large, and it is the same B no matter which representation you picked.
Now the only question that matters: how does that B get divided among the things in the scene? And the answer is: by whatever the loss weights. Not by importance. Not by your intentions. By the loss.
2 · Split the scene into three honest components
Take the running case for this whole part: a gripper inserting a connector into a socket. Divide the frame into three groups, and give each the two numbers that decide its fate.
| component | dimension n | variance σ² | decision weight ω |
|---|---|---|---|
| background — wall, bench, shadows, lighting | 700 | 1.00 | 0.02 |
| object appearance — the connector's body, texture, colour | 220 | 0.42 | 0.28 |
| contact and pose — the last few millimetres of alignment | 40 | 0.08 | 0.70 |
Look at the shape of that table, because it is the shape of nearly every embodied scene. Variance and decision-relevance are anti-correlated. The background has the most pixels and the most variance and decides almost nothing. The contact region has the fewest pixels and the least variance and decides everything — whether the insertion succeeds is a question about roughly 2 millimetres.
3 · Now do the allocation properly
This is a textbook problem: allocate rate across independent Gaussian components to minimise weighted distortion. With rate Ri bits per coefficient, the achievable distortion of component i is
Di(Ri) = σi² · 2−2RiEach bit buys a factor of 4 reduction — the familiar 6 dB per bit. We minimise a weighted total subject to the budget:
min Σi ni ci σi² 2−2Ri s.t. Σi ni Ri = B, Ri ≥ 0Set the derivatives equal through a Lagrange multiplier and something clean falls out. The optimum equalises the weighted distortion across components:
ci σi² 2−2Ri = θ ⟹ Ri = max( 0, ½ log₂( ci σi² / θ ) )with θ chosen so the budget is exactly spent. This is reverse water-filling, and the weight ci is the whole story:
Put in the numbers. Under ci=1, the term σi² alone decides the ranking, and contact has the smallest variance of the three — so it is the first component driven to Ri=0 when the budget tightens. The loss does not neglect contact by accident. Under a pixel-weighted objective, starving contact is the optimal thing to do, and your optimiser is good at its job.
At B = 1200 under pixel weighting the contact group receives zero rate — not a small share, none at all, because water-filling drives the lowest-ciσi² component below the threshold θ first — while the background takes about 85% of the bits. Switch to decision weighting: contact gets funded, the decision loss drops by more than half, and the pixel loss gets visibly worse. That trade — accept a worse-looking video to get a usable simulator — is the correct trade, and it is the one every practitioner has to make consciously because no metric will make it for them.
4 · The three prediction spaces, priced
Now the original question answers itself. The spaces differ in where the reweighting can happen.
| pixels | discrete tokens | latents (no decoder) | |
|---|---|---|---|
| who sets ci | the pixel metric — you fight it | the tokenizer, at training time | the encoder's own objective |
| contact bits | starved by construction | whatever the codebook resolves | whatever the encoder kept |
| can you look at it | yes, directly | yes, decode | no — no pixel view exists |
| cost per frame | highest — full decode | moderate | lowest |
| failure you will hit | blurry, uncontrollable, expensive | the codebook's ceiling (lesson 20) | collapse — a constant latent predicts perfectly |
The latent option deserves its warning label. Predicting in a learned latent space with no reconstruction term is the cheapest and sharpest choice — this is the JEPA family, and V-JEPA 2 shows it scales to over a million hours of video. But it removes the only cheap sanity check you had: you cannot look at a latent and see that the mug is on the floor. And it admits a degenerate solution the pixel objective never did — if the encoder maps everything to a constant, prediction is trivially perfect. Preventing collapse becomes a first-class training concern rather than an afterthought, and your acceptance tests from lesson 18 become your only evidence that anything was learned.
5 · What to actually do on Monday
- Write the three-column table for your own scene. Dimension, variance, decision weight. Estimate the weights by ablation: corrupt one group in the input, measure the drop in task success. Ten minutes of work that reorders your entire roadmap.
- Compute the implied allocation. Run the water-filling formula with ci=1. If the decision-critical group lands under 5% of the budget, you have found your bug and it is not in the model code.
- Add an explicit head for the decision-critical quantities. Not a reweighted pixel loss — a separate output with its own target. This is the highest return-per-line change available to you.
- Choose the rollout space for cost, the inspection space for debuggability. They do not have to be the same space, and in production they usually are not.
- Re-run the lesson-01 acceptance tests. Specifically calibration, which is the one that catches a collapsed latent.
Where this points next
Suppose we take the middle column and predict discrete tokens. Then the tokenizer is the thing that fixed ci before the dynamics model ever ran, and no amount of later training can recover a distinction the codebook threw away. Lesson 20 makes that ceiling quantitative: given a patch size and a field of view, we compute the smallest physical feature the model is capable of representing, and check it against the 2 millimetres our running task actually needs.
Interview prompts
- State the reverse water-filling solution and say what the weight ci represents. R_i = max(0, ½log₂(c_iσ_i²/θ)) with θ set by the budget; c_i is how much you care about that component's error, and it is 1 under a pixel loss. (§3)
- Why does a pixel loss starve contact specifically, rather than at random? With c_i=1 the allocation is ordered by variance alone, and the contact region has the smallest variance of the three groups, so it is the first driven to zero rate. (§3)
- What structural property of embodied scenes makes this a general problem? Variance and decision-relevance are anti-correlated: the background has most of the pixels and variance and decides nothing; contact has least of both and decides everything. (§2)
- Give one advantage and one danger of predicting in a latent space with no decoder. Cheapest and sharpest, and it scales — but you lose visual inspection and admit representational collapse, where a constant encoder makes prediction trivially perfect. (§4)
- What is the cheapest reliable fix for a starved decision-critical group? Give it an explicit low-dimensional head with its own loss term, so its bits are funded by construction rather than competing under a pixel-weighted objective. (§4, §5)
- How would you estimate the decision weights ωi empirically? Ablate: corrupt each group in the input and measure the resulting drop in downstream task success. (§5)
- Why might the rollout space and the inspection space differ in production? Rollout is priced on cost per step over many steps, inspection on debuggability at a few frames, so a latent rollout with occasional decoding optimises both. (§4)