all_lessons/ World Models/ 19 · the prediction spacelesson 19 / 31

The prediction space is a bit budget

"Should I predict pixels, tokens, or latents?" sounds like an architecture question. It is an accounting question. Your model has a finite number of bits to spend describing the future, and the loss function — not you — decides where they go. Here is the calculation that shows where they actually go.

Where we are
Lesson 18 fixed acceptance: controllability, consistency, calibration, inside a latency budget. Each of those is stated about the error in some quantity. Now we choose that quantity. The choice looks stylistic until you notice that every representation has the same fixed capacity and they differ only in how they allocate it.
Forced by 18Acceptance is measured on a quantity we have not yet chosen. This stepRun the rate–distortion allocation: a pixel-weighted loss starves the decision-relevant bits. Forces 20If we predict discrete tokens, the tokenizer becomes a hard ceiling we must design.

1 · A model has a bit rate, whether or not you chose one

Start with the picture. A camera is pointed at a bench: a wall behind, a connector in a gripper, and a socket two millimetres wider than the pin. Now imagine you are allowed to describe the next frame in a fixed number of words — say two hundred. How many do you spend on the wall? Because the wall is most of the image, and the two millimetres are the whole task, and a model that is scored on how well its description matches the picture will spend almost all two hundred words on the wall. This lesson is that intuition made exact.

Start from something uncontroversial. Your model, at some layer, carries a finite description of the future — a fixed number of activations at finite precision, or a fixed number of discrete tokens from a fixed codebook. Call the total B bits per frame. It is finite, it is not very large, and it is the same B no matter which representation you picked.

Now the only question that matters: how does that B get divided among the things in the scene? And the answer is: by whatever the loss weights. Not by importance. Not by your intentions. By the loss.

2 · Split the scene into three honest components

Take the running case for this whole part: a gripper inserting a connector into a socket. Divide the frame into three groups, and give each the two numbers that decide its fate.

componentdimension nvariance σ²decision weight ω
background — wall, bench, shadows, lighting7001.000.02
object appearance — the connector's body, texture, colour2200.420.28
contact and pose — the last few millimetres of alignment400.080.70

Look at the shape of that table, because it is the shape of nearly every embodied scene. Variance and decision-relevance are anti-correlated. The background has the most pixels and the most variance and decides almost nothing. The contact region has the fewest pixels and the least variance and decides everything — whether the insertion succeeds is a question about roughly 2 millimetres.

3 · Now do the allocation properly

This is a textbook problem: allocate rate across independent Gaussian components to minimise weighted distortion. With rate Ri bits per coefficient, the achievable distortion of component i is

Di(Ri) = σi² · 2−2Ri

Each bit buys a factor of 4 reduction — the familiar 6 dB per bit. We minimise a weighted total subject to the budget:

min Σi ni ci σi² 2−2Ri   s.t.   Σi ni Ri = B,   Ri ≥ 0

Set the derivatives equal through a Lagrange multiplier and something clean falls out. The optimum equalises the weighted distortion across components:

ci σi² 2−2Ri = θ   ⟹   Ri = max( 0, ½ log₂( ci σi² / θ ) )

with θ chosen so the budget is exactly spent. This is reverse water-filling, and the weight ci is the whole story:

pixel-weighted loss
ci = 1
All coefficients equally important. Bits flow to whatever has the most variance — the background.
decision-weighted loss
ci = ωi
Bits flow to whatever changes the decision. The 40 contact coefficients get funded.

Put in the numbers. Under ci=1, the term σi² alone decides the ranking, and contact has the smallest variance of the three — so it is the first component driven to Ri=0 when the budget tightens. The loss does not neglect contact by accident. Under a pixel-weighted objective, starving contact is the optimal thing to do, and your optimiser is good at its job.

Where a pixel loss actually spends its bits
Solves the water-filling equation above for θ and reports the resulting bits per coefficient. Switch the weighting and watch the contact bar appear from nothing. The two loss readouts show why the pixel objective is not "wrong" — it is winning its own game while losing yours.
Interactive rate allocation across three scene components.
background
—
contact
—
decision loss
—
pixel loss
—
diagnosis
—

At B = 1200 under pixel weighting the contact group receives zero rate — not a small share, none at all, because water-filling drives the lowest-ciσi² component below the threshold θ first — while the background takes about 85% of the bits. Switch to decision weighting: contact gets funded, the decision loss drops by more than half, and the pixel loss gets visibly worse. That trade — accept a worse-looking video to get a usable simulator — is the correct trade, and it is the one every practitioner has to make consciously because no metric will make it for them.

4 · The three prediction spaces, priced

Now the original question answers itself. The spaces differ in where the reweighting can happen.

pixelsdiscrete tokenslatents (no decoder)
who sets cithe pixel metric — you fight itthe tokenizer, at training timethe encoder's own objective
contact bitsstarved by constructionwhatever the codebook resolveswhatever the encoder kept
can you look at ityes, directlyyes, decodeno — no pixel view exists
cost per framehighest — full decodemoderatelowest
failure you will hitblurry, uncontrollable, expensivethe codebook's ceiling (lesson 20)collapse — a constant latent predicts perfectly

The latent option deserves its warning label. Predicting in a learned latent space with no reconstruction term is the cheapest and sharpest choice — this is the JEPA family, and V-JEPA 2 shows it scales to over a million hours of video. But it removes the only cheap sanity check you had: you cannot look at a latent and see that the mug is on the floor. And it admits a degenerate solution the pixel objective never did — if the encoder maps everything to a constant, prediction is trivially perfect. Preventing collapse becomes a first-class training concern rather than an afterthought, and your acceptance tests from lesson 18 become your only evidence that anything was learned.

The choice is not permanent, and should not be uniform
Real systems mix. Predict in a latent for the dynamics rollout (cheap, sharp, many steps), decode to pixels only at the few frames a human or a detector needs to inspect, and keep an explicit low-dimensional head for the decision-critical quantities — contact state, object pose, gripper force — so their bits are funded by construction instead of by hope. That last head is the cheapest fix in this lesson: 40 coefficients with ω=0.70, given their own loss term, cannot be starved.

5 · What to actually do on Monday

  1. Write the three-column table for your own scene. Dimension, variance, decision weight. Estimate the weights by ablation: corrupt one group in the input, measure the drop in task success. Ten minutes of work that reorders your entire roadmap.
  2. Compute the implied allocation. Run the water-filling formula with ci=1. If the decision-critical group lands under 5% of the budget, you have found your bug and it is not in the model code.
  3. Add an explicit head for the decision-critical quantities. Not a reweighted pixel loss — a separate output with its own target. This is the highest return-per-line change available to you.
  4. Choose the rollout space for cost, the inspection space for debuggability. They do not have to be the same space, and in production they usually are not.
  5. Re-run the lesson-01 acceptance tests. Specifically calibration, which is the one that catches a collapsed latent.
Takeaway
Every prediction space has the same finite bit budget; they differ only in who decides the allocation. Reverse water-filling gives the optimum, Ri = max(0, ½ log₂(ciσi²/θ)), and shows that under a pixel-weighted loss (ci=1) the ranking is set by variance alone — so the low-variance, high-decision-weight contact region is the first thing defunded. Fix it by moving ci: a tokenizer that resolves what matters, a latent objective that keeps it, or — cheapest and most reliable — an explicit head for the decision-critical quantities.

Where this points next

Suppose we take the middle column and predict discrete tokens. Then the tokenizer is the thing that fixed ci before the dynamics model ever ran, and no amount of later training can recover a distinction the codebook threw away. Lesson 20 makes that ceiling quantitative: given a patch size and a field of view, we compute the smallest physical feature the model is capable of representing, and check it against the 2 millimetres our running task actually needs.

Interview prompts