all_lessons/ World Models/ 20 · the tokenizer is the ceilinglesson 20 / 31

The tokenizer is the ceiling

The tokenizer is usually treated as preprocessing — a thing you download and move past. But it runs before the dynamics model exists and it throws information away permanently. Whatever it discards, no later training run, no larger transformer, and no amount of post-training can ever recover.

Where we are
Lesson 19 showed that the loss weighting decides which parts of the world get bits, and that predicting discrete tokens moves that decision into the tokenizer. So the tokenizer is now the component holding the ci we care about. This lesson computes what it costs and what it forbids — in millimetres, against a real task.
Forced by 19Predicting tokens delegates the bit allocation to the tokenizer. This stepConvert patch size into a physical blur floor and compare it to the task's clearance. Forces 21With a representation fixed, we can finally train dynamics — and the schedule is the next trap.

1 · A tokenizer is a lossy codec with a job description

Strip away the terminology. A visual tokenizer takes a frame (or a short stack of frames) and returns a small grid of discrete symbols, each drawn from a learned codebook. Then a decoder turns symbols back into pixels. It is a lossy codec, trained with a reconstruction objective, and it has exactly three knobs:

spatial compression
patch size p
How many pixels collapse into one symbol. Sets the spatial resolution of everything downstream.
temporal compression
stride s
How many frames collapse together. Sets whether fast events survive at all.
codebook
size K
How many distinguishable states one symbol can name. log₂K bits per token.

Two of those knobs set the token count, which sets the dynamics model's entire cost structure:

tokens per frame = (res / p)²   ·   tokens per second = (res/p)² · (fps / s)

And because attention over a context is quadratic in tokens, halving the patch size quadruples the tokens and multiplies the attention cost by sixteen. That is why every practical system compresses aggressively. The question this lesson answers is: how aggressively can you compress before the physics you need stops existing?

2 · Convert the compression into millimetres

Here is the move that makes the ceiling concrete, and it is embarrassingly simple. A token summarises a p × p pixel patch. Whatever spatial structure lives inside that patch is represented only as an average — the codebook can say "edge, roughly here, roughly this orientation," not "the pin is 0.4 mm left of the hole." So the tokenizer imposes a blur floor: a smallest physical feature it can reliably distinguish.

Convert pixels to millimetres through the camera's field of view. If the frame's 256 pixels span 400 mm of workspace:

scale = 400 mm / 256 px = 1.5625 mm per pixel blur floor = p · scale = 16 · 1.5625 = 25 mm

Now put that next to the task. Our running case is inserting a connector, and the clearance between pin and socket is about 2 mm. The tokenizer's floor is 25 mm. It is 12.5× too coarse to represent the distinction the task is entirely about.

What this means, stated plainly
No dynamics model trained on these tokens can learn that the insertion succeeds when the pin is within 2 mm and fails when it is 3 mm off — because in this representation those two worlds are the same token. Train for a year. Add parameters. Add data. Post-train on a million careful demonstrations. The distinction is not in the input. This is the single most common silent failure in visual world-model projects, and it is diagnosable in one line of arithmetic before you write any training code.
Compression against the physics you need
Cyan is tokens per second (your entire cost structure); red is the blur floor in millimetres; green is the 2 mm clearance the insertion actually requires. The verdict tells you whether the task is representable at all at this compression — before any training.
Interactive tokenizer compression versus physical resolution.
tokens / frame
—
tokens / s
—
bit rate
—
blur floor
—
is the task expressible
—

3 · Three ways out, and their prices

Drag the field-of-view slider to 120 mm and the verdict flips. That is the first and best answer, and it is not a modelling change at all.

  1. Change the optics, not the network. A wrist camera looking at a 120 mm workspace gives 0.47 mm per pixel, so a patch-16 tokenizer has a 7.5 mm floor — still coarse, but drop to patch 4 and you are at 1.9 mm, under the clearance. Moving the camera closer is orders of magnitude cheaper than moving the loss. Most "we need a bigger model" conclusions in manipulation are actually "we need a wrist camera."
  2. Spend tokens non-uniformly. Nothing requires one patch size across the frame. Tokenize the gripper neighbourhood at patch 4 and the background at patch 32. Cost scales with the area you refine, and the contact region is a small fraction of the frame — this is the water-filling of lesson 19, implemented in the tokenizer instead of the loss.
  3. Take the decision-critical quantity out of the image entirely. The insertion clearance is a number. Joint encoders and force sensors report it at 1 kHz with micrometre precision and zero tokens. Feeding proprioception and force as their own channels alongside coarse visual tokens is usually the correct architecture, and it dissolves the ceiling rather than negotiating with it.
Note the order
These are ranked by cost-effectiveness, and the ranking is stable across projects: optics → non-uniform tokens → non-visual channels → bigger model. "Bigger model" is last because it is the only one that does not change what information is present.

4 · Training the tokenizer: the two failures worth naming

If you train your own, two things go wrong reliably.

Codebook collapse. Vector-quantised tokenizers pass gradients through a nearest-neighbour lookup, and entries that are never selected receive no gradient — so they are never improved, so they are never selected. A 16,384-entry codebook can decay to a few hundred live entries, and your log₂K = 14 bits per token is really 8. Diagnose it directly: log codebook usage (fraction of entries hit per epoch) and perplexity of the code distribution, every run, always. Do not infer it from reconstruction loss, which stays deceptively fine. The standard mitigations — entries re-initialised toward unused encoder outputs, or replacing the learned codebook with a fixed scalar-quantisation grid so no entry can starve — both work by removing the winner-take-all dynamic.

Reconstruction is the wrong objective, slightly. A tokenizer trained purely to reconstruct will spend its capacity exactly the way lesson 19 predicted — on high-variance background. If you have downstream labels, add a term. Even a crude auxiliary loss (predict gripper-object distance from the tokens) reweights the codebook toward the physics, at negligible cost. You are moving ci; the mechanism does not matter much, only that you move it.

5 · The temporal knob is a separate trap

Spatial compression gets all the attention; temporal compression quietly decides whether events exist. A tokenizer with temporal stride s = 4 at 8 fps represents the world in 500 ms chunks. A slip lasts perhaps 30 ms. A contact transition lasts less. Both are averaged into their neighbourhood and cease to be observable.

The same arithmetic applies, and it is a Nyquist argument: to represent an event of duration τ you need an effective sample interval below τ/2.

s / fps < τ / 2   ⟹   for τ = 30 ms at fps = 8:   s < 0.12

Which is impossible — 8 fps cannot see a 30 ms event at any stride. The conclusion is not "increase the frame rate to 500 fps," which is unaffordable across the whole pipeline. It is the same conclusion as §3: fast contact events belong on a non-visual channel, sampled at the rate the physics demands, fused with visual tokens sampled at the rate the budget allows. Lesson 25 returns to this as a hierarchy question and lesson 10 derives the matching of each channel's rate to the decision's rate in full.

Takeaway
A tokenizer is a lossy codec whose patch size imposes a physical blur floor — p × (FOV/res) millimetres — and whose temporal stride imposes a Nyquist limit on which events exist. Compute both before training: patch 16 over a 400 mm field of view gives a 25 mm floor against a 2 mm clearance, a 12.5× shortfall that no later training can repair. Fix it in cost order: optics first, then non-uniform tokenization, then non-visual channels, and only then a bigger model. If you train the tokenizer yourself, log codebook usage and perplexity every run — collapse is invisible in reconstruction loss.

Where this points next

We now have a representation whose ceiling we have checked. Time to train the dynamics on top of it. And immediately there is a choice that looks like an implementation detail and is in fact the difference between a model that scores well and a model that works: at each training step, do you feed the model the real previous token or its own previous prediction? Lesson 21 shows that the convenient answer is the one that hides the only error a planner will ever encounter — and that a recent alternative gets rollout-grade robustness at teacher-forcing prices.

Interview prompts