The tokenizer is the ceiling
The tokenizer is usually treated as preprocessing — a thing you download and move past. But it runs before the dynamics model exists and it throws information away permanently. Whatever it discards, no later training run, no larger transformer, and no amount of post-training can ever recover.
1 · A tokenizer is a lossy codec with a job description
Strip away the terminology. A visual tokenizer takes a frame (or a short stack of frames) and returns a small grid of discrete symbols, each drawn from a learned codebook. Then a decoder turns symbols back into pixels. It is a lossy codec, trained with a reconstruction objective, and it has exactly three knobs:
Two of those knobs set the token count, which sets the dynamics model's entire cost structure:
tokens per frame = (res / p)² · tokens per second = (res/p)² · (fps / s)And because attention over a context is quadratic in tokens, halving the patch size quadruples the tokens and multiplies the attention cost by sixteen. That is why every practical system compresses aggressively. The question this lesson answers is: how aggressively can you compress before the physics you need stops existing?
2 · Convert the compression into millimetres
Here is the move that makes the ceiling concrete, and it is embarrassingly simple. A token summarises a p × p pixel patch. Whatever spatial structure lives inside that patch is represented only as an average — the codebook can say "edge, roughly here, roughly this orientation," not "the pin is 0.4 mm left of the hole." So the tokenizer imposes a blur floor: a smallest physical feature it can reliably distinguish.
Convert pixels to millimetres through the camera's field of view. If the frame's 256 pixels span 400 mm of workspace:
scale = 400 mm / 256 px = 1.5625 mm per pixel blur floor = p · scale = 16 · 1.5625 = 25 mmNow put that next to the task. Our running case is inserting a connector, and the clearance between pin and socket is about 2 mm. The tokenizer's floor is 25 mm. It is 12.5× too coarse to represent the distinction the task is entirely about.
3 · Three ways out, and their prices
Drag the field-of-view slider to 120 mm and the verdict flips. That is the first and best answer, and it is not a modelling change at all.
- Change the optics, not the network. A wrist camera looking at a 120 mm workspace gives 0.47 mm per pixel, so a patch-16 tokenizer has a 7.5 mm floor — still coarse, but drop to patch 4 and you are at 1.9 mm, under the clearance. Moving the camera closer is orders of magnitude cheaper than moving the loss. Most "we need a bigger model" conclusions in manipulation are actually "we need a wrist camera."
- Spend tokens non-uniformly. Nothing requires one patch size across the frame. Tokenize the gripper neighbourhood at patch 4 and the background at patch 32. Cost scales with the area you refine, and the contact region is a small fraction of the frame — this is the water-filling of lesson 19, implemented in the tokenizer instead of the loss.
- Take the decision-critical quantity out of the image entirely. The insertion clearance is a number. Joint encoders and force sensors report it at 1 kHz with micrometre precision and zero tokens. Feeding proprioception and force as their own channels alongside coarse visual tokens is usually the correct architecture, and it dissolves the ceiling rather than negotiating with it.
4 · Training the tokenizer: the two failures worth naming
If you train your own, two things go wrong reliably.
Codebook collapse. Vector-quantised tokenizers pass gradients through a nearest-neighbour lookup, and entries that are never selected receive no gradient — so they are never improved, so they are never selected. A 16,384-entry codebook can decay to a few hundred live entries, and your log₂K = 14 bits per token is really 8. Diagnose it directly: log codebook usage (fraction of entries hit per epoch) and perplexity of the code distribution, every run, always. Do not infer it from reconstruction loss, which stays deceptively fine. The standard mitigations — entries re-initialised toward unused encoder outputs, or replacing the learned codebook with a fixed scalar-quantisation grid so no entry can starve — both work by removing the winner-take-all dynamic.
Reconstruction is the wrong objective, slightly. A tokenizer trained purely to reconstruct will spend its capacity exactly the way lesson 19 predicted — on high-variance background. If you have downstream labels, add a term. Even a crude auxiliary loss (predict gripper-object distance from the tokens) reweights the codebook toward the physics, at negligible cost. You are moving ci; the mechanism does not matter much, only that you move it.
5 · The temporal knob is a separate trap
Spatial compression gets all the attention; temporal compression quietly decides whether events exist. A tokenizer with temporal stride s = 4 at 8 fps represents the world in 500 ms chunks. A slip lasts perhaps 30 ms. A contact transition lasts less. Both are averaged into their neighbourhood and cease to be observable.
The same arithmetic applies, and it is a Nyquist argument: to represent an event of duration τ you need an effective sample interval below τ/2.
s / fps < τ / 2 ⟹ for τ = 30 ms at fps = 8: s < 0.12Which is impossible — 8 fps cannot see a 30 ms event at any stride. The conclusion is not "increase the frame rate to 500 fps," which is unaffordable across the whole pipeline. It is the same conclusion as §3: fast contact events belong on a non-visual channel, sampled at the rate the physics demands, fused with visual tokens sampled at the rate the budget allows. Lesson 25 returns to this as a hierarchy question and lesson 10 derives the matching of each channel's rate to the decision's rate in full.
Where this points next
We now have a representation whose ceiling we have checked. Time to train the dynamics on top of it. And immediately there is a choice that looks like an implementation detail and is in fact the difference between a model that scores well and a model that works: at each training step, do you feed the model the real previous token or its own previous prediction? Lesson 21 shows that the convenient answer is the one that hides the only error a planner will ever encounter — and that a recent alternative gets rollout-grade robustness at teacher-forcing prices.
Interview prompts
- Compute the blur floor for patch 16 over a 400 mm field of view at 256 px, and interpret it. 400/256 = 1.5625 mm per pixel, times 16 = 25 mm — twelve times coarser than a 2 mm insertion clearance, so the task's decisive distinction is not representable. (§2)
- Why can no amount of downstream training fix a tokenizer ceiling? The tokenizer runs first and is lossy, so two physically different worlds map to identical tokens; the distinction is absent from the dynamics model's input. (§2)
- Rank the fixes for an insufficient blur floor by cost-effectiveness. Change the optics (wrist camera), then tokenize non-uniformly, then move the decision-critical quantity onto a non-visual channel, and only then enlarge the model. (§3)
- What is codebook collapse, and why won't reconstruction loss reveal it? Unselected codebook entries get no gradient so they never become selectable, shrinking effective bits per token; the surviving entries still reconstruct acceptably, so only usage and perplexity logs expose it. (§4)
- Why does halving patch size cost 16× in attention, not 4×? Tokens per frame scale as (res/p)², so halving p quadruples tokens, and attention is quadratic in tokens. (§1)
- Give the Nyquist argument for why an 8 fps pipeline cannot represent a 30 ms slip. Representing duration τ needs an effective interval below τ/2 = 15 ms, but 8 fps gives 125 ms even at stride 1, so the event is averaged away regardless of stride. (§5)
- A tokenizer trained purely for reconstruction under-represents contact. Give a cheap correction. Add an auxiliary head predicting a decision-relevant scalar such as gripper-object distance from the tokens, which reweights the codebook toward the physics at negligible cost. (§4)