Anatomy of a robot model: two clocks
Lesson 12 ended with a policy that starts nearly right, and a list of what it must do at once: read language and images, feel force, emit chunks, and still answer within one control step of 50 ms. This lesson prices the last demand. A model that reads costs 2 · parameters · tokens operations plus a pass over its weights for every token it writes, about 2.9 steps at 7 billion parameters, and a loop that acts on a reading d steps old either rings or learns of a shove late. The way out is two networks on two clocks: a slow reasoner hands a chunk of targets to a fast loop that feels the arm every step. The chunk must last at least the reasoner's latency plus its interval. What each part is trained on is the next lesson's problem.
New idea: a robot model runs on two clocks: a slow reasoner that looks every K steps and hands over a chunk of H targets L steps later (its latency), and a fast loop that feels the arm at every step; the chunk must last at least L + K steps, H ≥ L + K. What each clock handles is set by who can see the signal and how long it can wait.
Forces next: A large model that reads and reasons is too slow to steer an arm directly, so it runs slowly and hands a plan to a small fast network that runs at the control rate, and the split decides what each part must be trained on. Training the parts in order, on data of very different size and quality, has its own failure: fine-tuning on a few dozen tasks raises every number we track while erasing what the language model knew. In what order, on what mixture, should the pieces be trained, and how do we notice the loss?
1 · The price of a control step
A control step is 0.05 s. When something slow sits in the loop, lesson 6's delay appears: the command computed now comes from a reading d steps old, the error obeys et+1 = et − k·et−d (k is the share of the error one step removes), and the loop is stable only for k < 2 sin(π / (4d + 2)). The ceiling is 2.00 at d = 0, 1.00 at 1, 0.62 at 2, 0.45 at 3, 0.35 at 4, 0.24 at 6 and 0.15 at 10. The fast loop of this lesson takes k = 0.5 (a 7 cm shove is 1.75 cm two steps later), which is under the ceiling at d = 2 and over it from d = 3, a reading 150 ms old. Latency is paid twice: in the reaction, and in the stiffness the loop may have.
How old are the readings of real models? Published decision times, with lesson 1's planner, in steps of 50 ms:
| system | what runs | time per decision | steps |
|---|---|---|---|
| RT-2 (Brohan et al., 2023) | 55B at 1 to 3 Hz on cloud TPUs; 5B at about 5 Hz | 333 to 1000 ms; 200 ms | 6.7 to 20; 4 |
| OpenVLA, 7B (Kim et al., 2024) | about 6 Hz on one RTX 4090 | 167 ms | 3.3 |
| π0, 3.3B (Black et al., 2024) | on board, three cameras, a chunk of 50 | 73 ms | 1.5 |
| Gemini Robotics (2025) | cloud backbone, local action decoder | about 250 ms | 5 |
| the planner of lesson 1's exercise | 200 sequences × 25 steps × 0.1 ms | 500 ms | 10 |
None is under one step and most are several. "Answer within one control step" is a demand the models that read are not built to meet.
2 · The price of reading
A model with P parameters that reads N tokens (image patches and words) pushes each token through every weight, a multiply and an add: 2P operations per token, 2PN in all, which a device doing F operations a second takes 2PN/F to do. A model that writes its action as tokens, as RT-2 and OpenVLA do with 8 and 7 per action, must also read all of its weights again for each of the n tokens it writes, 2P bytes at two bytes a weight; at a batch of one that is almost no arithmetic per byte, so the time is set by the memory bandwidth B:
t = 2PN / F + n · 2P / B
Take a ruler of round numbers, not a product: F = 100 TFLOP/s, B = 1 TB/s, N = 250 tokens (three camera frames of 64 tokens each, the per-frame count of GR00T N1 and SmolVLA, and an instruction) and n = 8 tokens written. π0 and SmolVLA write no action tokens; for them the last columns are what their backbones would cost if they did.
| parameters | read, 2PN/F | per token written, 2P/B | read + 8 tokens | steps of 50 ms |
|---|---|---|---|---|
| 0.45 B (SmolVLA's size) | 2.25 ms | 0.90 ms | 9.5 ms | 0.19 |
| 3 B (π0's backbone) | 15 ms | 6.0 ms | 63 ms | 1.3 |
| 7 B (OpenVLA's size) | 35 ms | 14 ms | 147 ms | 2.9 |
The ruler is round numbers, not a calibration, and it runs under the printed times: 147 ms where OpenVLA's 6 Hz is 167, and 6 ms for π0's ten flow steps through its 0.3 B action expert, the small network beside the backbone that writes the chunk (0.6 GB read per step, 0.6 ms at 1 TB/s), where Black et al. print 27. Turned around, it gives the largest model that fits a clock of τ seconds, P ≤ τ / (2N/F + 2n/B): 2.4 B at 50 ms (20 Hz), 0.95 B at 20 ms (50 Hz), 0.24 B at 5 ms (200 Hz) and 0.05 B at 1 ms. Writing a whole chunk in ten flow steps, each over all 50 action tokens at once, removes most of the second term: Driess et al. (2025) put an autoregressive VLA at 1.3 Hz and π0 at 10 Hz.
Does reading need the billions? RT-2 trained a 5B model on its robot data from scratch and it scored 9 % on unseen objects, backgrounds and environments; started from a web-pretrained model and fine-tuned on the same robot data it scored 42 %, and the 55B model 52 % (Brohan et al., 2023). The generalisation came with the web pretraining and grew with the size, and the size sets the clock.
3 · One model on one clock, three ways to fail
Put the question on the Bench: the five-post course with lesson 1's gust. The policy is lesson 6's, a chunk of absolute joint targets along the path, tracked from the joints it reads at every step with k = 0.5. A run ends at the mat, at a post, or at step 235 (a timeout).
Two things go wrong, each when the cup is 8 to 16 cm (0.5 to 1.1 s) before the middle post. A shove moves the cup 7 cm toward the post, as a bump would: the path passes 10 cm from the post's centre, so unless something pulls the cup back it passes 3 cm from it, and cup and post touch when their centres are closer than 5 cm, which leaves 5 cm of clearance. A moved vase moves the post 7 cm into the path, and the expert's path bends away from its new place. The joints feel the shove at the next step; only a model that looks at the scene sees the vase.
The reasoner on the Bench is a stand-in, the expert's own path planner, so that what is measured is when a chunk arrives and not how it was learned. Each cell is 300 runs under the same luck, good to within 5.7 points at 95 %, at a latency of L = 6 steps (300 ms, between Gemini Robotics and RT-2 55B):
| the model | a shove | a vase moves | nothing happens | copies of the reasoner |
|---|---|---|---|---|
| one model in the loop, naive: it acts on its stale reading | 0 % | 0 % | 0 % | 6 |
| one model in the loop, best case: it adds up its own commands | 74 % | 75 % | 100 % | 6 |
| fast loop alone, on the path it was trained on | 100 % | 0 % | 100 % | 0 |
| two clocks (look every 4 steps, chunk of 12) | 100 % | 66 % | 100 % | 2 |
The model in the loop, naively. It acts on a reading L steps old as if it were current: lesson 6's loop with d = 6, whose ceiling is 0.24, at a stiffness of 0.5. It rings, and it fails even when nothing happens: 100 % at L = 2, 0 % at every latency from 3 to 8 steps with every run ending at a post, and at most 24 % beyond, where it flails. The model in the loop, at its best. Let it add up the commands it has issued since its reading and it is stable at every latency. It stays blind to what it did not cause: it learns of a shove L steps after the joints feel it, L × 0.05 s × 15 cm/s = 4.5 cm of travel against 5 cm of clearance, and it survives 74 % of the shoves. It also decides at every step and each decision takes L steps, so L decisions are always in flight: L copies of the model run at once. The fast loop alone. A network small enough for its clock (§2) feels every shove. It does not read the scene, so a vase that has moved is news it never gets: 0 %.
Each failure has a cause. Rows one and two act on a reading that is old; row three has a loop that cannot read. The cure for the first is a loop whose own readings are never old, and for the second something that reads. One network cannot be both, and two can.
4 · Two clocks, and the rule the chunk obeys
Give what feels the arm its clock, and what reads its own. Real systems call the parts a backbone and an action expert; here they are the reasoner and the fast loop. The reasoner looks at the scene every K steps; what it saw becomes a chunk of H targets along the path of the course as it saw it, and the chunk lands L steps after the look. The fast loop runs at every step on fresh joints: it takes the next target of the newest chunk to have landed, commands k times the distance to it, and holds when no chunk is usable. Its own reading is never old, so §1's ceiling is not in play: the reasoner's latency costs the plan its age, not the loop its stability.
A chunk made at step s is usable from s + L, its first L targets having fallen due while it was made, to s + H − 1, and the next chunk lands at s + K + L. The steps from s + H to s + K + L − 1 have no usable chunk, so there is no gap exactly when
H ≥ L + K
in steps: the chunk must last at least the reasoner's latency plus its interval. Otherwise L + K − H steps of every K are spent holding, the arm moves for a share (H − L)/K of the time, and a run that needs 145.5 steps at full speed needs 145.5 · K/(H − L).
Two more consequences. A change in the world is first seen at the next look, 0 to K − 1 steps away and equally likely, and acted on L steps after that: the arm acts on a picture L + (K − 1)/2 steps old on average, and travels blind meanwhile. And the reasoner is busy L steps of every K, so it needs ⌈L/K⌉ copies running, against L for a model that decides at every step.
The widget
What to try. At the defaults (a vase moves, two clocks, L = 4, K = 4, H = 12, 15 cm/s) 84.0 % of the 300 runs reach the mat (interval 79.4 to 87.7 %); the event is acted on 5.4 steps (269 ms) after it happens, 4.0 cm of blind travel against 5 cm of clearance, and the chunk margin is +4. Move L first, to 6 and to 8: 66.3 % and 44.7 % of the vases are survived, and the margin falls to +2 and 0. Back at 4, switch the event to a shove: 100 %, the fast loop feels it at once. Switch the design to one model in the loop: shoves fall to 92.0 % with 4 copies running, and vases do better, 94.7 %, because there is no K to wait for. The fast loop alone: shoves 100 %, vases 0 %. The naive loop, with nothing happening: 100 % at L = 2 and 0 % at L = 3, every run ending at a post. Back to two clocks and the vase, set L = 10 (500 ms, the planner of lesson 1) and H = 18: 100 % of the shoves and 25.3 % of the vases survive, against 44.7 % and 38.7 % for the model in the loop. Back at L = 4, vases and H = 16, slide K to 1 and to 8: 97.0 % with 4 copies running, 84.0 % at 4, 63.3 % at 8 with one copy and 2.5 passes a second; every look skipped is a step more of blind travel. Set the event to nothing, K = 8, and shorten H: at 12 there is no gap and a run takes 145.5 steps; at 10 the gap is 2 steps of every 8, 48.5 steps are held per run and 193.6 steps are needed (the rule says 194); at 8 every run times out, 0 % reach the mat. Last, with the vase and the defaults, raise the arm speed: 96.7 % at 10 cm/s, 63.7 % at 20 and 22.0 % at 30.
5 · What each clock is for
The deadline of a signal is a distance: how far the cup can travel on a stale picture before the event is lost. With the model in the loop, where the age is exactly L, success on the vase falls to 90 % at an age of 7.4 steps at 10 cm/s, 4.5 at 15 and 3.3 at 20: 3.7, 3.4 and 3.3 cm of travel, two thirds to three quarters of the 5 cm clearance each time. A faster arm or a tighter clearance shrinks the budget in steps in proportion; at 1 m/s a reasoner a quarter of a second late has been blind for 25 cm. Two clocks spend the same budget on L + (K − 1)/2 and not on L: at K = 4 the 90 % point is at L = 3.3, a mean age of 4.8 steps, so the curve moves left by about a step.
That sorts the signals. One whose deadline is shorter than L cannot go through the reasoner, whatever K is; one longer than L + K can. What the fast loop can do alone it must: the shove, and contact. What only the reasoner understands rides the chunk: language, which object, which vase, which subtask. Real systems split the same way, and where the numbers are printed the rule can be checked:
| system | reasoner and action part | what the numbers say |
|---|---|---|
| π0 (Black et al., 2024) | one 3.3B pass makes a chunk of 50, run open loop, 16 actions at 20 Hz or 25 at 50 Hz before the next look | 73 ms is 1.5 steps at 20 Hz: 17.5 needed of 50, 32.5 to spare; at 50 Hz 3.65 + 25 = 28.65, 21.35 to spare. Only 32 % or 50 % of a chunk is run. |
| GR00T N1 (NVIDIA, 2025) | a 1.34B model reads at 10 Hz; a flow-matching network samples a chunk of 16 in 63.9 ms | 15.6 chunks and 250 actions a second, against the 20 a second of its own teleoperated demonstrations |
| Helix (Figure, vendor-reported, 2025) | a 7B model at 7 to 9 Hz, an 80M network at 200 Hz | 22 to 29 fast steps per slow update |
| Gemini Robotics (2025) | cloud backbone under 160 ms, local decoder, about 250 ms end to end, 50 Hz effective | a chunk of at least 12.5 actions; its length is not printed |
| SmolVLA (Shukor et al., 2025) | 0.45B, a chunk of 50 at 30 Hz (1.67 s) | asynchronous inference, the next chunk requested before the last runs out, scored 73.3 % on its three real tasks against 78.3 % synchronously and finished a pick-and-place 30 % faster. In simulation, running all 50 of 50 actions before looking again scored 51.8 % on LIBERO, against 80.3 % for looking after every action. |
It is one bargain: the more often the reasoner looks, the fresher the picture and the more copies it takes; the less often, the longer the chunk it must leave behind.
6 · What the split asks of training
The split fixes what each part must be trained on, and the two are nothing alike. The reasoner learns to read, from data that mostly come from somewhere else: RT-2 (Brohan et al., 2023) filtered about 10B image-text pairs to a billion training examples, and π0.5 (Physical Intelligence, 2025) reports that 97.6 % of the examples of its first phase do not come from the mobile robots doing household tasks. The action network learns to act, from robot data that costs robot time (over 10,000 hours for π0, Black et al., 2024), and it is small: 9.1 % of π0's parameters, 1.1 % of Helix's, on its maker's figures. The hand-off decides how far they can be trained apart. π0.5 first writes the subtask as text ("pick up the plate"), and its mix has data that teaches only that. π0 hands over features and the action loss reaches back into the backbone; Helix, by its maker's account, does the same; GR00T N1 keeps its language model frozen and trains the vision encoder and the action network (NVIDIA, 2025). Driess et al. (2025) find that adding a randomly initialised action expert naively harms both training speed and the transfer of knowledge.
And the numbers we track will not see it. The shove column of §3 reads 100 % with a working reasoner and 100 % with none; it takes a probe that needs the reasoner, here the moved vase, 66 % against 0 %, to tell them apart. The real cases look alike. A pre-trained GR00T N1 checkpoint scored 76.6 % on a left-to-right hand-over; after post-training on right-hand demonstrations the policy had lost it (NVIDIA, 2025). Removing π0.5's web data left its full-task benchmark within noise and hurt following language about new objects (Physical Intelligence, 2025). A model trained in parts can lose what one part knew while every number we watch holds.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
A model that reads is too slow to steer an arm: at a latency of 6 steps the best whole-model loop keeps 74 % of the shoves, and the fast loop alone feels everything and sees nothing. Two clocks give each what it can do, a chunk that lasts L + K and a loop at every step, for 2 copies of the reasoner instead of 6. But the parts differ in size (a 0.3B expert beside a 3B backbone in π0) and in what an example costs, the action loss reaches the reasoner through the hand-off, and a number like 100 % under shoves cannot tell a reasoner that works from one that has been written over; GR00T N1 lost a skill its pre-trained checkpoint had. In what order, on what mixture, should the pieces be trained, and how do we notice the loss?
Interview prompts
- Why can a 7B model that writes its action as tokens not sit in a 20 Hz feedback loop? (§1, §2 — it takes about 2.9 steps on a round-number ruler, and a loop acting on a reading d steps old is stable only for k below 2 sin(π / (4d + 2)).)
- Why does reading cost operations and writing cost bandwidth? (§2 — reading pushes N tokens through every weight, 2PN operations at the device's rate; each token written needs all 2P bytes of weights once, and at a batch of one that is bound by memory.)
- State the rule for the length of a chunk. (§4 — the chunk must last at least the latency plus the interval, H ≥ L + K; otherwise L + K − H steps of every K are spent holding.)
- A reasoner takes 180 ms at 50 Hz and looks every 25 steps, and its chunk has 50 steps. Is the rule met? (§4, §5 — L is 9 steps, so L + K = 34 against 50: yes, with 16 steps to spare.)
- What does the fast loop do with a stale plan? (§3 — it tracks it faithfully: with no reasoner no moved vase is survived; staleness is the reasoner's cost, not the loop's.)
- What does looking less often cost, and what does it save? (§4 — the picture is L + (K − 1)/2 steps old on average, and the reasoner needs ⌈L/K⌉ copies instead of L.)
- Why is the deadline of a signal a distance and not a time? (§5 — what is lost is clearance, which the arm uses up at its speed: about 3.5 cm of stale travel before one event in ten is lost, at 10 to 20 cm/s.)
- How can fine-tuning raise every number you watch and still damage the model? (§6 — the watched tasks exercise the fast loop; the action loss reaches the reasoner through the hand-off and nothing watched measures it.)
Companion reads: System ML · 15 Disaggregated prefill / decode (reading is compute-bound, writing is memory-bound), World Models · 29 Distilling into a real-time loop (latency as part of the specification) and World Models · 09 Plan with it (the slow planner of the world-model track).