all_lessons/Robot Model Training/13 · Two clockslesson 13 / 24

Anatomy of a robot model: two clocks

Lesson 12 ended with a policy that starts nearly right, and a list of what it must do at once: read language and images, feel force, emit chunks, and still answer within one control step of 50 ms. This lesson prices the last demand. A model that reads costs 2 · parameters · tokens operations plus a pass over its weights for every token it writes, about 2.9 steps at 7 billion parameters, and a loop that acts on a reading d steps old either rings or learns of a shove late. The way out is two networks on two clocks: a slow reasoner hands a chunk of targets to a fast loop that feels the arm every step. The chunk must last at least the reasoner's latency plus its interval. What each part is trained on is the next lesson's problem.

The thesis, here
Understanding and feeling have different deadlines. Understanding (language, which object, a moved vase) needs billions of parameters and so tens to hundreds of milliseconds; feeling (a shove, a contact) must be answered within a step. A model on one clock loses one of them: with the whole model in the loop at a latency of 6 steps, 74 % of shoves are survived, and with the fast loop alone 0 % of moved vases. Two clocks keep 100 % of the shoves and 66 % of the vases, on 2 copies of the reasoner running at once instead of 6. The Bench's reasoner is a stand-in that reads the scene without error, so these numbers measure timing alone.
Linear position
Forced by: Learning from a score alone works, slowly, on a small task, and works in an afternoon of robot time only when it starts from a policy that is already nearly right and learns a small correction. Every source so far feeds that starting point: demonstrations, other bodies, video, force, simulation. It is one model that must read language and images, feel force, emit chunks, and still answer within one control step. How is such a model built, and what does it cost to run it at the speed of the robot?
New idea: a robot model runs on two clocks: a slow reasoner that looks every K steps and hands over a chunk of H targets L steps later (its latency), and a fast loop that feels the arm at every step; the chunk must last at least L + K steps, H ≥ L + K. What each clock handles is set by who can see the signal and how long it can wait.
Forces next: A large model that reads and reasons is too slow to steer an arm directly, so it runs slowly and hands a plan to a small fast network that runs at the control rate, and the split decides what each part must be trained on. Training the parts in order, on data of very different size and quality, has its own failure: fine-tuning on a few dozen tasks raises every number we track while erasing what the language model knew. In what order, on what mixture, should the pieces be trained, and how do we notice the loss?
The plan
Six moves. (1) Price the control step: what a stale reading does to a loop, and how long real models take. (2) Price reading: a ruler of operations and bytes. (3) Run the whole model on one clock three ways and watch each fail. (4) Split it, and derive the rule the chunk obeys. (5) Say which signals belong to which clock, and check real systems against the rule. (6) Read off what the split asks of training.

1 · The price of a control step

A control step is 0.05 s. When something slow sits in the loop, lesson 6's delay appears: the command computed now comes from a reading d steps old, the error obeys et+1 = et − k·et−d (k is the share of the error one step removes), and the loop is stable only for k < 2 sin(π / (4d + 2)). The ceiling is 2.00 at d = 0, 1.00 at 1, 0.62 at 2, 0.45 at 3, 0.35 at 4, 0.24 at 6 and 0.15 at 10. The fast loop of this lesson takes k = 0.5 (a 7 cm shove is 1.75 cm two steps later), which is under the ceiling at d = 2 and over it from d = 3, a reading 150 ms old. Latency is paid twice: in the reaction, and in the stiffness the loop may have.

How old are the readings of real models? Published decision times, with lesson 1's planner, in steps of 50 ms:

systemwhat runstime per decisionsteps
RT-2 (Brohan et al., 2023)55B at 1 to 3 Hz on cloud TPUs; 5B at about 5 Hz333 to 1000 ms; 200 ms6.7 to 20; 4
OpenVLA, 7B (Kim et al., 2024)about 6 Hz on one RTX 4090167 ms3.3
π0, 3.3B (Black et al., 2024)on board, three cameras, a chunk of 5073 ms1.5
Gemini Robotics (2025)cloud backbone, local action decoderabout 250 ms5
the planner of lesson 1's exercise200 sequences × 25 steps × 0.1 ms500 ms10

None is under one step and most are several. "Answer within one control step" is a demand the models that read are not built to meet.

2 · The price of reading

A model with P parameters that reads N tokens (image patches and words) pushes each token through every weight, a multiply and an add: 2P operations per token, 2PN in all, which a device doing F operations a second takes 2PN/F to do. A model that writes its action as tokens, as RT-2 and OpenVLA do with 8 and 7 per action, must also read all of its weights again for each of the n tokens it writes, 2P bytes at two bytes a weight; at a batch of one that is almost no arithmetic per byte, so the time is set by the memory bandwidth B:

t = 2PN / F + n · 2P / B

Take a ruler of round numbers, not a product: F = 100 TFLOP/s, B = 1 TB/s, N = 250 tokens (three camera frames of 64 tokens each, the per-frame count of GR00T N1 and SmolVLA, and an instruction) and n = 8 tokens written. π0 and SmolVLA write no action tokens; for them the last columns are what their backbones would cost if they did.

parametersread, 2PN/Fper token written, 2P/Bread + 8 tokenssteps of 50 ms
0.45 B (SmolVLA's size)2.25 ms0.90 ms9.5 ms0.19
3 B (π0's backbone)15 ms6.0 ms63 ms1.3
7 B (OpenVLA's size)35 ms14 ms147 ms2.9

The ruler is round numbers, not a calibration, and it runs under the printed times: 147 ms where OpenVLA's 6 Hz is 167, and 6 ms for π0's ten flow steps through its 0.3 B action expert, the small network beside the backbone that writes the chunk (0.6 GB read per step, 0.6 ms at 1 TB/s), where Black et al. print 27. Turned around, it gives the largest model that fits a clock of τ seconds, P ≤ τ / (2N/F + 2n/B): 2.4 B at 50 ms (20 Hz), 0.95 B at 20 ms (50 Hz), 0.24 B at 5 ms (200 Hz) and 0.05 B at 1 ms. Writing a whole chunk in ten flow steps, each over all 50 action tokens at once, removes most of the second term: Driess et al. (2025) put an autoregressive VLA at 1.3 Hz and π0 at 10 Hz.

Does reading need the billions? RT-2 trained a 5B model on its robot data from scratch and it scored 9 % on unseen objects, backgrounds and environments; started from a web-pretrained model and fine-tuned on the same robot data it scored 42 %, and the 55B model 52 % (Brohan et al., 2023). The generalisation came with the web pretraining and grew with the size, and the size sets the clock.

3 · One model on one clock, three ways to fail

Put the question on the Bench: the five-post course with lesson 1's gust. The policy is lesson 6's, a chunk of absolute joint targets along the path, tracked from the joints it reads at every step with k = 0.5. A run ends at the mat, at a post, or at step 235 (a timeout).

Two things go wrong, each when the cup is 8 to 16 cm (0.5 to 1.1 s) before the middle post. A shove moves the cup 7 cm toward the post, as a bump would: the path passes 10 cm from the post's centre, so unless something pulls the cup back it passes 3 cm from it, and cup and post touch when their centres are closer than 5 cm, which leaves 5 cm of clearance. A moved vase moves the post 7 cm into the path, and the expert's path bends away from its new place. The joints feel the shove at the next step; only a model that looks at the scene sees the vase.

The reasoner on the Bench is a stand-in, the expert's own path planner, so that what is measured is when a chunk arrives and not how it was learned. Each cell is 300 runs under the same luck, good to within 5.7 points at 95 %, at a latency of L = 6 steps (300 ms, between Gemini Robotics and RT-2 55B):

the modela shovea vase movesnothing happenscopies of the reasoner
one model in the loop, naive: it acts on its stale reading0 %0 %0 %6
one model in the loop, best case: it adds up its own commands74 %75 %100 %6
fast loop alone, on the path it was trained on100 %0 %100 %0
two clocks (look every 4 steps, chunk of 12)100 %66 %100 %2

The model in the loop, naively. It acts on a reading L steps old as if it were current: lesson 6's loop with d = 6, whose ceiling is 0.24, at a stiffness of 0.5. It rings, and it fails even when nothing happens: 100 % at L = 2, 0 % at every latency from 3 to 8 steps with every run ending at a post, and at most 24 % beyond, where it flails. The model in the loop, at its best. Let it add up the commands it has issued since its reading and it is stable at every latency. It stays blind to what it did not cause: it learns of a shove L steps after the joints feel it, L × 0.05 s × 15 cm/s = 4.5 cm of travel against 5 cm of clearance, and it survives 74 % of the shoves. It also decides at every step and each decision takes L steps, so L decisions are always in flight: L copies of the model run at once. The fast loop alone. A network small enough for its clock (§2) feels every shove. It does not read the scene, so a vase that has moved is news it never gets: 0 %.

Each failure has a cause. Rows one and two act on a reading that is old; row three has a loop that cannot read. The cure for the first is a loop whose own readings are never old, and for the second something that reads. One network cannot be both, and two can.

Road not taken · make the one model fast enough
Quantise it, skip layers, distil it: no hand-off, no staleness, one network to train. It moves the ceiling and not the structure. Time is proportional to P, so halving the latency halves the parameters, and a faster device lifts every ceiling of §2 by the same factor until a signal that needs a kilohertz takes the model out of the loop again (1 ms fits 0.05 B parameters). Nor is a faster model automatically a better loop: OpenVLA's int8 model ran at 1.2 Hz on an A5000, against a controller that expected 5 Hz, and lost 13 points of success (71.3 to 58.1 %); when the controller waited for the model, all three precisions scored between 68.8 and 74.4 % (Kim et al., 2024). The loss was latency, not precision.

4 · Two clocks, and the rule the chunk obeys

Give what feels the arm its clock, and what reads its own. Real systems call the parts a backbone and an action expert; here they are the reasoner and the fast loop. The reasoner looks at the scene every K steps; what it saw becomes a chunk of H targets along the path of the course as it saw it, and the chunk lands L steps after the look. The fast loop runs at every step on fresh joints: it takes the next target of the newest chunk to have landed, commands k times the distance to it, and holds when no chunk is usable. Its own reading is never old, so §1's ceiling is not in play: the reasoner's latency costs the plan its age, not the loop its stability.

A chunk made at step s is usable from s + L, its first L targets having fallen due while it was made, to s + H − 1, and the next chunk lands at s + K + L. The steps from s + H to s + K + L − 1 have no usable chunk, so there is no gap exactly when

H ≥ L + K

in steps: the chunk must last at least the reasoner's latency plus its interval. Otherwise L + K − H steps of every K are spent holding, the arm moves for a share (H − L)/K of the time, and a run that needs 145.5 steps at full speed needs 145.5 · K/(H − L).

Two more consequences. A change in the world is first seen at the next look, 0 to K − 1 steps away and equally likely, and acted on L steps after that: the arm acts on a picture L + (K − 1)/2 steps old on average, and travels blind meanwhile. And the reasoner is busy L steps of every K, so it needs ⌈L/K⌉ copies running, against L for a model that decides at every step.

The widget

Look slowly, feel quickly
Top: 12 of the 300 runs on the table from above, the cup starting at the left and aiming at the mat on the right (cyan reached it, red touched a post), the vase before (dashed) and after, the band where the event fires. Middle: the two clocks over run 1; amber the reasoner reading, light teal a chunk that can still be used, dark teal the chunk in force, red a step with none. Bottom: success against latency L for two clocks (this K and H), for one model in the loop and for the fast loop alone; the red band is where H < L + K, the amber dot is the setting, the ticks are latencies from §1. K and H act on the two-clocks design only. The first control is L. The reasoner is the expert's own path planner, a stand-in that reads the scene without error: the widget measures timing, not understanding.
reaches the mat
—
95 % interval
—
touches a post
—
times out
—
event acted on after
—
travelled blind meanwhile
—
steps held per run
—
steps to the mat
—
chunk margin H − (L + K)
—
copies of the reasoner
—
reasoner passes per second
—
Show the core JS
CK.ceiling = function (d) { return 2 * Math.sin(Math.PI / (4 * d + 2)); };
CK.readTime = function (P, N, n, F, B) { F = F || CK.F; B = B || CK.B; return 2 * P * N / F + n * 2 * P / B; };
...
var s = t - L, seen = (wMoved && s >= fired) ? wMoved : w0;
qs = hist[Math.max(0, s)].slice();
if (!o.naive) for (j = Math.max(0, s); j < t; j++) { qs[0] += dt * uapp[j][0]; qs[1] += dt * uapp[j][1]; }
tgt = CK.target(seen, ++c, vs);
u = [(k / dt) * (tgt[0] - qs[0]), (k / dt) * (tgt[1] - qs[1])];
...
while (nextObs <= t) { plans.push({ s: nextObs, land: nextObs + L, seen: (wMoved && nextObs >= fired) ? wMoved : w0, moved: !!(wMoved && nextObs >= fired) }); nextObs += K; }
active = null; for (j = plans.length - 1; j >= 0; j--) if (plans[j].land <= t) { active = plans[j]; break; }
if (!active || t - active.s >= H) { held++; if (act) act.push(-1); tgt = last; }
else { if (act) act.push(active.s); tgt = last = CK.target(active.seen, ++c, vs); }
u = [(k / dt) * (tgt[0] - q[0]), (k / dt) * (tgt[1] - q[1])];
...
CK.gap = function (L, K, H) { return Math.max(0, L + K - H); };
CK.moving = function (L, K, H) { return Math.min(1, Math.max(0, (H - L) / K)); };
CK.accel = function (L, K) { return Math.ceil(L / K); };
CK.meanAge = function (L, K) { return L + (K - 1) / 2; };

What to try. At the defaults (a vase moves, two clocks, L = 4, K = 4, H = 12, 15 cm/s) 84.0 % of the 300 runs reach the mat (interval 79.4 to 87.7 %); the event is acted on 5.4 steps (269 ms) after it happens, 4.0 cm of blind travel against 5 cm of clearance, and the chunk margin is +4. Move L first, to 6 and to 8: 66.3 % and 44.7 % of the vases are survived, and the margin falls to +2 and 0. Back at 4, switch the event to a shove: 100 %, the fast loop feels it at once. Switch the design to one model in the loop: shoves fall to 92.0 % with 4 copies running, and vases do better, 94.7 %, because there is no K to wait for. The fast loop alone: shoves 100 %, vases 0 %. The naive loop, with nothing happening: 100 % at L = 2 and 0 % at L = 3, every run ending at a post. Back to two clocks and the vase, set L = 10 (500 ms, the planner of lesson 1) and H = 18: 100 % of the shoves and 25.3 % of the vases survive, against 44.7 % and 38.7 % for the model in the loop. Back at L = 4, vases and H = 16, slide K to 1 and to 8: 97.0 % with 4 copies running, 84.0 % at 4, 63.3 % at 8 with one copy and 2.5 passes a second; every look skipped is a step more of blind travel. Set the event to nothing, K = 8, and shorten H: at 12 there is no gap and a run takes 145.5 steps; at 10 the gap is 2 steps of every 8, 48.5 steps are held per run and 193.6 steps are needed (the rule says 194); at 8 every run times out, 0 % reach the mat. Last, with the vase and the defaults, raise the arm speed: 96.7 % at 10 cm/s, 63.7 % at 20 and 22.0 % at 30.

5 · What each clock is for

The deadline of a signal is a distance: how far the cup can travel on a stale picture before the event is lost. With the model in the loop, where the age is exactly L, success on the vase falls to 90 % at an age of 7.4 steps at 10 cm/s, 4.5 at 15 and 3.3 at 20: 3.7, 3.4 and 3.3 cm of travel, two thirds to three quarters of the 5 cm clearance each time. A faster arm or a tighter clearance shrinks the budget in steps in proportion; at 1 m/s a reasoner a quarter of a second late has been blind for 25 cm. Two clocks spend the same budget on L + (K − 1)/2 and not on L: at K = 4 the 90 % point is at L = 3.3, a mean age of 4.8 steps, so the curve moves left by about a step.

That sorts the signals. One whose deadline is shorter than L cannot go through the reasoner, whatever K is; one longer than L + K can. What the fast loop can do alone it must: the shove, and contact. What only the reasoner understands rides the chunk: language, which object, which vase, which subtask. Real systems split the same way, and where the numbers are printed the rule can be checked:

systemreasoner and action partwhat the numbers say
π0 (Black et al., 2024)one 3.3B pass makes a chunk of 50, run open loop, 16 actions at 20 Hz or 25 at 50 Hz before the next look73 ms is 1.5 steps at 20 Hz: 17.5 needed of 50, 32.5 to spare; at 50 Hz 3.65 + 25 = 28.65, 21.35 to spare. Only 32 % or 50 % of a chunk is run.
GR00T N1 (NVIDIA, 2025)a 1.34B model reads at 10 Hz; a flow-matching network samples a chunk of 16 in 63.9 ms15.6 chunks and 250 actions a second, against the 20 a second of its own teleoperated demonstrations
Helix (Figure, vendor-reported, 2025)a 7B model at 7 to 9 Hz, an 80M network at 200 Hz22 to 29 fast steps per slow update
Gemini Robotics (2025)cloud backbone under 160 ms, local decoder, about 250 ms end to end, 50 Hz effectivea chunk of at least 12.5 actions; its length is not printed
SmolVLA (Shukor et al., 2025)0.45B, a chunk of 50 at 30 Hz (1.67 s)asynchronous inference, the next chunk requested before the last runs out, scored 73.3 % on its three real tasks against 78.3 % synchronously and finished a pick-and-place 30 % faster. In simulation, running all 50 of 50 actions before looking again scored 51.8 % on LIBERO, against 80.3 % for looking after every action.

It is one bargain: the more often the reasoner looks, the fresher the picture and the more copies it takes; the less often, the longer the chunk it must leave behind.

6 · What the split asks of training

The split fixes what each part must be trained on, and the two are nothing alike. The reasoner learns to read, from data that mostly come from somewhere else: RT-2 (Brohan et al., 2023) filtered about 10B image-text pairs to a billion training examples, and π0.5 (Physical Intelligence, 2025) reports that 97.6 % of the examples of its first phase do not come from the mobile robots doing household tasks. The action network learns to act, from robot data that costs robot time (over 10,000 hours for π0, Black et al., 2024), and it is small: 9.1 % of π0's parameters, 1.1 % of Helix's, on its maker's figures. The hand-off decides how far they can be trained apart. π0.5 first writes the subtask as text ("pick up the plate"), and its mix has data that teaches only that. π0 hands over features and the action loss reaches back into the backbone; Helix, by its maker's account, does the same; GR00T N1 keeps its language model frozen and trains the vision encoder and the action network (NVIDIA, 2025). Driess et al. (2025) find that adding a randomly initialised action expert naively harms both training speed and the transfer of knowledge.

And the numbers we track will not see it. The shove column of §3 reads 100 % with a working reasoner and 100 % with none; it takes a probe that needs the reasoner, here the moved vase, 66 % against 0 %, to tell them apart. The real cases look alike. A pre-trained GR00T N1 checkpoint scored 76.6 % on a left-to-right hand-over; after post-training on right-hand demonstrations the policy had lost it (NVIDIA, 2025). Removing π0.5's web data left its full-task benchmark within noise and hurt following language about new objects (Physical Intelligence, 2025). A model trained in parts can lose what one part knew while every number we watch holds.

What this lesson did not do
The reasoner on the Bench is the expert's own path planner: it reads no picture and no word and learns nothing, so what was measured is timing, not understanding (lessons 1 to 12 are how a policy learns). The fast loop is a plain tracker; a real action network makes its chunk by iterating denoising steps, whose own latency enters the same rule one level down (π0 uses ten, GR00T N1 four), and that is arithmetic here, not simulation. The ruler is round numbers for one device and a fixed latency, where a cloud link varies, and the Bench has no signal faster than its 20 Hz step. The Bench's deadline comes from its slow arm and 5 cm clearance. Which data trains which part, in what order and on what mixture, and how a loss like that is noticed, is the next lesson; what a verdict on any of it costs is lesson 15.

Common mistakes / failure modes

"a bigger model is a better controller"
With the whole model in the loop at L = 6 steps, 74 % of shoves are survived; acting on its stale reading it scores 0 % even when nothing happens (§3).
"the chunk length is the planning horizon"
It only has to last L + K. π0 predicts 50 steps and runs 16 or 25 before looking again (§4, §5).
"asynchronous inference is free"
The picture is on average L + (K − 1)/2 steps old: 5.5 at L = 4, K = 4 (the runs measure 5.4), which is 4.0 cm of blind travel (§4).
"the fast network corrects a stale plan"
It tracks the chunk it is given: with no reasoner 0 % of moved vases are survived (§3).
"a faster GPU removes the problem"
Every ceiling rises by the same factor, and a faster signal needs a smaller model: 0.24 B at 5 ms (§2).
"latency only delays the reaction"
In a loop it caps the stiffness: k must stay below 0.45 at 3 steps and 0.15 at 10 (§1).
"if success is high the reasoner is fine"
Shoves score 100 % with no reasoner at all (§6).

Checkpoint exercise

Try it
A reasoner takes 200 ms to answer and looks every 300 ms, on a 20 Hz robot that carries at 0.2 m/s. (a) What are L and K in steps, and the shortest chunk that leaves no gap? (b) How old is the picture on average when a change is acted on, and how far does the arm travel blind meanwhile? (c) How many copies of the reasoner does it keep busy, against a model that decides at every step? Answer: (a) 200 ms is 4 steps and 300 ms is 6, so H ≥ L + K = 10 steps, half a second. (b) L + (K − 1)/2 = 6.5 steps, 325 ms, and 0.325 s × 0.2 m/s = 6.5 cm. On the Bench's 5 cm clearance that is too much: it loses one event in ten at 3.5 cm (§5). (c) ⌈L/K⌉ = 1 copy, against 4 for a decision at every step.

Where this points next

A model that reads is too slow to steer an arm: at a latency of 6 steps the best whole-model loop keeps 74 % of the shoves, and the fast loop alone feels everything and sees nothing. Two clocks give each what it can do, a chunk that lasts L + K and a loop at every step, for 2 copies of the reasoner instead of 6. But the parts differ in size (a 0.3B expert beside a 3B backbone in π0) and in what an example costs, the action loss reaches the reasoner through the hand-off, and a number like 100 % under shoves cannot tell a reasoner that works from one that has been written over; GR00T N1 lost a skill its pre-trained checkpoint had. In what order, on what mixture, should the pieces be trained, and how do we notice the loss?

Takeaway
A model that reads costs 2 · parameters · tokens operations plus a pass over its weights per token written: at 7B, 147 ms on a round-number ruler, 2.9 control steps, and a loop that acts on a reading d steps old must keep its stiffness below 2 sin(π / (4d + 2)). So the model that reads cannot be the loop that feels: with it in the loop at a latency of 6 steps 74 % of shoves are survived, and with the fast loop alone 0 % of moved vases. Two clocks fix both: a reasoner that looks every K steps and hands over a chunk L steps later, and a fast loop on fresh joints at every step, provided the chunk lasts at least L + K, or the arm holds for the difference. The arm then acts on a picture L + (K − 1)/2 steps old, which costs distance (about 3.5 cm of stale travel before one event in ten is lost), and the reasoner needs ⌈L/K⌉ copies. What each part is trained on follows from the split, and so does its first failure: tasks that exercise only the fast loop can raise the numbers we watch and say nothing about the reasoner.

Interview prompts

Companion reads: System ML · 15 Disaggregated prefill / decode (reading is compute-bound, writing is memory-bound), World Models · 29 Distilling into a real-time loop (latency as part of the specification) and World Models · 09 Plan with it (the slow planner of the world-model track).