The recipe: stages, mixtures and forgetting
Lesson 13 split the robot model into a slow reader and a fast actor and left open how to train them, because the parts often share weights and the data that train them differ in size, kind and price. This lesson trains a small network on a knowledge task and then fine-tunes it on 36 layouts of the Bench's course. Success rises to 72.9 % and the copy error falls, while a probe of the knowledge task falls from 98.5 % to 38.8 %, starting before success moves. Five cheaper repairs are computed and fail. What works is to keep the old data in every batch of every later stage, to order the stages by trust, and to log a probe the mixture cannot see. It costs steps, and the price is smaller than the error bar of an evaluation a robot can afford.
New idea: keep the old data in every batch of every later stage, in a share ρ, and order the stages by trust. The old task's own loss then acts as a spring that is stiff where the old task is sensitive and slack elsewhere, and its price on the Bench is about 1/(1 − ρ) in steps.
Forces next: Staging the training and mixing the old data into the new keeps the language the model started with, at a measurable price in speed and in the tasks it learns. We now have a recipe and a checkpoint, and need to say whether the checkpoint is better than the last one. An evaluation is a count of trials, and each trial costs a robot and a person for minutes. How many trials does it take to tell one policy from another, and what does that do to the cost of the whole programme?
1 · One set of weights, two kinds of data
Several of the systems lesson 13 describes do not train their two parts apart. RT-2 (Brohan et al., 2023) and OpenVLA (Kim et al., 2024) fine-tune the vision-language model itself to write actions. In π0 (Black et al., 2024) the action expert attends to the backbone's keys and values, so unless the gradient is stopped its loss reaches the backbone's weights, and Knowledge Insulating VLAs (Driess et al., 2025) report that attaching a new expert naively harms both training speed and knowledge transfer. One set of weights has to serve what the model learned before it saw a robot and what the robot now teaches it.
The Bench cannot hold a billion weights. It can hold the mechanism, as an analogue of that harm: one small network, two kinds of data, and a measurement of what each costs the other. K stands for the web-scale data a robot model starts from, B for the robot tasks it is then taught.
| K, the knowledge task | B, the robot tasks | |
|---|---|---|
| what it is | 120 words, each a code of 8 numbers, in one of 10 categories by an arbitrary rule | 36 layouts of a two-post course: 12 shifts of the row, ±10 cm, times 3 tilts |
| reads, writes | a rendition of a word (its code plus noise of 0.15 per number); 10 category scores | the cup's position and the layout; the velocity it wants for the cup (the Bench's solver turns it into joint velocities) |
| data | an endless stream, every example new | 3 calm demonstrations per layout: 8884 frames |
| score | the probe: 600 fresh renditions no step trained on; the share right (chance 10 %) | success: 4 rollouts per layout under gust 0.08 rad/s (144); copy error on 2964 held-out expert frames |
The network reads 12 numbers and writes 12, through two hidden layers of 24 tanh units: 1212 weights. A K example fills the first 8 inputs and trains the first 10 outputs, a B example fills the last 4 inputs and trains the last 2, and the hidden layers belong to both. The head that writes B's two numbers starts at zero, so the pretrained model does nothing with an arm. K is a table of arbitrary facts: nothing in a word predicts its category, so the rule can only be stored, and it is stored in the weights. The symbols: ρ is the share of every batch of 40 examples drawn from K; θ is all the weights and θK their values when stage 1 ends; LK and LB are the mean squared errors on K and on B examples; d is the Euclidean distance the hidden-layer weights have moved from θK.
Stage 1 trains on K alone for 4000 steps with the Adam optimiser. The probe reads 98.5 % and success on B is 0 %. This checkpoint is what the language model started with.
2 · Fine-tune on the tasks, and every number we track improves
Stage 2, post-training, runs 600 steps of 40 B examples with a fresh optimiser, the learning rate decaying from 0.003 to 5 % of that inside the stage. Every 50 steps we log what a training dashboard shows, the copy error and the success, and the probe. Seed 1, no K in the batches:
| step of post-training | copy error, held-out frames | success on the layouts | probe of K |
|---|---|---|---|
| 0 | 100.0 % | 0.0 % | 98.5 % |
| 50 | 57.4 % | 0.0 % | 80.8 % |
| 200 | 34.1 % | 2.1 % | 53.0 % |
| 300 | 16.7 % | 68.1 % | 39.0 % |
| 600 | 14.1 % | 72.9 % | 38.8 % |
The copy error falls from 100 %, a model that does nothing, to 14.1 %. Success crosses 50 % at step 300 and ends at 72.9 %; one round of corrections (stage 3, §5) lifts it to 92.4 %. Every number we track says the fine-tuning worked. The probe falls from 98.5 % to 38.8 %, and it falls early: at step 50 it has lost 17.7 points while success is still 0 %, and when success crosses 50 % it has lost 59.5. Five seeds, each with its own vocabulary, initial weights and batches, agree: after post-training the probe is 42 % (39 to 48) and success 69 % (57 to 76).
Why? Stage 2 descends LB alone, and nothing in LB mentions K. Stage 1 left LK at a minimum, so its slope there is zero and, for small moves,
LK(θ) ≈ LK(θK) + ½ (θ − θK)T HK (θ − θK)
where HK is the matrix of second derivatives of LK. The damage is the distance moved, weighed along the directions where HK is large. Fitting B needs distance: the new head starts at zero and the hidden layers must make features it can read. Adam also steps each weight by up to about the learning rate whatever the size of its gradient, so the hidden layers move at once: after 10 steps d = 0.60 and the probe has lost 3.3 points. After post-training d = 4.90 and the K loss has risen from 0.026 by 0.255. To test the expansion, walk in a straight line from θK to those weights and read LK on the way: it rises as the distance to a power between 1.96 and 2.14 over the first 30 % of the walk (five seeds) and between 1.82 and 1.95 over all of it, with r² above 0.99.
3 · Five repairs, computed
The expansion says what a repair must do: keep the weights from moving along K's sensitive directions while B still moves them. Five cheaper ideas come to mind. Each runs on five seeds for 600 steps (the learning-rate rows scale the steps so that rate × steps stays within 2 %), with success counted on 360 rollouts, because a probe that survives is worth little if success does not.
| repair | distance d | probe of K | success | what goes wrong |
|---|---|---|---|---|
| fine-tune everything on B (the failure) | 5.03 | 42 % | 69 % | K is lost |
| train only the new head | 0.00 | 97 % | 0 % | nothing moves, nothing is learned |
| train layer 1 and the head | 2.80 | 88 % | 4 % | learns later, still loses K |
| learning rate × 0.3, 2004 steps | 5.50 | 39 % | 74 % | the same place, reached later |
| learning rate × 3, 204 steps | 4.29 | 50 % | 41 % | the same road, stopped short |
| pull towards θK: add λ/2 ‖θ − θK‖² to LB, λ = 0.03 | 1.12 | 93 % | 0 % | a round spring holds every direction, B's too |
| the same, λ = 0.01 | 2.20 | 78 % | 21 % | K and B are lost together |
| K alone for 300 steps after the failure | — | 95 % | 6 % | the last stage wins |
Freezing keeps K exactly, because nothing it uses is trained, and nothing frozen is about the arm. Training layer 1 and the head learns, but not within the budget, and the probe still falls. The learning rate sets the speed along the road, not the road: × 0.3, × 1 and × 3 end at distance 5.50, 5.03 and 4.29, with the probe at 39, 42 and 50 %, in the order of the distance. The damage is the price of the distance, and B needs the distance. K alone at the end brings K back and takes B away: whatever is trained last decides the end state, and §5 uses that.
The pull towards θK is the instructive failure, because it is the first repair aimed at the right quantity. Its spring has the same stiffness λ in every direction, but K is sensitive along a few directions and indifferent along the rest, and B needs some of the indifferent ones: a round spring stiff enough to protect K (λ = 0.03) stops B, and a slack one loses both. What is wanted is a spring that is stiff where HK is large and slack where it is zero. The K loss is exactly that spring, its own expansion above, and its gradient can be had whenever we like, because K's data are a stream.
4 · The move: a share ρ of K in every batch
A batch of 40 examples holds round(40ρ) from K and the rest from B, and its expected gradient is
(1 − ρ) ∇LB + ρ ∇LK, with ∇LK ≈ HK (θ − θK) near θK
which is the pull of B plus a spring of stiffness ρHK back to θK. The spring is stiff along the directions K is sensitive to and slack along the others, so B can move the weights wherever K does not mind. It steers the weights and does not stop them. Measured at ρ = 0.3, five seeds: the hidden layers end at distance 3.82, 0.76 times the 5.03 they travel without the mixture, and the K loss rises by 0.0047 against 0.220, 47 times less.
The price is steps. Each batch carries only a share 1 − ρ of B, so if progress depended on nothing but the number of B examples seen, B would need 1/(1 − ρ) times the steps, 1.43 at ρ = 0.3. Measured as the step at which success first reaches 50 %, over the four seeds that get there at both shares (the fifth does not at ρ = 0.3 within 600 steps), it grows from 300 to 438, a factor 1.46. At a fixed budget the same price reads as tasks not yet learned: after the corrections of §5 success is 89 % without the mixture and 82 % with it, 7.3 points lower (4.4 to 10.6, positive in all five seeds). With 2004 steps instead of 600 the gap closes, 88.6 % against 90.6 % when post-training ends, and the probe stands at 35.9 % against 97.5 %: on the Bench the price is time. The table gives the five-seed means for each share.
| share ρ | probe after post-training | largest fall of the probe | success after post-training | success after corrections |
|---|---|---|---|---|
| 0 | 42 % | 54.8 points | 69 % | 89 % |
| 0.1 | 91 % | 8.4 | 65 % | 85 % |
| 0.2 | 93 % | 5.0 | 57 % | 84 % |
| 0.3 | 95 % | 3.4 | 54 % | 82 % |
| 0.5 | 96 % | 1.7 | 25 % | 66 % |
Which share? The alarm of §6 is a fall of 5 points in the probe. Of the shares run (the table's, and 0.05, 0.4 and 0.7) the smallest for which no seed trips it is 30 %: the largest falls are 2.0 to 4.7 points, against up to 7.2 at 20 %. The grid, the budget and the 5-point alarm are choices, and the table is the evidence, not the 0.3.
The widget
What to try. Leave ρ = 30 % and the recipe: the probe reads 98.5 % after stage 1 and 95.7 % at the end, success is 86.1 % after the corrections, and the alarm never trips. Success shows one early blip, 23.6 % at step 50 and 0.0 % at step 100: a half-trained policy that clears some layouts and loses them. Slide ρ to 0: the probe falls to 38.8 % after post-training and 34.5 % at the end, success reaches 92.4 %, the alarm trips at step 50 and the weights move 4.90. Now keep ρ = 0 and open the menu, which reruns the repairs of §3 on seed 1. Only the new head trains: the probe stays at 98.5 % and success at 0.0 %. K alone at the end takes the probe from 38.8 % to 96.3 % and success from 72.9 % to 0.0 %. Back at ρ = 30 %: the mixture that holds 60 of 120 words reads 99.7 % on those and 60.7 % on the others, and the alarm trips at step 50; no K in the corrections takes the probe from 95.7 % to 86.8 %.
5 · Stages: the order, and what each one carries
Three kinds of data differ in size and in what an example costs, and the order of the stages follows from that and from the result of §3 that the last stage wins.
| stage | data | size | what an example costs |
|---|---|---|---|
| 1 pretrain | K: broad, noisy renditions | an endless stream | almost nothing: scraped |
| 2 post-train | B: calm, curated demonstrations | 8884 frames | a person and a robot |
| 3 correct | states the model itself reaches, labelled by the expert (lesson 3) | 2740 frames, 31 % of the demonstrations | the most: a labeller for every state |
Whatever is trained last decides the end state, so the data you trust most and pay most for go last, and the cheap abundant data go first, where they build what the later data will use. The same result says what each later stage must carry: one that does not carry the earlier data erases them. The widget's stage 3 keeps drawing K at the share ρ and adds the corrections to the demonstrations, as lesson 3 aggregated; dropping the mixture for those 300 short steps takes the probe from 95 % to 88 % (7.1 points over five seeds) and buys success only 3 points. Every stage is another chance to forget.
Stage 3 is lesson 3's recipe with one round: one rollout of the current model on each layout under the same gust, every visited state labelled by the expert, then 300 more steps at a learning rate of 0.002. It lifts success from 72.9 % to 92.4 % without the mixture and from 66.0 % to 86.1 % with ρ = 0.3 (seed 1; over five seeds by 20 and 28 points), and the probe stays at 95.7 %. Published recipes have the same shape. π0 (Black et al., 2024) pretrains on over 10,000 hours of robot data, diverse and of lower quality, which lets the model recover from mistakes, then post-trains on 5 hours for the simplest tasks and 100 or more for the hardest, data that show a consistent and fluent strategy. π0.5 (Physical Intelligence, 2025) pretrains for 280,000 steps and post-trains for 80,000 more. RT-2 (Brohan et al., 2023) co-fine-tunes with web data and, in its ablation, reaches 63 % on unseen objects, backgrounds and environments for its 55B model against 52 % for plain fine-tuning.
6 · Noticing the loss, and what the verdict costs
A probe is a set of questions about what the model used to know, asked at every checkpoint. Three rules follow from the measurements. Log it often. Without the mixture the alarm, a probe 5 points below its pretrained value, trips at step 50, 250 steps before success crosses 50 %; success is still 0 % then, and the copy error, at 57.4 %, is falling as it should: nothing tracked warns. Make it fresh and broad. The Bench's probe is 600 renditions no step trained on, from all 120 words. Keep it out of the mixture. In the widget's variant mixture holds 60 of 120 words, the mixture draws only words 0 to 59. The probe reads 99 % on those and 62 % on words 60 to 119 (five seeds): a probe made of the mixture's own words would have called this model undamaged. The mixture protects what it contains and, for a table of arbitrary facts, little else.
The probe is cheap: 600 forward passes, about 0.6 million multiply-adds. What the recipe costs is success, which is not. The price of ρ = 0.3 is 7.3 points after the corrections (89 % against 82 %, five seeds of 360 rollouts each). A laboratory that runs one trial per layout on each checkpoint, 36 trials, measures the default model's 86 % as 31 of 36, a 95 % Wilson interval of 71.3 to 93.9 % with a half-width of 11.3 points: the price sits inside the error bar. Even the widget's 144 rollouts (79.5 to 90.8 %) have a half-width of 5.7 points, most of the price. The probe separates 96 % from 41 % in milliseconds; the success that justifies what the recipe costs is invisible at the number of trials a robot can afford.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
Staging the training and keeping the old data in every batch preserves what the model started with: at ρ = 0.3 the probe ends at 96 % where plain fine-tuning leaves 41 % (five-seed means). The price is steps, a factor of about 1.5, and 7.3 points of success after the corrections at a fixed budget. The probe that finds the damage costs 0.6 million multiply-adds, but the evaluation that would show the price cannot resolve it at the trials a robot can afford: 36 trials give 71.3 to 93.9 %. We now have a recipe and a checkpoint, and need to say whether the checkpoint is better than the last one. An evaluation is a count of trials, and each trial costs a robot and a person for minutes. How many trials does it take to tell one policy from another, and what does that do to the cost of the whole programme?
Interview prompts
- Why does fine-tuning a pretrained model on a few tasks erase what it knew, even though the new data never contradict it? (§2 — the new loss does not mention the old task, the old loss is at a minimum so it rises roughly as the square of the distance moved, and the new task needs distance.)
- A colleague proposes a learning rate three times smaller to avoid forgetting. What do you expect? (§3 — with three times the steps, the same road: about the same distance and the same probe; the rate sets the speed along it.)
- Why is a pull towards the pretrained weights worse than a share of the old data? (§3, §4 — it is a round spring that also holds the directions the new task needs; the old data's loss is stiff only where the old task is sensitive.)
- What does a share ρ of old data cost, and how would you predict it? (§4 — on the Bench about 1/(1 − ρ) times the steps, since the new data are 1 − ρ of every batch; the budget decides what that costs in tasks.)
- Why train the broad, cheap data first and the curated, costly data last? (§5 — the last stage decides the end state.)
- Your forgetting probe reads 99 % while users report lost abilities. What do you check? (§6 — whether the probe is made of data the mixture replays: the words it never drew sat at 62 %.)
Companion reads: World Models · 26 The curriculum (when staging beats joint training), World Models · 27 The mixture is the model (how mixture weights are set) and World Models · 28 Post-training a world model (what post-training can and cannot add).