all_lessons/Robot Model Training/14 · Recipelesson 14 / 24

The recipe: stages, mixtures and forgetting

Lesson 13 split the robot model into a slow reader and a fast actor and left open how to train them, because the parts often share weights and the data that train them differ in size, kind and price. This lesson trains a small network on a knowledge task and then fine-tunes it on 36 layouts of the Bench's course. Success rises to 72.9 % and the copy error falls, while a probe of the knowledge task falls from 98.5 % to 38.8 %, starting before success moves. Five cheaper repairs are computed and fail. What works is to keep the old data in every batch of every later stage, to order the stages by trust, and to log a probe the mixture cannot see. It costs steps, and the price is smaller than the error bar of an evaluation a robot can afford.

The thesis, here
A model's weights serve everything it was ever trained on, and the loss of a new task does not mention the old ones. Fine-tuning therefore moves the weights as far as the new task needs and charges the old task for the distance. On the Bench, moving less does not help, because the new task needs the distance; moving where the old task does not look does, and the old task's own data, a share of every batch, are what point there.
Linear position
Forced by: A large model that reads and reasons is too slow to steer an arm directly, so it runs slowly and hands a plan to a small fast network that runs at the control rate, and the split decides what each part must be trained on. Training the parts in order, on data of very different size and quality, has its own failure: fine-tuning on a few dozen tasks raises every number we track while erasing what the language model knew. In what order, on what mixture, should the pieces be trained, and how do we notice the loss?
New idea: keep the old data in every batch of every later stage, in a share ρ, and order the stages by trust. The old task's own loss then acts as a spring that is stiff where the old task is sensitive and slack elsewhere, and its price on the Bench is about 1/(1 − ρ) in steps.
Forces next: Staging the training and mixing the old data into the new keeps the language the model started with, at a measurable price in speed and in the tasks it learns. We now have a recipe and a checkpoint, and need to say whether the checkpoint is better than the last one. An evaluation is a count of trials, and each trial costs a robot and a person for minutes. How many trials does it take to tell one policy from another, and what does that do to the cost of the whole programme?
The plan
Six moves. (1) Build a stand-in for shared weights and for the two kinds of data that train them. (2) Fine-tune it, and compare what a dashboard tracks with what the old task loses. (3) Say what the damage depends on and rule out five cheaper repairs by computing them. (4) Derive the mixture, measure its price, choose a share. (5) Order the stages. (6) Log a probe the mixture cannot see, and ask what judging the result costs.

1 · One set of weights, two kinds of data

Several of the systems lesson 13 describes do not train their two parts apart. RT-2 (Brohan et al., 2023) and OpenVLA (Kim et al., 2024) fine-tune the vision-language model itself to write actions. In π0 (Black et al., 2024) the action expert attends to the backbone's keys and values, so unless the gradient is stopped its loss reaches the backbone's weights, and Knowledge Insulating VLAs (Driess et al., 2025) report that attaching a new expert naively harms both training speed and knowledge transfer. One set of weights has to serve what the model learned before it saw a robot and what the robot now teaches it.

The Bench cannot hold a billion weights. It can hold the mechanism, as an analogue of that harm: one small network, two kinds of data, and a measurement of what each costs the other. K stands for the web-scale data a robot model starts from, B for the robot tasks it is then taught.

K, the knowledge taskB, the robot tasks
what it is120 words, each a code of 8 numbers, in one of 10 categories by an arbitrary rule36 layouts of a two-post course: 12 shifts of the row, ±10 cm, times 3 tilts
reads, writesa rendition of a word (its code plus noise of 0.15 per number); 10 category scoresthe cup's position and the layout; the velocity it wants for the cup (the Bench's solver turns it into joint velocities)
dataan endless stream, every example new3 calm demonstrations per layout: 8884 frames
scorethe probe: 600 fresh renditions no step trained on; the share right (chance 10 %)success: 4 rollouts per layout under gust 0.08 rad/s (144); copy error on 2964 held-out expert frames

The network reads 12 numbers and writes 12, through two hidden layers of 24 tanh units: 1212 weights. A K example fills the first 8 inputs and trains the first 10 outputs, a B example fills the last 4 inputs and trains the last 2, and the hidden layers belong to both. The head that writes B's two numbers starts at zero, so the pretrained model does nothing with an arm. K is a table of arbitrary facts: nothing in a word predicts its category, so the rule can only be stored, and it is stored in the weights. The symbols: ρ is the share of every batch of 40 examples drawn from K; θ is all the weights and θK their values when stage 1 ends; LK and LB are the mean squared errors on K and on B examples; d is the Euclidean distance the hidden-layer weights have moved from θK.

Stage 1 trains on K alone for 4000 steps with the Adam optimiser. The probe reads 98.5 % and success on B is 0 %. This checkpoint is what the language model started with.

2 · Fine-tune on the tasks, and every number we track improves

Stage 2, post-training, runs 600 steps of 40 B examples with a fresh optimiser, the learning rate decaying from 0.003 to 5 % of that inside the stage. Every 50 steps we log what a training dashboard shows, the copy error and the success, and the probe. Seed 1, no K in the batches:

step of post-trainingcopy error, held-out framessuccess on the layoutsprobe of K
0100.0 %0.0 %98.5 %
5057.4 %0.0 %80.8 %
20034.1 %2.1 %53.0 %
30016.7 %68.1 %39.0 %
60014.1 %72.9 %38.8 %

The copy error falls from 100 %, a model that does nothing, to 14.1 %. Success crosses 50 % at step 300 and ends at 72.9 %; one round of corrections (stage 3, §5) lifts it to 92.4 %. Every number we track says the fine-tuning worked. The probe falls from 98.5 % to 38.8 %, and it falls early: at step 50 it has lost 17.7 points while success is still 0 %, and when success crosses 50 % it has lost 59.5. Five seeds, each with its own vocabulary, initial weights and batches, agree: after post-training the probe is 42 % (39 to 48) and success 69 % (57 to 76).

Why? Stage 2 descends LB alone, and nothing in LB mentions K. Stage 1 left LK at a minimum, so its slope there is zero and, for small moves,

LK(θ) ≈ LK(θK) + ½ (θ − θK)T HK (θ − θK)

where HK is the matrix of second derivatives of LK. The damage is the distance moved, weighed along the directions where HK is large. Fitting B needs distance: the new head starts at zero and the hidden layers must make features it can read. Adam also steps each weight by up to about the learning rate whatever the size of its gradient, so the hidden layers move at once: after 10 steps d = 0.60 and the probe has lost 3.3 points. After post-training d = 4.90 and the K loss has risen from 0.026 by 0.255. To test the expansion, walk in a straight line from θK to those weights and read LK on the way: it rises as the distance to a power between 1.96 and 2.14 over the first 30 % of the walk (five seeds) and between 1.82 and 1.95 over all of it, with r² above 0.99.

3 · Five repairs, computed

The expansion says what a repair must do: keep the weights from moving along K's sensitive directions while B still moves them. Five cheaper ideas come to mind. Each runs on five seeds for 600 steps (the learning-rate rows scale the steps so that rate × steps stays within 2 %), with success counted on 360 rollouts, because a probe that survives is worth little if success does not.

repairdistance dprobe of Ksuccesswhat goes wrong
fine-tune everything on B (the failure)5.0342 %69 %K is lost
train only the new head0.0097 %0 %nothing moves, nothing is learned
train layer 1 and the head2.8088 %4 %learns later, still loses K
learning rate × 0.3, 2004 steps5.5039 %74 %the same place, reached later
learning rate × 3, 204 steps4.2950 %41 %the same road, stopped short
pull towards θK: add λ/2 ‖θ − θK‖² to LB, λ = 0.031.1293 %0 %a round spring holds every direction, B's too
the same, λ = 0.012.2078 %21 %K and B are lost together
K alone for 300 steps after the failure—95 %6 %the last stage wins

Freezing keeps K exactly, because nothing it uses is trained, and nothing frozen is about the arm. Training layer 1 and the head learns, but not within the budget, and the probe still falls. The learning rate sets the speed along the road, not the road: × 0.3, × 1 and × 3 end at distance 5.50, 5.03 and 4.29, with the probe at 39, 42 and 50 %, in the order of the distance. The damage is the price of the distance, and B needs the distance. K alone at the end brings K back and takes B away: whatever is trained last decides the end state, and §5 uses that.

The pull towards θK is the instructive failure, because it is the first repair aimed at the right quantity. Its spring has the same stiffness λ in every direction, but K is sensitive along a few directions and indifferent along the rest, and B needs some of the indifferent ones: a round spring stiff enough to protect K (λ = 0.03) stops B, and a slack one loses both. What is wanted is a spring that is stiff where HK is large and slack where it is zero. The K loss is exactly that spring, its own expansion above, and its gradient can be had whenever we like, because K's data are a stream.

Road not taken · wall off the weights
The cleanest repair makes forgetting impossible: freeze the language backbone, or let the action expert read it without sending gradients back, as Knowledge Insulating VLAs (Driess et al., 2025) do with a stop-gradient. K cannot move because nothing it uses is trained, and on the Bench the arm learns nothing (the head-only row). The real systems agree in kind: fine-tuning OpenVLA (Kim et al., 2024) on a new setup by its last layer alone gives 30.3 % success against 69.7 % for full fine-tuning, and Driess et al. report 0 % for a frozen backbone on two tasks. What works there is not a wall alone: their recipe also co-trains the backbone on vision-language data, and the expert has its own weights, the split of lesson 13. A wall protects K by teaching the arm nothing the frozen part cannot already say; the mixture lets the shared part learn what the robot needs.

4 · The move: a share ρ of K in every batch

A batch of 40 examples holds round(40ρ) from K and the rest from B, and its expected gradient is

(1 − ρ) ∇LB + ρ ∇LK, with ∇LK ≈ HK (θ − θK) near θK

which is the pull of B plus a spring of stiffness ρHK back to θK. The spring is stiff along the directions K is sensitive to and slack along the others, so B can move the weights wherever K does not mind. It steers the weights and does not stop them. Measured at ρ = 0.3, five seeds: the hidden layers end at distance 3.82, 0.76 times the 5.03 they travel without the mixture, and the K loss rises by 0.0047 against 0.220, 47 times less.

The price is steps. Each batch carries only a share 1 − ρ of B, so if progress depended on nothing but the number of B examples seen, B would need 1/(1 − ρ) times the steps, 1.43 at ρ = 0.3. Measured as the step at which success first reaches 50 %, over the four seeds that get there at both shares (the fifth does not at ρ = 0.3 within 600 steps), it grows from 300 to 438, a factor 1.46. At a fixed budget the same price reads as tasks not yet learned: after the corrections of §5 success is 89 % without the mixture and 82 % with it, 7.3 points lower (4.4 to 10.6, positive in all five seeds). With 2004 steps instead of 600 the gap closes, 88.6 % against 90.6 % when post-training ends, and the probe stands at 35.9 % against 97.5 %: on the Bench the price is time. The table gives the five-seed means for each share.

share ρprobe after post-traininglargest fall of the probesuccess after post-trainingsuccess after corrections
042 %54.8 points69 %89 %
0.191 %8.465 %85 %
0.293 %5.057 %84 %
0.395 %3.454 %82 %
0.596 %1.725 %66 %

Which share? The alarm of §6 is a fall of 5 points in the probe. Of the shares run (the table's, and 0.05, 0.4 and 0.7) the smallest for which no seed trips it is 30 %: the largest falls are 2.0 to 4.7 points, against up to 7.2 at 20 %. The grid, the budget and the 5-point alarm are choices, and the table is the evidence, not the 0.3.

The widget

Pretrain, post-train, correct: what does the model keep?
Top: the probe of K (teal), success (amber) and copy error (purple, dashed) against steps: stage 1 trains on K alone (compressed), stage 2 post-trains with a share ρ of K in every batch, stage 3 adds one round of corrections. The red line is 5 points under the probe's pretrained value, the red dot where the probe first crosses it. Bottom: the end of the recipe for every share (seed 1; dashed: success when post-training ends).
probe, after stage 1
—
probe, after post-training
—
probe, end of recipe
—
probe, words 0-59, end
—
probe, words 60-119, end
—
success, after post-training
—
success, end of recipe
—
copy error, after post-training
—
steps until success 50 %
—
probe alarm trips
—
weights moved, d
—
K-loss rise, ×10⁻³
—
Show the core JS
FL.zero(g);
if (nB > 0 && tab) for (b = 0; b < nB; b++) { var id = floor(rng() * tab.n); FL.grad(net, g, w, tab.X, id * FL.DIN, 1, 0, tab.Y, id * 2); }
for (b = 0; b < nK; b++) { var k = FL.draw(s1.voc, rng, x, nW); FL.grad(net, g, w, x, 0, 0, s1.voc.cls[k]); }
FL.adam(net, g, FL.lrAt(lr0, s, S), 1 / B, train);
...
for (var i = lo; i < hi; i++) { var gi = gw[i] * scale; m[i] = b1 * m[i] + (1 - b1) * gi; v[i] = b2 * v[i] + (1 - b2) * gi * gi; W[i] -= lr * (m[i] / c1) / (sqrt(v[i] / c2) + 1e-8); }
...
if (c[i].stage === 2 && out.alarm === null && c[i].s > 0 && c[i].k < k0 - FL.ALARM) out.alarm = c[i].s;

What to try. Leave ρ = 30 % and the recipe: the probe reads 98.5 % after stage 1 and 95.7 % at the end, success is 86.1 % after the corrections, and the alarm never trips. Success shows one early blip, 23.6 % at step 50 and 0.0 % at step 100: a half-trained policy that clears some layouts and loses them. Slide ρ to 0: the probe falls to 38.8 % after post-training and 34.5 % at the end, success reaches 92.4 %, the alarm trips at step 50 and the weights move 4.90. Now keep ρ = 0 and open the menu, which reruns the repairs of §3 on seed 1. Only the new head trains: the probe stays at 98.5 % and success at 0.0 %. K alone at the end takes the probe from 38.8 % to 96.3 % and success from 72.9 % to 0.0 %. Back at ρ = 30 %: the mixture that holds 60 of 120 words reads 99.7 % on those and 60.7 % on the others, and the alarm trips at step 50; no K in the corrections takes the probe from 95.7 % to 86.8 %.

5 · Stages: the order, and what each one carries

Three kinds of data differ in size and in what an example costs, and the order of the stages follows from that and from the result of §3 that the last stage wins.

stagedatasizewhat an example costs
1 pretrainK: broad, noisy renditionsan endless streamalmost nothing: scraped
2 post-trainB: calm, curated demonstrations8884 framesa person and a robot
3 correctstates the model itself reaches, labelled by the expert (lesson 3)2740 frames, 31 % of the demonstrationsthe most: a labeller for every state

Whatever is trained last decides the end state, so the data you trust most and pay most for go last, and the cheap abundant data go first, where they build what the later data will use. The same result says what each later stage must carry: one that does not carry the earlier data erases them. The widget's stage 3 keeps drawing K at the share ρ and adds the corrections to the demonstrations, as lesson 3 aggregated; dropping the mixture for those 300 short steps takes the probe from 95 % to 88 % (7.1 points over five seeds) and buys success only 3 points. Every stage is another chance to forget.

Stage 3 is lesson 3's recipe with one round: one rollout of the current model on each layout under the same gust, every visited state labelled by the expert, then 300 more steps at a learning rate of 0.002. It lifts success from 72.9 % to 92.4 % without the mixture and from 66.0 % to 86.1 % with ρ = 0.3 (seed 1; over five seeds by 20 and 28 points), and the probe stays at 95.7 %. Published recipes have the same shape. π0 (Black et al., 2024) pretrains on over 10,000 hours of robot data, diverse and of lower quality, which lets the model recover from mistakes, then post-trains on 5 hours for the simplest tasks and 100 or more for the hardest, data that show a consistent and fluent strategy. π0.5 (Physical Intelligence, 2025) pretrains for 280,000 steps and post-trains for 80,000 more. RT-2 (Brohan et al., 2023) co-fine-tunes with web data and, in its ablation, reaches 63 % on unseen objects, backgrounds and environments for its 55B model against 52 % for plain fine-tuning.

6 · Noticing the loss, and what the verdict costs

A probe is a set of questions about what the model used to know, asked at every checkpoint. Three rules follow from the measurements. Log it often. Without the mixture the alarm, a probe 5 points below its pretrained value, trips at step 50, 250 steps before success crosses 50 %; success is still 0 % then, and the copy error, at 57.4 %, is falling as it should: nothing tracked warns. Make it fresh and broad. The Bench's probe is 600 renditions no step trained on, from all 120 words. Keep it out of the mixture. In the widget's variant mixture holds 60 of 120 words, the mixture draws only words 0 to 59. The probe reads 99 % on those and 62 % on words 60 to 119 (five seeds): a probe made of the mixture's own words would have called this model undamaged. The mixture protects what it contains and, for a table of arbitrary facts, little else.

The probe is cheap: 600 forward passes, about 0.6 million multiply-adds. What the recipe costs is success, which is not. The price of ρ = 0.3 is 7.3 points after the corrections (89 % against 82 %, five seeds of 360 rollouts each). A laboratory that runs one trial per layout on each checkpoint, 36 trials, measures the default model's 86 % as 31 of 36, a 95 % Wilson interval of 71.3 to 93.9 % with a half-width of 11.3 points: the price sits inside the error bar. Even the widget's 144 rollouts (79.5 to 90.8 %) have a half-width of 5.7 points, most of the price. The probe separates 96 % from 41 % in milliseconds; the success that justifies what the recipe costs is invisible at the number of trials a robot can afford.

What this lesson did not do
The network has 1212 weights. It shows what the damage depends on and what a mixture does about it, not how much a model of billions of weights forgets, nor that 0.3 transfers. K is arbitrary facts, the hardest case for a mixture that holds only a sample; knowledge with structure may share more between words, and that was not tested. K does not stand for what RT-2's co-training is for, generalisation to unseen objects and instructions. The mixture needs the old data to still exist; with only the old weights, the wall and the penalty of §3 are what is left. One optimiser, one schedule and one budget were used, and the price in tasks depends on the budget. Low-rank adapters, averaged weights and penalties that estimate HK were not computed. How to judge a checkpoint is lesson 15; what each hour of data is worth and costs is lessons 16 and 17.

Common mistakes / failure modes

"the loss fell and success rose, so the fine-tune worked"
Both did, and the probe fell from 98.5 % to 38.8 % (§2).
"a smaller learning rate prevents forgetting"
At × 0.3 the weights end at distance 5.50 and the probe at 39 %: the damage is the distance, not the speed (§3).
"freeze the backbone and K is safe"
It is (97 %), and success is 0 %: nothing frozen is about the arm (§3).
"a few percent of old data is enough"
At ρ = 0.1 the probe still falls 8.4 points at its worst checkpoint (five-seed mean), and at 0.05 12.0 (§4).
"the probe reads 99 %, so nothing was lost"
Not if it is made of words the mixture draws: 99 % there, 62 % on the rest (§6).

Checkpoint exercise

Try it
A batch holds 40 examples. (a) With ρ = 0.3, how many are K and how many are B? (b) Without the mixture, seed 1 first reaches 50 % success at step 300, after 300 × 40 B examples. If only the number of B examples mattered, at what step would it get there with ρ = 0.3? (c) Compare (b) with the default widget run. Answer: (a) 12 K and 28 B. (b) 300 × 40 / 28 = 429 steps. (c) The widget's default run gets there at step 500, a factor 1.67 on 300 against the count's 1.43; over the four seeds that reach it at both shares the factor is 1.46, about the dilution 1/(1 − ρ), and curves read every 50 steps cannot say more.

Where this points next

Staging the training and keeping the old data in every batch preserves what the model started with: at ρ = 0.3 the probe ends at 96 % where plain fine-tuning leaves 41 % (five-seed means). The price is steps, a factor of about 1.5, and 7.3 points of success after the corrections at a fixed budget. The probe that finds the damage costs 0.6 million multiply-adds, but the evaluation that would show the price cannot resolve it at the trials a robot can afford: 36 trials give 71.3 to 93.9 %. We now have a recipe and a checkpoint, and need to say whether the checkpoint is better than the last one. An evaluation is a count of trials, and each trial costs a robot and a person for minutes. How many trials does it take to tell one policy from another, and what does that do to the cost of the whole programme?

Takeaway
Fine-tuning moves the weights as far as the new task needs and charges the old task for the distance, to second order, along the directions the old task is sensitive to. On the Bench it takes success to 72.9 % and the probe from 98.5 % to 38.8 %, and the probe starts falling before success moves. Freezing, a smaller learning rate, K alone at the end and a pull towards the old weights fail by computation: they change the speed, stop the learning, lose B, or hold B as hard as K. What works is a share ρ of the old data in every batch of every later stage, a spring made of the old task's own loss: at ρ = 0.3 the weights move 0.76 times as far and the K loss rises 47 times less, for a price of about 1/(1 − ρ) in steps. Order the stages by trust, the dearest data last and the earlier data carried along, and log a probe the mixture cannot see: it trips the alarm at step 50, long before success moves.

Interview prompts

Companion reads: World Models · 26 The curriculum (when staging beats joint training), World Models · 27 The mixture is the model (how mixture weights are set) and World Models · 28 Post-training a world model (what post-training can and cannot add).