all_lessons/Robot Model Training/08 · Other bodieslesson 8 / 24

Other bodies: write the action where the hand goes

The last lesson left a policy that needs about a hundred layouts to reach ninety percent, a price one arm cannot pay for a real scene. Other laboratories have recorded demonstrations on other arms, and the file format invites you to concatenate them. This lesson shows why that is worse than leaving them out: a joint angle is a measurement in the units of the body that took it, and a two percent difference in a link length moves the hand by the whole margin of the task. It rules out raw pooling and a body tag by computation, writes the action as the place the hand should go with each arm supplying its own inverse kinematics and tracker, prices a foreign layout, and shows what an older arm's tempo does to that price. It ends at the size of the pool.

The thesis, here
A stored action means something only relative to the body that produced it. Written in joint angles, a label is in that body's units, and the task's margin is smaller than the error a slightly different arm makes in them. Written as the place the hand should go, the label means the same to every arm, and each arm's own inverse kinematics and tracker turn it into its own joint motions. Then a layout another laboratory recorded is, to a good approximation, a layout of yours.
Linear position
Forced by: For a policy that generalises only to what resembles its data, the error on a new layout is the distance to the nearest demonstrated layout, so success grows with the number of different layouts, with a power set by how many numbers describe a layout, and hardly at all with more demonstrations of a layout already seen. Ninety percent costs about a hundred layouts when two numbers describe a layout, and every further number multiplies that by the range it varies over divided by the width one layout covers; a real scene is described by many more than two numbers, and one arm cannot pay for that. Other laboratories have already collected such data, on other arms. What can a robot use from demonstrations made on a body that is not its own?
New idea: a demonstration made on another body is a correct label for yours only if the action is written in a space every body shares, the place the hand should go, with each arm supplying its own inverse kinematics and tracker. Pooled that way a foreign layout is worth about one of your own; pooled in joint angles it is worth less than none.
Forces next: Demonstrations from other bodies help when the action is written in a space every body shares, the place the hand should go, and each arm turns that into its own joint motions; written in joint angles they are worse than no data. Pooled this way, all the robot demonstrations ever recorded are a small fraction of what people have filmed themselves doing, and the video has no actions in it. What can be learned from watching, and what is missing from the video that no amount of it can supply?
The plan
Six moves. (1) Count what your own layouts buy and what the others have. (2) Measure how wrong a joint-angle label from another arm is. (3) Compare three ways to pool and rule two out. (4) Write the label where the hand goes, give each body its own inverse kinematics and tracker, and price a foreign layout. (5) Let the other body move differently and put a price on trust. (6) Count what the pool cannot buy.

1 · What your own layouts buy

Everything runs on the Bench of lessons 1 to 7. The policy is the lookup of lesson 7: given the layout θ = (s, τ), shift in metres and tilt in metres per metre of the row of posts, copy the stored layout nearest to it, 16 absolute targets at a time, tracked by the stiff controller of lesson 6, u = K (q* − q), where q are the joint angles, q* the target, u the joint velocity (K = 20 /s, clipped at 1.5 rad/s). The distance between two layouts is how far the farthest of the five posts moves, |Δs| + 0.4 |Δτ| metres. Arm A is the worn arm of lesson 6 (it realises 85 % of a command) in a gust of 0.08 rad/s, and layouts are drawn uniformly from |s| ≤ 10 cm, |τ| ≤ 0.2. Success is the share of 1,000 fresh layouts, one run each, that A completes without touching a post.

With 16 layouts of its own, one demonstration each, A completes the course on 34.0 % of fresh layouts. That draw of 16 is typical (it is the one of twelve whose curve lies nearest the mean): over the twelve random draws the success runs from 25.3 to 40.7 %, mean 34.4 %. Ninety percent takes about a hundred layouts: 104 in lesson 7 (64 random sets, a grid of 320 test layouts), 107 here on the mean curve of twelve draws tested on 1,000 random layouts, 85 to 124 on single draws. At assumed prices, 9.3 s for the expert's run (lesson 1), 20 s to put the cup back and 120 s to rebuild the layout, 149.3 s a layout, the 16 layouts cost 40 minutes and the 107 cost 4.4 hours of one operator on one arm.

That is for two numbers per layout. With the shift alone ninety percent takes 15.2 layouts (14 in lesson 7), so the second number multiplied the count by 7.0, and every further number multiplies it again by the range it varies over divided by the width one layout covers. A real scene has many more than two. The layouts somebody else has already recorded are a way out, and §6 puts a number on how far that goes. First, what does a demonstration made on another arm say to this one?

2 · A joint angle is a measurement in someone else's units

A frame in the other laboratory's file is a pair: what its arm sensed, and where it was heading. Lesson 6 made that label a target, an absolute position tracked by a stiff controller, and in the other file the target is what its controller was given, joint angles q*. Arm A can track them. Its hand then goes to FKA(q*) (forward kinematics: the hand's position for given joint angles) while the other arm's hand went to FKB(q*), and the two differ whenever the arms do. For two-link arms with the same total reach whose links differ by δ, B = (L1 + δ, L2 − δ) against A = (L1, L2), the difference of the two hands is one line:

FKB(q) − FKA(q) = δ · [ (cos q1, sin q1) − (cos(q1 + q2), sin(q1 + q2)) ], length 2 |δ| sin(q2 / 2)

with q2 the elbow angle. Along the course it averages 101°, so the factor 2 sin(q2/2) averages 1.52: the hand misses by about one and a half times the difference in link length. The margin is small. A's own demonstrations clear the posts by 1.32 cm on average over 64 layouts (1.55 cm on the nominal layout of lesson 6) and by 1.04 cm at the closest, so a joint label stays usable only while 1.52 |δ| is under 1.32 cm: |δ| below 0.87 cm, 1.7 % of a 50 cm link. Measured on the first 16 demonstrations of each arm in the other laboratory's file, with A tracking the stored joint angles:

other armlinks (m)A's hand against theirs: mean (largest)
same make as A0.50 + 0.500.0 cm
a hair different0.51 + 0.491.5 cm (1.8)
another arm of the same size0.55 + 0.457.6 cm (9.3)

A velocity label does no better: at the same hand position the other arm's joint command moves A's hand in a direction rotated by 1.0° (links 0.51 + 0.49) to 5.0° (0.55 + 0.45), a drift of 1.9 to 9.6 cm over the 1.1 m of the course once it integrates (lessons 5 and 6). A joint angle is a number in the units of the body that took it, and two arms that look the same to a person disagree by more than the task allows when their links differ by two percent.

3 · Three ways to pool, and what each costs

The constraints are sharp. A foreign frame must be a correct label for A in the same situation, meaning that A's hand does what theirs did; the pooled file must stay something one lookup can read; and nothing that tells A where to go may need the other body's dimensions when A runs. Three designs, with A's 16 layouts and 64 layouts of the arm with links 0.55 + 0.45:

designwhat it needsA completes the coursewhat happens
(a) concatenate the files: joint-angle labels, joint-angle keysnothing17.9 % (A alone: 34.2 %)a foreign layout is picked in 80 % of the runs and 82 % of all runs end in a collision
(b) add the arm's name to every key, so a foreign frame is never the nearesta tag on every frame34.2 %, exactly A aloneno foreign frame is ever read: safe, and their hours buy nothing
(c) write the label as the place the hand goes (§4)an inverse kinematics per body83.2 %the label is correct for A by construction

Why (a) is worse than nothing, not merely useless. Let each stored layout cover a fraction f of the box: n random layouts leave a fresh one uncovered with probability (1 − f)n ≈ e−nf, so success is about 1 − e−nf. Your N layouts and the other laboratory's M are random points in the box, and among random points of two kinds the kind of the nearest does not depend on how near it is: it is yours with probability N/(N + M) when the lookup cannot tell the kinds apart. Foreign joint labels fail, so only that share can succeed:

successjoint(M) ≈ (1 − e−(N+M)f) · N/(N + M) < 1 − e−Nf for every M > 0

The inequality holds because (1 − e−x)/x falls as x grows: the more raw foreign data, the worse. The Bench agrees. At N = 16, M = 64 the nearest layout is one of A's for 20.3 % of fresh layouts against N/(N + M) = 0.2, and the hand-labelled pool's 83.2 % × 0.2 = 16.6 % against 17.9 % measured in joint angles; for every M from 4 to 256 that product lies within 3.3 points of the joint-angle result. Design (b) removes the harm and the benefit together: with the name in the key no foreign frame is ever nearest, and success is A's 34.2 % to the digit (the same 16 layouts give 34.0 % in §1 and §4, where the lookup keys on the hand's place).

Road not taken · concatenate the files and let a large network sort out the bodies
At scale, concatenating is not absurd. π0 (Physical Intelligence, 2024) keeps each of its seven robot configurations' own configuration and action vectors, zero-padded to 18 numbers, and trains one 3.3-billion-parameter network on over 10,000 hours of its own robots' data, leaving the network to learn what a vector means on which body. That pays for the separation in data and capacity: a lookup cannot take that road (design (b)), and a laboratory with sixteen layouts cannot pay for it.

4 · Write the label where the hand goes

What every arm shares is the table. The place the hand should go, p* = (x, y) in metres, origin at the arm's base, x to the right and y up, means the same to every arm: it names a point, not a configuration. So write the label that way, and let each body turn the point into its own joint motions with two pieces it already owns. The first is inverse kinematics; for a two-link arm with links L1, L2 and the elbow on the branch it is on now,

cos q2* = (|p*|² − L1² − L2²) / (2 L1 L2), q1* = atan2(p*y, p*x) − atan2(L2 sin q2*, L1 + L2 cos q2*)

The second is the stiff tracker of lesson 6, u = K (q* − q), which stays specific to the body: its stiffness must respect its own delay (lesson 6). The stored frame now holds the hand's place and the layout, and nothing of the other arm's joints, links or commands; the lookup keys on the hand's place too. A point is usable when it lies inside the arm's reach, |L1 − L2| ≤ |p*| ≤ L1 + L2, which the course satisfies for every arm here.

The same experiment with hand labels: at 64 foreign layouts A completes the course on 83.2 % of fresh layouts (95 % interval 80.8 to 85.4), up from 34.0 %. The arm with links 0.51 + 0.49 gives 83.1 % and the same make gives 83.1 %: an expert draws the same path with the same hand whichever arm holds it, so in the hand's space these are the same data. At 256 foreign layouts success is 98.9 %.

What is a foreign layout worth? Put it in the unit the laboratory pays in. Your own curve gives success as a function of your number of layouts N (the dashed line in the widget, on the same 1,000 test layouts). A pool of N own and M foreign layouts reaches some success; read off the own curve the number of own layouts Neq that would have reached it, and define the exchange rate

r = (Neq − N) / M own layouts that one foreign layout is worth

so that r = 1 means as good as your own, 0 worthless and below 0 harmful. (A layout costs the same time on either arm here, so r is also own hours per foreign hour.) A's own curve reaches 83.2 % at about 76 layouts, so hand labels give r = (76 − 16)/64 = 0.94. The same pool in joint angles gives 17.9 %, which A's own curve reaches with 0.8 layouts: the 64 foreign layouts have cost A nearly all of its 16 (r = −0.24; the floor is −16/64 = −0.25, A with no layouts of its own). An arm of the same make is fine in either space (r = 0.88 in joint angles), which is why a fleet of identical robots, like the more than 100 behind AgiBot World (2025), can stay in joint space. Over twelve draws of both layout sets the hand-label rate averages 1.06 (0.72 to 1.24) and the joint-label rate −0.16: in the shared space a foreign layout is worth one of yours, in joint angles it is worth less than none. With 16 own layouts A needs 88 foreign ones to reach ninety percent, where its own curve needs 98 in all: the 88 stand in for 82 of A's own.

Open X-Embodiment (2023) does this at scale: 60 datasets from 22 robot embodiments, over a million trajectories, every action converted to a seven-number end-effector vector in a coarsely aligned space. RT-1-X, trained on the pool, beat each laboratory's own method on four of five small-data datasets. The authors also mark the limit: the frames are not aligned across datasets, so the same vector can mean different motions on different robots. A shared space is shared only as far as it is calibrated.

5 · Bodies differ in more than their joints

Even in the hand's space the other arm's demonstrations carry two things that are its own: its clock and its servo's wobble. The stored targets are indexed by its clock, so A replays them at that arm's pace, wobble included. The Bench has an older arm with the same links as A: it realises 85 % of a command, answers 0.1 s late, and its servo has a gust of 0.12 rad/s. Of 256 attempted layouts, 197 yield a demonstration that completes the course; the other 59 touch a post and a laboratory discards them. Take the two differences apart. The slower, later plant alone gives demonstrations of 226 frames against 185 for A's own, inside A's allowance of 235 steps, and A completes 100 % of the runs that follow them. The extra gust alone gives 187 frames and 79 % of A's runs complete: A copies the wobble and the wobble eats the margin. With both, 19 % of the kept demonstrations are longer than 235 steps (mean 227), and A cannot finish when it replays those.

Now measure the effect on A. Let A follow one stored layout on test layouts at a layout distance below 1.2 cm, where it should succeed: with A's own layouts 99 % of 237 runs complete the course; with the older arm's layouts only 63 % do, 11 % end in a collision and 26 % run out of steps. The exchange rate follows: r = 0.32 at 64 attempted layouts on this draw and 0.25 averaged over twelve (between 0.15 and 0.33): an hour with the older arm buys about a quarter of an hour of A's own.

The lookup cannot see quality. When a layout of the older arm is nearer than one of A's by a hair, it wins, even though A's own would have completed the course. The remedy is a price on trust. Let a foreign layout count 1/w times as far away as an own one: w = 1 trusts it like your own and w = 0 never uses it, which is the tag of design (b). With 64 own layouts A completes the course on 79.0 % of fresh layouts. Adding 64 attempted layouts of the older arm at w = 1 gives 77.4 %: they cost 1.6 points. At w = 0.6, the best weight in a sweep from 0 to 1, success is 81.8 %, 4.4 points above full trust and 2.8 above ignoring the arm. With 16 own layouts the best weight is 1.0: when your own layouts are thin, even a weak foreign layout is worth following. The weight is set on held-out layouts, and it moves with how much data you already have.

The widget

Pool another arm's demonstrations: which space, which arm, how much trust
Top left: one demonstration of the other arm (purple) and arm A (dark) tracking its labels; with joint-angle labels A's hand goes where the red dashed path is, with the hand's place A's own inverse kinematics put it on the purple path. Top right: success on 1,000 fresh layouts against the other laboratory's layouts for both kinds of label (dashed: A's own layouts alone; amber dot: this setting). Bottom: success against the trust weight w. The first control is the number of foreign layouts.
A completes the course
—
95 % interval
—
A's own layouts alone
—
gain over own layouts alone
—
exchange rate r
—
joint-label gap (hand vs theirs)
—
runs following a foreign layout
—
foreign demonstrations kept
—
runs touching a post
—
runs out of steps
—
best weight in the sweep
—
Show the core JS
BL.pick = function (own, foreign, wf, theta) {
  var best = null, bd = Infinity, i, d;
  for (i = 0; i < own.length; i++) { d = BL.dist(theta, own[i].theta); if (d < bd) { bd = d; best = own[i]; } }
  if (wf > 0) for (i = 0; i < foreign.length; i++) { d = BL.dist(theta, foreign[i].theta) / wf; if (d < bd) { bd = d; best = foreign[i]; } }
  return best;
};
...
    if (space === 'joint') tq = [Q[2 * i], Q[2 * i + 1]];
    else { el = q[1] >= 0 ? 1 : -1; tq = BN.arm.ik([PP[2 * i], PP[2 * i + 1]], el, me); }
    u0 = Math.max(-VMAX, Math.min(VMAX, K * (tq[0] - q[0]))); u1 = Math.max(-VMAX, Math.min(VMAX, K * (tq[1] - q[1])));
    q[0] += DT * (u0 * g + gust * BN.randn(rng)); q[1] += DT * (u1 * g + gust * BN.randn(rng));
...
BL.exchangeRate = function (curve, nOwn, m, sPooled) { return m > 0 ? (BL.equivalentOwn(curve, sPooled) - nOwn) / m : NaN; };

What to try. Leave the defaults: 16 own layouts, the arm with links 0.55 + 0.45, 64 of its layouts, hand labels, w = 1. A completes the course on 83.2 % of fresh layouts (80.8 to 85.4 %) against 34.0 % alone, a gain of 49.2 points at an exchange rate of 0.94. Switch to joint angles at 64: 17.9 % against 34.2 % alone, 80 % of the runs follow a foreign layout, 82 % end in a collision, the gap reads 7.6 cm against a margin of 1.32 cm, and the rate is −0.24; at 256 layouts it falls to 9.1 %, near the 16/272 = 5.9 % of runs that follow A's own layouts. Turn the weight to 0, the tag of design (b): back to 34.2 %. Links 0.51 + 0.49 in joint angles still fail, 21.9 % with a gap of 1.5 cm; the same make works, 82.0 % against 83.1 % with hand labels. Choose the older arm with hand labels: 58.8 % and a rate of 0.32. With A's own layouts at 64 and 64 older-arm layouts: w = 1 gives 77.4 % against 79.0 % alone, w = 0.6 gives 81.8 %.

6 · What pooling cannot buy

Pooling turns other laboratories' hours into yours at close to one for one when their arms move like yours. It does not change how fast the need grows. If each further number per layout multiplies the layouts needed for ninety percent by the same factor, 7.0 as measured between one number and two, the need for scenes of d numbers is:

numbers per layoutlayouts for 90 %one operator at Bench pricesshare of Open X-Embodiment's million trajectories
21074.4 h0.011 %
45,233217 h (5.4 working weeks)0.5 %
6256,42710,635 h26 %
71,794,98874,442 h179 %

Take the most generous reading, one different layout per trajectory and every trajectory worth a layout of your own (r = 1, which §5 shows is generous). All of Open X-Embodiment (2023), over a million trajectories, then covers scenes of 6 numbers and not seven; at the factor of 5.8 per number that lesson 7 derives from footprints it covers 7 and not eight. Five objects on a table are already ten numbers. In hours the pool is smaller than it sounds: over 10,000 hours of π0's own data (2024), 2,976 hours in the 1,001,552 trajectories of AgiBot World (2025), about 2,000 hours in Open X-Embodiment by the AgiBot authors' count. Add them for a rough total of 14,976 hours: arithmetic on three published figures, not a census. As a yardstick it is the footage of 41 people each recording one hour of their day for a year. This lesson measures no video, but 41 people are a handful beside everyone who films what they do. And a film has no action channel: it shows where a hand went, not the command that sent it there or the force behind it.

What this lesson did not do
It used two-link planar arms that differ in link lengths and servo quality. Real bodies also differ in degrees of freedom (a redundant arm needs a choice of inverse kinematics), grippers, cameras and control rates. It assumed the hand's frame is shared exactly; base placement and camera calibration shift it, and an offset eats the margin one for one. Its policy reads the hand's place and the layout directly, while one that reads pictures must also learn to see another arm and camera as the same task (lesson 14 treats mixtures). One fixed draw gives the numbers, and the spread between draws is real: 4.6 points at 16 own layouts. A lookup cannot learn what tagged bodies share, so the Bench cannot show what a network recovers. Time, force and effort are not in a place the hand went: tempo hurt in §5, and what a camera cannot see is lesson 10.

Common mistakes / failure modes

"more demonstrations from any robot can only help"
Joint-angle labels from the 0.55 + 0.45 arm take A from 34.2 % to 17.9 % at 64 layouts and 9.1 % at 256 (§3).
"tag the data with the robot's name and pooling is safe"
A lookup then never reads the foreign frames: 34.2 %, exactly no data. A network can learn to share across tags and pays in data and capacity (§3).
"end-effector actions make all bodies the same"
They fix where, not when or how steadily: the older arm's layouts are worth 0.25 of yours, against 1.06 for an arm that moves like yours (§4, §5).
"pool every laboratory's data and the layout problem is solved"
The need multiplies by about 7.0 per number: six numbers need 256,427 layouts, 26 % of a million trajectories (§6).

Checkpoint exercise

Try it
Arm A has links 0.50 + 0.50 and arm B 0.53 + 0.47, so δ = 3 cm, and the elbow is at 90° at the moment of interest. (a) A tracks B's stored joint angles: how far does A's hand miss B's? (b) Is that inside a margin of 1.3 cm? (c) What link-length difference does this elbow angle allow, in cm and as a share of a 50 cm link? (d) A pool of 16 own and 64 foreign layouts reaches the success that A's own curve gives at 70 layouts: what is the exchange rate? Answer: (a) 2 δ sin(q2/2) = 2 × 3 cm × sin 45° = 4.2 cm. (b) No: it is more than three times the margin. (c) δ = 1.3 cm / (2 sin 45°) = 0.92 cm, 1.8 % of the link. (d) r = (70 − 16)/64 = 0.84: a foreign layout is worth a little under one of A's own.

Where this points next

Written in joint angles, another arm's demonstrations take A from 34.2 % to 17.9 % at 64 layouts. Written as the place the hand should go, with A's own inverse kinematics and tracker, they take it to 83.2 %, and a foreign layout is worth 1.06 of A's own on average. That is still not enough for a real scene: every further number per layout multiplies the need by about 7.0, six numbers need 256,427 layouts, and the robot collections named here add up to roughly 14,976 hours, the footage of 41 people filming one hour a day for a year. People film themselves doing things all the time; that film shows where a hand went, and has no command and no force in it. What can be learned from watching, and what is missing from the video that no amount of it can supply?

Takeaway
A joint angle is a measurement in the units of the body that took it, so a label in joint angles from another arm is wrong by 2|δ| sin(q2/2) for a link-length difference δ, which is more than the task's margin once δ passes 1.7 % of a link. Concatenating such files is worse than leaving them out (17.9 % against 34.2 % at 16 own and 64 foreign layouts), because the nearest layout is a foreign one with probability M/(N + M); a body tag removes the harm and the benefit together. Writing the label as the place the hand should go, and letting each arm's own inverse kinematics and tracker do the rest, makes a foreign layout worth about one of yours (83.2 %, exchange rate 0.94). A body that moves differently is worth less, the older arm 0.25, and its data need a weight below one once you have layouts of your own. Pooling does not change how fast the need grows, about 7.0 times per number of the layout, and the robot collections named here add up to roughly 14,976 hours.

Interview prompts

Companion reads: Reinforcement Learning · 56 Robot control (kinematics and the Jacobian: the per-body part), Computer Vision 3D · 03 Rigid motion (frames and offsets: why a shared space is only as shared as its calibration) and World Models · 27 Data mixture (weights on sources, as in §5).