Other bodies: write the action where the hand goes
The last lesson left a policy that needs about a hundred layouts to reach ninety percent, a price one arm cannot pay for a real scene. Other laboratories have recorded demonstrations on other arms, and the file format invites you to concatenate them. This lesson shows why that is worse than leaving them out: a joint angle is a measurement in the units of the body that took it, and a two percent difference in a link length moves the hand by the whole margin of the task. It rules out raw pooling and a body tag by computation, writes the action as the place the hand should go with each arm supplying its own inverse kinematics and tracker, prices a foreign layout, and shows what an older arm's tempo does to that price. It ends at the size of the pool.
New idea: a demonstration made on another body is a correct label for yours only if the action is written in a space every body shares, the place the hand should go, with each arm supplying its own inverse kinematics and tracker. Pooled that way a foreign layout is worth about one of your own; pooled in joint angles it is worth less than none.
Forces next: Demonstrations from other bodies help when the action is written in a space every body shares, the place the hand should go, and each arm turns that into its own joint motions; written in joint angles they are worse than no data. Pooled this way, all the robot demonstrations ever recorded are a small fraction of what people have filmed themselves doing, and the video has no actions in it. What can be learned from watching, and what is missing from the video that no amount of it can supply?
1 · What your own layouts buy
Everything runs on the Bench of lessons 1 to 7. The policy is the lookup of lesson 7: given the layout θ = (s, τ), shift in metres and tilt in metres per metre of the row of posts, copy the stored layout nearest to it, 16 absolute targets at a time, tracked by the stiff controller of lesson 6, u = K (q* − q), where q are the joint angles, q* the target, u the joint velocity (K = 20 /s, clipped at 1.5 rad/s). The distance between two layouts is how far the farthest of the five posts moves, |Δs| + 0.4 |Δτ| metres. Arm A is the worn arm of lesson 6 (it realises 85 % of a command) in a gust of 0.08 rad/s, and layouts are drawn uniformly from |s| ≤ 10 cm, |τ| ≤ 0.2. Success is the share of 1,000 fresh layouts, one run each, that A completes without touching a post.
With 16 layouts of its own, one demonstration each, A completes the course on 34.0 % of fresh layouts. That draw of 16 is typical (it is the one of twelve whose curve lies nearest the mean): over the twelve random draws the success runs from 25.3 to 40.7 %, mean 34.4 %. Ninety percent takes about a hundred layouts: 104 in lesson 7 (64 random sets, a grid of 320 test layouts), 107 here on the mean curve of twelve draws tested on 1,000 random layouts, 85 to 124 on single draws. At assumed prices, 9.3 s for the expert's run (lesson 1), 20 s to put the cup back and 120 s to rebuild the layout, 149.3 s a layout, the 16 layouts cost 40 minutes and the 107 cost 4.4 hours of one operator on one arm.
That is for two numbers per layout. With the shift alone ninety percent takes 15.2 layouts (14 in lesson 7), so the second number multiplied the count by 7.0, and every further number multiplies it again by the range it varies over divided by the width one layout covers. A real scene has many more than two. The layouts somebody else has already recorded are a way out, and §6 puts a number on how far that goes. First, what does a demonstration made on another arm say to this one?
2 · A joint angle is a measurement in someone else's units
A frame in the other laboratory's file is a pair: what its arm sensed, and where it was heading. Lesson 6 made that label a target, an absolute position tracked by a stiff controller, and in the other file the target is what its controller was given, joint angles q*. Arm A can track them. Its hand then goes to FKA(q*) (forward kinematics: the hand's position for given joint angles) while the other arm's hand went to FKB(q*), and the two differ whenever the arms do. For two-link arms with the same total reach whose links differ by δ, B = (L1 + δ, L2 − δ) against A = (L1, L2), the difference of the two hands is one line:
FKB(q) − FKA(q) = δ · [ (cos q1, sin q1) − (cos(q1 + q2), sin(q1 + q2)) ], length 2 |δ| sin(q2 / 2)
with q2 the elbow angle. Along the course it averages 101°, so the factor 2 sin(q2/2) averages 1.52: the hand misses by about one and a half times the difference in link length. The margin is small. A's own demonstrations clear the posts by 1.32 cm on average over 64 layouts (1.55 cm on the nominal layout of lesson 6) and by 1.04 cm at the closest, so a joint label stays usable only while 1.52 |δ| is under 1.32 cm: |δ| below 0.87 cm, 1.7 % of a 50 cm link. Measured on the first 16 demonstrations of each arm in the other laboratory's file, with A tracking the stored joint angles:
| other arm | links (m) | A's hand against theirs: mean (largest) |
|---|---|---|
| same make as A | 0.50 + 0.50 | 0.0 cm |
| a hair different | 0.51 + 0.49 | 1.5 cm (1.8) |
| another arm of the same size | 0.55 + 0.45 | 7.6 cm (9.3) |
A velocity label does no better: at the same hand position the other arm's joint command moves A's hand in a direction rotated by 1.0° (links 0.51 + 0.49) to 5.0° (0.55 + 0.45), a drift of 1.9 to 9.6 cm over the 1.1 m of the course once it integrates (lessons 5 and 6). A joint angle is a number in the units of the body that took it, and two arms that look the same to a person disagree by more than the task allows when their links differ by two percent.
3 · Three ways to pool, and what each costs
The constraints are sharp. A foreign frame must be a correct label for A in the same situation, meaning that A's hand does what theirs did; the pooled file must stay something one lookup can read; and nothing that tells A where to go may need the other body's dimensions when A runs. Three designs, with A's 16 layouts and 64 layouts of the arm with links 0.55 + 0.45:
| design | what it needs | A completes the course | what happens |
|---|---|---|---|
| (a) concatenate the files: joint-angle labels, joint-angle keys | nothing | 17.9 % (A alone: 34.2 %) | a foreign layout is picked in 80 % of the runs and 82 % of all runs end in a collision |
| (b) add the arm's name to every key, so a foreign frame is never the nearest | a tag on every frame | 34.2 %, exactly A alone | no foreign frame is ever read: safe, and their hours buy nothing |
| (c) write the label as the place the hand goes (§4) | an inverse kinematics per body | 83.2 % | the label is correct for A by construction |
Why (a) is worse than nothing, not merely useless. Let each stored layout cover a fraction f of the box: n random layouts leave a fresh one uncovered with probability (1 − f)n ≈ e−nf, so success is about 1 − e−nf. Your N layouts and the other laboratory's M are random points in the box, and among random points of two kinds the kind of the nearest does not depend on how near it is: it is yours with probability N/(N + M) when the lookup cannot tell the kinds apart. Foreign joint labels fail, so only that share can succeed:
successjoint(M) ≈ (1 − e−(N+M)f) · N/(N + M) < 1 − e−Nf for every M > 0
The inequality holds because (1 − e−x)/x falls as x grows: the more raw foreign data, the worse. The Bench agrees. At N = 16, M = 64 the nearest layout is one of A's for 20.3 % of fresh layouts against N/(N + M) = 0.2, and the hand-labelled pool's 83.2 % × 0.2 = 16.6 % against 17.9 % measured in joint angles; for every M from 4 to 256 that product lies within 3.3 points of the joint-angle result. Design (b) removes the harm and the benefit together: with the name in the key no foreign frame is ever nearest, and success is A's 34.2 % to the digit (the same 16 layouts give 34.0 % in §1 and §4, where the lookup keys on the hand's place).
4 · Write the label where the hand goes
What every arm shares is the table. The place the hand should go, p* = (x, y) in metres, origin at the arm's base, x to the right and y up, means the same to every arm: it names a point, not a configuration. So write the label that way, and let each body turn the point into its own joint motions with two pieces it already owns. The first is inverse kinematics; for a two-link arm with links L1, L2 and the elbow on the branch it is on now,
cos q2* = (|p*|² − L1² − L2²) / (2 L1 L2), q1* = atan2(p*y, p*x) − atan2(L2 sin q2*, L1 + L2 cos q2*)
The second is the stiff tracker of lesson 6, u = K (q* − q), which stays specific to the body: its stiffness must respect its own delay (lesson 6). The stored frame now holds the hand's place and the layout, and nothing of the other arm's joints, links or commands; the lookup keys on the hand's place too. A point is usable when it lies inside the arm's reach, |L1 − L2| ≤ |p*| ≤ L1 + L2, which the course satisfies for every arm here.
The same experiment with hand labels: at 64 foreign layouts A completes the course on 83.2 % of fresh layouts (95 % interval 80.8 to 85.4), up from 34.0 %. The arm with links 0.51 + 0.49 gives 83.1 % and the same make gives 83.1 %: an expert draws the same path with the same hand whichever arm holds it, so in the hand's space these are the same data. At 256 foreign layouts success is 98.9 %.
What is a foreign layout worth? Put it in the unit the laboratory pays in. Your own curve gives success as a function of your number of layouts N (the dashed line in the widget, on the same 1,000 test layouts). A pool of N own and M foreign layouts reaches some success; read off the own curve the number of own layouts Neq that would have reached it, and define the exchange rate
r = (Neq − N) / M own layouts that one foreign layout is worth
so that r = 1 means as good as your own, 0 worthless and below 0 harmful. (A layout costs the same time on either arm here, so r is also own hours per foreign hour.) A's own curve reaches 83.2 % at about 76 layouts, so hand labels give r = (76 − 16)/64 = 0.94. The same pool in joint angles gives 17.9 %, which A's own curve reaches with 0.8 layouts: the 64 foreign layouts have cost A nearly all of its 16 (r = −0.24; the floor is −16/64 = −0.25, A with no layouts of its own). An arm of the same make is fine in either space (r = 0.88 in joint angles), which is why a fleet of identical robots, like the more than 100 behind AgiBot World (2025), can stay in joint space. Over twelve draws of both layout sets the hand-label rate averages 1.06 (0.72 to 1.24) and the joint-label rate −0.16: in the shared space a foreign layout is worth one of yours, in joint angles it is worth less than none. With 16 own layouts A needs 88 foreign ones to reach ninety percent, where its own curve needs 98 in all: the 88 stand in for 82 of A's own.
Open X-Embodiment (2023) does this at scale: 60 datasets from 22 robot embodiments, over a million trajectories, every action converted to a seven-number end-effector vector in a coarsely aligned space. RT-1-X, trained on the pool, beat each laboratory's own method on four of five small-data datasets. The authors also mark the limit: the frames are not aligned across datasets, so the same vector can mean different motions on different robots. A shared space is shared only as far as it is calibrated.
5 · Bodies differ in more than their joints
Even in the hand's space the other arm's demonstrations carry two things that are its own: its clock and its servo's wobble. The stored targets are indexed by its clock, so A replays them at that arm's pace, wobble included. The Bench has an older arm with the same links as A: it realises 85 % of a command, answers 0.1 s late, and its servo has a gust of 0.12 rad/s. Of 256 attempted layouts, 197 yield a demonstration that completes the course; the other 59 touch a post and a laboratory discards them. Take the two differences apart. The slower, later plant alone gives demonstrations of 226 frames against 185 for A's own, inside A's allowance of 235 steps, and A completes 100 % of the runs that follow them. The extra gust alone gives 187 frames and 79 % of A's runs complete: A copies the wobble and the wobble eats the margin. With both, 19 % of the kept demonstrations are longer than 235 steps (mean 227), and A cannot finish when it replays those.
Now measure the effect on A. Let A follow one stored layout on test layouts at a layout distance below 1.2 cm, where it should succeed: with A's own layouts 99 % of 237 runs complete the course; with the older arm's layouts only 63 % do, 11 % end in a collision and 26 % run out of steps. The exchange rate follows: r = 0.32 at 64 attempted layouts on this draw and 0.25 averaged over twelve (between 0.15 and 0.33): an hour with the older arm buys about a quarter of an hour of A's own.
The lookup cannot see quality. When a layout of the older arm is nearer than one of A's by a hair, it wins, even though A's own would have completed the course. The remedy is a price on trust. Let a foreign layout count 1/w times as far away as an own one: w = 1 trusts it like your own and w = 0 never uses it, which is the tag of design (b). With 64 own layouts A completes the course on 79.0 % of fresh layouts. Adding 64 attempted layouts of the older arm at w = 1 gives 77.4 %: they cost 1.6 points. At w = 0.6, the best weight in a sweep from 0 to 1, success is 81.8 %, 4.4 points above full trust and 2.8 above ignoring the arm. With 16 own layouts the best weight is 1.0: when your own layouts are thin, even a weak foreign layout is worth following. The weight is set on held-out layouts, and it moves with how much data you already have.
The widget
What to try. Leave the defaults: 16 own layouts, the arm with links 0.55 + 0.45, 64 of its layouts, hand labels, w = 1. A completes the course on 83.2 % of fresh layouts (80.8 to 85.4 %) against 34.0 % alone, a gain of 49.2 points at an exchange rate of 0.94. Switch to joint angles at 64: 17.9 % against 34.2 % alone, 80 % of the runs follow a foreign layout, 82 % end in a collision, the gap reads 7.6 cm against a margin of 1.32 cm, and the rate is −0.24; at 256 layouts it falls to 9.1 %, near the 16/272 = 5.9 % of runs that follow A's own layouts. Turn the weight to 0, the tag of design (b): back to 34.2 %. Links 0.51 + 0.49 in joint angles still fail, 21.9 % with a gap of 1.5 cm; the same make works, 82.0 % against 83.1 % with hand labels. Choose the older arm with hand labels: 58.8 % and a rate of 0.32. With A's own layouts at 64 and 64 older-arm layouts: w = 1 gives 77.4 % against 79.0 % alone, w = 0.6 gives 81.8 %.
6 · What pooling cannot buy
Pooling turns other laboratories' hours into yours at close to one for one when their arms move like yours. It does not change how fast the need grows. If each further number per layout multiplies the layouts needed for ninety percent by the same factor, 7.0 as measured between one number and two, the need for scenes of d numbers is:
| numbers per layout | layouts for 90 % | one operator at Bench prices | share of Open X-Embodiment's million trajectories |
|---|---|---|---|
| 2 | 107 | 4.4 h | 0.011 % |
| 4 | 5,233 | 217 h (5.4 working weeks) | 0.5 % |
| 6 | 256,427 | 10,635 h | 26 % |
| 7 | 1,794,988 | 74,442 h | 179 % |
Take the most generous reading, one different layout per trajectory and every trajectory worth a layout of your own (r = 1, which §5 shows is generous). All of Open X-Embodiment (2023), over a million trajectories, then covers scenes of 6 numbers and not seven; at the factor of 5.8 per number that lesson 7 derives from footprints it covers 7 and not eight. Five objects on a table are already ten numbers. In hours the pool is smaller than it sounds: over 10,000 hours of π0's own data (2024), 2,976 hours in the 1,001,552 trajectories of AgiBot World (2025), about 2,000 hours in Open X-Embodiment by the AgiBot authors' count. Add them for a rough total of 14,976 hours: arithmetic on three published figures, not a census. As a yardstick it is the footage of 41 people each recording one hour of their day for a year. This lesson measures no video, but 41 people are a handful beside everyone who films what they do. And a film has no action channel: it shows where a hand went, not the command that sent it there or the force behind it.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
Written in joint angles, another arm's demonstrations take A from 34.2 % to 17.9 % at 64 layouts. Written as the place the hand should go, with A's own inverse kinematics and tracker, they take it to 83.2 %, and a foreign layout is worth 1.06 of A's own on average. That is still not enough for a real scene: every further number per layout multiplies the need by about 7.0, six numbers need 256,427 layouts, and the robot collections named here add up to roughly 14,976 hours, the footage of 41 people filming one hour a day for a year. People film themselves doing things all the time; that film shows where a hand went, and has no command and no force in it. What can be learned from watching, and what is missing from the video that no amount of it can supply?
Interview prompts
- Why can two arms of almost the same size not share joint-angle labels? (§2 — a joint angle is in the body's own units; a link difference δ moves the hand by 2|δ| sin(q2/2), about 1.5 δ here, and the margin allows 1.7 % of a link.)
- With actions written as the hand's place, what does the file store and what does each body supply? (§4 — the hand's place and the layout; each body supplies inverse kinematics for its own links and a stiff tracker for its own delay.)
- Why is raw joint-space pooling worse than leaving the data out? (§3 — the nearest layout is a foreign one with probability M/(N + M), its labels fail, and success falls to about the pooled success times N/(N + M): 17.9 % against 34.2 %.)
- How do you measure what another laboratory's hour is worth? (§4 — read off your own curve the own layouts that give the pooled success; the extra per foreign layout is the exchange rate, 0.94 here.)
- A foreign dataset in the right space is still worth a quarter of yours: why, and what do you do? (§5 — its targets carry its tempo and its servo's wobble (only 63 % of the runs that follow its layouts complete the course); weight it, 0.6 with 64 own layouts and 1.0 with 16.)
- Why does pooling every laboratory's data not solve the layout problem of a real scene? (§6 — the need multiplies by about 7.0 per number; six numbers need 256,427 layouts, a quarter of Open X-Embodiment.)
Companion reads: Reinforcement Learning · 56 Robot control (kinematics and the Jacobian: the per-body part), Computer Vision 3D · 03 Rigid motion (frames and offsets: why a shared space is only as shared as its calibration) and World Models · 27 Data mixture (weights on sources, as in §5).