How much data: the coverage law
Lesson 6 made the policy exact about the layout it was shown: it completes the course on that layout every time, and walks into the vases when the row has moved 2 cm. This lesson gives the policy the layout as an input. A policy that generalises only to what resembles its data is wrong on a new layout by the distance to the nearest layout it was shown, so success is a coverage problem. One demonstrated layout serves a footprint, N layouts serve a share 1 − exp(−N·A/V) of the box, and the error falls as N−1/d when d numbers describe a layout. A repeat of a layout adds nothing; a new layout adds a footprint. Ninety percent is priced by measurement for one number and two, and by derivation for more. Where more layouts come from is left open.
New idea: a policy that generalises only to what resembles its data serves a footprint of layouts around each layout it was shown, so success is coverage, and it is bought with different layouts and not with repeats of one. The count is the range of the layouts divided by the footprint, once for every number that describes a layout.
Forces next: For a policy that generalises only to what resembles its data, the error on a new layout is the distance to the nearest demonstrated layout, so success grows with the number of different layouts, with a power set by how many numbers describe a layout, and hardly at all with more demonstrations of a layout already seen. Ninety percent costs about a hundred layouts when two numbers describe a layout, and every further number multiplies that by the range it varies over divided by the width one layout covers; a real scene is described by many more than two numbers, and one arm cannot pay for that. Other laboratories have already collected such data, on other arms. What can a robot use from demonstrations made on a body that is not its own?
1 · The layout is an input, and a local learner copies the nearest one
Lesson 6 ended on a table of shifts, and this lesson starts from it, finer. One layout of the course demonstrated, the row of posts then shifted by s and the policy run 200 times, with the cup starting where the new row starts (lesson 6 left it where the demonstrations began, so a few percentages differ):
| shift s (cm) | −2 | −1.5 | −1 | 0 | +1 | +1.5 | +2 |
|---|---|---|---|---|---|---|---|
| reaches the mat (%) | 5.0 | 91.5 | 100.0 | 100.0 | 100.0 | 88.0 | 0.0 |
The policy serves the shifts from −1.75 to +1.85 cm around the layout it was shown: 3.6 cm of the 20 cm range we will vary the row over, 18.0 %. The path clears the posts by 1.55 cm (lesson 6), and the policy fails once the row has moved far enough that a post stands in the path. We call the set of layouts that a demonstrated layout serves, the ones where it succeeds at least half the time, its footprint.
A layout is the numbers that say where the posts are. Here there are two: the shift s of the whole row, up, in cm, and its tilt τ, the height gained per metre along the row, so that post i stands at 0.55 m + s + τ·(xi − xmid). The layouts fill a box: s within ±10 cm and τ within ±0.2, so that an end post stands up to 8 cm higher or lower. At the top corner the goal is 0.95 m from the base, close to the arm's 1.0 m reach, so the box cannot grow upward. With one number (d = 1) only the shift varies; with two (d = 2) the tilt varies as well. The expert reads the posts, so it can demonstrate any layout in the box, and it does, without a touch, on every one of them.
The data are N different layouts and m demonstrations of each, calm, from starts jittered by 1 cm (lessons 1 and 6). The test is the success on layouts the policy has never seen: a grid of 320 spread evenly over the box (100 for one number), each run once under the gust. A lab pays for both counts: a layout costs the rebuilding of the row and a demonstration costs a run of the expert.
What does the policy do with a layout it has not seen? Lesson 1's learner copies what the expert did in the most similar moment, and a moment from another layout is not similar, even at the same joint angles: the same pose of the arm is a different path among different posts. So the lookup first takes the demonstrated layout nearest to the new one, and then does what lesson 6 did inside it: the stored frame nearest to the arm, the 16 poses that followed it, tracked by the stiff controller. Nearest in what? In how far the posts moved. The distance between two layouts is how far the farthest post moves, and for a shift and a tilt that is
D(θ, θ′) = |s − s′| + 0.4 m · |τ − τ′| (the end posts are 0.4 m from the middle of the row)
with s in cm, 0.4 m being 40 cm per unit of tilt. This is the simplest rule that generalises only to what resembles its data.
2 · The error of a copy is a distance, and a layout serves a footprint
Run the policy that copies one layout on another, with no gust and with the posts taken away so that nothing stops the run, and compare the cup's path with the path the expert takes on the new layout. The stored path is the old path. The right path is the old path with every post's height changed by Δs + Δτ·(x − xmid): a rigid move, largest at an end post, where it is |Δs| + 0.4|Δτ|, which is D. So the error of a copy, the worst gap between the path it follows and the right one, is the distance. Measured, with demonstrations that start on the nominal start:
| new layout, relative to the copied one | distance D (cm) | worst gap measured (cm) |
|---|---|---|
| 1 cm of shift | 1.0 | 1.00 |
| 0.05 of tilt | 2.0 | 2.00 |
| 2 cm of shift and 0.05 of tilt | 4.0 | 3.99 |
| −6 cm of shift and 0.10 of tilt | 10.0 | 10.01 |
With the demonstrations the lessons use, whose starts are jittered by 1 cm, the gap and D differ by at most 0.55 cm in nine random pairs of ten, and the median ratio is 1.00 over 193 pairs.
Whether an error is a failure depends on the posts. The stored path passes above some posts and below others, so a move of the row up threatens only the posts it passes above, and a move down only the posts it passes below; the footprint is where the move stays under the clearance. For one number it is the interval above. For two it is a skewed diamond of area 0.24 (cm · tilt), 3.0 % of the box of 20 cm × 0.4 = 8.0: the widget draws it when N = 1. Its size carries the argument, and its shape does not.
3 · N layouts: the share served, and how the error falls
Draw N layouts at random in the box, whose size is V (20 cm for one number, 8.0 cm · tilt for two), and centre a footprint of size A on each (3.6 cm, 0.24). A given new layout lies inside one footprint with probability A/V, setting the box's edges aside, and inside none of N with probability (1 − A/V)N ≈ exp(−N·A/V). Success on a new layout is the share served:
S(N) = 1 − exp(−N / Nc), Nc = V / A, S = 90 % at N90 = ln 10 · Nc = 2.30 · V / A
The footprints give Nc = 5.6 for one number and 33.7 for two, so N90 = 12.8 and 77.6 layouts. The edges and corners of the box, which fewer footprints reach, raise both; §4 measures by how much.
The error follows from the same count. The distance δ from a new layout to the nearest of N exceeds r only if no layout lies in the ball of radius r, whose size is c·rd, so P(δ > r) ≈ exp(−N·c·rd/V), and the mean distance is Γ(1 + 1/d)·(V/cN)1/d. The error falls as N−1/d. The balls of our distance are an interval of 2r cm (c = 2) and a diamond of area r²/20 in cm · tilt (c = 1/20), so the mean distance is 10/N cm for one number and 11.2/√N cm for two. How many numbers describe a layout sets the exponent; the box sets the constant.
The widget
What to try. Leave the defaults: two numbers, 16 layouts, one demonstration each. The set on the left serves 36.6 % of the box; the mean over 64 sets is 34.7 %, with 3.3 points between sets, and the mean distance from a new layout to its nearest demonstrated one is 3.07 cm. Slide to one layout: the teal patch is the footprint, 3.0 % of the box, and the mean over sets reads 7.7 %, above that share, because a layout near one end of the box also serves cells at the other end: a row moved down by 20 cm lies wholly below the old path, which clears every post (100 % success, against 0 % at 10 cm). Layouts far apart in shift serve each other this way, so it matters at N = 1 and 2 and not later. Slide up: at 64 layouts the mean is 78.9 %, at 128 it is 93.6 % and at 256 it is 98.9 %, while the mean distance falls from 11.79 cm at one layout to 0.71 cm at 256, the fitted exponent reading −0.51 against the guide −0.5; the layouts needed for 90 % read 104. Switch to one number: 16 layouts now give 92.8 %, the exponent is −0.95 against −1, and 90 % takes 14 layouts. Back to two numbers and 16 layouts, raise the second slider to 8: the total goes from 16 to 128 demonstrations and the set on the left goes from 36.6 % to 36.3 %. Put the same 128 demonstrations into 128 layouts, one each: 93.6 % against 34.7 % for 16 layouts of 8, at 5.3 operator hours against 1.6.
4 · What the Bench measures
The same experiment over 64 random sets of layouts, success as the mean over the sets, distance in cm against the formula of §3:
| N | one number (shift) | two numbers (shift, tilt) | ||||
|---|---|---|---|---|---|---|
| success (%) | distance, formula | measured | success (%) | distance, formula | measured | |
| 4 | 50.3 | 2.50 | 2.40 | 10.9 | 5.6 | 6.3 |
| 16 | 92.8 | 0.63 | 0.63 | 34.7 | 2.8 | 3.1 |
| 64 | 100.0 | 0.16 | 0.15 | 78.9 | 1.4 | 1.5 |
| 256 | 100.0 | 0.04 | 0.04 | 98.9 | 0.70 | 0.71 |
The distance is the formula to within ten percent from 16 layouts up (the box's edges raise it a little at small N), and the fitted exponent over N = 1 to 256 is −0.95 and −0.51, against −1 and −1/2. The success is not a power law. It is the share served, and ninety percent is reached at N90 = 14 layouts for one number and 104 for two (groups of 16 sets give 13 to 16 and 102 to 108). The formula of §3 said 12.8 and 77.6; counting the box's edges, the footprint predicts 13 and 85, within 3.7 and 3.9 points of the measured curve everywhere from N = 3 up. The measured curve runs below the prediction, in two steps of about two points. At N = 104, 94.0 % of the cells lie in some footprint, but the policy copies the nearest layout, which is not always the one whose footprint holds the cell, and only 92.2 % lie in the footprint of the nearest. Then 90.2 % are served, because no two footprints are quite alike (at the four corners of the box they measure 94 to 112 % of the central one, each resting on one demonstration) and their edges are soft (§1's table, at ±1.5 cm).
Ninety percent is a mean over sets. A set of 104 layouts has a spread of 2.1 points, and 39 of the 64 sets reach 90 %. The last points are the dear ones: 95 % takes 144 layouts for two numbers, and 256 layouts reach 98.9 %. Real policies also improve as a power of the number of environments: Lin et al. (2024) fitted the gap in a graded score against the number of training environments, objects or pairs, from 1 to 32, with exponents from −0.47 to −0.84 on six fits of six points each. Theirs is a score and ours a distance, so compare the shape and not the exponents.
5 · Which data buy it
Everything in §3 counted layouts. A second demonstration of a layout puts a second set of frames inside a footprint that is already there: the nearest demonstrated layout is the same one, and it serves the same cells. The widget's second slider tests this, and the table gives the answer on 64 sets, two numbers, N = 64:
| demonstrations per layout | 1 | 2 | 3 | 5 | 8 |
|---|---|---|---|---|---|
| success (%) | 78.9 | 78.9 | 78.8 | 78.8 | 78.8 |
The spread is 0.09 points, and with one number at N = 16 it is 0.11. The tracker follows the stored path, and a calm expert repeats itself to within the 1 cm jitter of its start. A person does not repeat themselves (lesson 4), so on real data a repeat averages scatter, which is worth something up to a point: Lin et al. (2024) find no clear power law in the number of demonstrations once the environments are fixed, and a plateau near 50 demonstrations per environment–object pair.
Now spend the same 128 demonstrations in different ways, two numbers:
| layouts × demonstrations each | 1 × 128 | 16 × 8 | 32 × 4 | 64 × 2 | 128 × 1 |
|---|---|---|---|---|---|
| success (%) | 7.7 | 34.7 | 56.3 | 78.9 | 93.6 |
| operator hours (Bench prices, §6) | 1.1 | 1.6 | 2.1 | 3.2 | 5.3 |
Layouts cost more time than repeats, so compare at equal time as well: the 5.3 hours of 128 single demonstrations buy 53 layouts of 8 (5.2 hours), which serve 72.6 % against 93.6 %. Real programmes spend the same way. DROID (Khazatsky et al., 2024) moves on after up to 100 trajectories in a scene, across 564 scenes, and in its comparison of 7,362 trajectories from the 20 scenes with most data against 7,362 drawn across all scenes, the spread-out draw did better out of distribution. At equal totals of about 120 demonstrations, Lin et al. (2024) score 0.05 for one environment–object pair and 0.44 for 32 pairs on pouring water.
6 · What ninety percent costs
Put prices on a layout. A demonstration takes the expert 9.3 s (lesson 1). We add 20 s to put the cup back and 120 s to rebuild the row; these two are assumptions, chosen to be kind to the lab. A layout then costs 149.3 s with one demonstration and 354.4 s with eight. One number takes 14 layouts for 90 %: 35 minutes. Two numbers take 104: 4.3 hours at one demonstration each, 6.0 at three and 10.2 at eight. A real programme is slower. RT-1 (Brohan et al., 2022) collected about 130k demonstrations with 13 robots in 17 months, which is 588 per robot-month; at that pace 104 demonstrations are 0.8 weeks of one robot. A day is not what stops one arm. The growth is.
The second number multiplied the layouts by 7.4. Along the shift the box is 20 cm and one layout serves 3.6 cm, a ratio of 5.6, which is Nc for one number. Along the tilt the box is 0.40 and the footprint's width, its area 0.24 divided by its 3.6 cm of shift, is 0.066, a ratio of 6.1. The product is 33.7, the Nc of two numbers, and N90 is at least 2.30 times it, the edges adding more (×1.09 for one number, ×1.34 for two). So Nc is the product, over the numbers that describe a layout, of the range divided by the width one layout covers, the width of a number being the footprint's size with that number free divided by its size with the number held. For two numbers that is measured. For more it is derived, and what the Bench cannot tell is the width of the numbers of a real scene. A number the task ignores has a width equal to its range and costs a factor of 1; a number it depends on costs, here, about 5.8, the geometric mean of the two ratios. Taking that factor for every further number, N90 ≥ 2.30 · 5.8k for k numbers:
| numbers k | layouts for 90 % (at least) | operator hours, one demonstration each | 40-hour weeks | robot-months at RT-1's pace |
|---|---|---|---|---|
| 3 | 449 | 18.6 | 0.5 | 0.8 |
| 4 | 2,606 | 108.1 | 2.7 | 4.4 |
| 5 | 15,113 | 626.8 | 15.7 | 25.7 |
| 6 | 87,656 | 3,635 | 90.9 | 149 |
A real scene has many more than two numbers: a single object lying on a table already has three, where it is and which way it faces. At this Bench's factor, three numbers ask for 449 layouts, a little under the 564 scenes of DROID, and six ask for 87,656, 0.67 of the 130k demonstrations that 13 robots collected for RT-1 in 17 months. The factor 5.8 belongs to this Bench; the form does not change, and it is why one arm cannot pay for a real scene.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
The error on a new layout is the distance to the nearest layout it was shown, and success is the share of the box that some footprint covers: 78.9 % at 64 layouts, and 104 layouts for 90 % when two numbers describe a layout, 14 when one does. Repeats of a layout buy nothing (0.09 points from 1 to 8). Two numbers cost 4.3 hours of an operator at the Bench's prices, under a working day; every further number multiplies the count by its range over its width, about 5.8 here, so five numbers cost 15,113 layouts and 15.7 working weeks, and a real scene has more than five. One arm cannot pay for that. Other laboratories have already collected such data, on other arms. What can a robot use from demonstrations made on a body that is not its own?
Interview prompts
- Why is the error of a policy that copies the nearest demonstrated layout equal to the distance to it? (§2 — the right path is the stored path moved rigidly by how far the posts moved, largest at an end post.)
- What shape does success take as the number of layouts grows, and why? (§3 — 1 − exp(−N·A/V): each layout serves a footprint of size A in a box of size V, so a new layout is missed by all N with probability (1 − A/V)N.)
- What sets the exponent of the error against the number of layouts? (§3 — the ball of radius r holds c·rd of the box, so the nearest of N is at a distance of order N−1/d; measured −0.95 and −0.51 for d = 1 and 2.)
- You have a budget of 128 demonstrations: 16 layouts of 8 or 128 of 1? (§5 — 128 of 1: 93.6 % against 34.7 %, because repeats land in a footprint already covered.)
- How would you estimate the layouts a task needs before collecting them? (§3, §6 — measure the footprint of one layout, divide the box by it, multiply by 2.3 for ninety percent, and add the edges.)
- What does one more variable of the scene cost in data? (§6 — a factor of its range over the width one layout covers; 1 if the task ignores it, about 5.8 for the Bench's numbers.)
Companion reads: Reinforcement Learning · 19 Offline RL (a fixed dataset supports only what it covers).