all_lessons/Robot Model Training/18 · Diminishing returnslesson 18 / 24

The thousandth hour: diminishing returns

Lesson 17 priced every source per useful hour, and that is the price of its first hours. The hours after them teach the policy less, and this lesson measures how much less. On the Bench the failures a pool leaves shrink by about the same share with every layout added, and the exchange rate of a poorer source falls as a power of the hours bought, with exponents from 0.11 to 0.40. A budget is spent well when the next dollar buys the same success from every source bought. The policy then scores that rule, and hours do not add: the same 256 simulated layouts raise a pool of 8 own layouts by 60.6 points and lower a pool that holds 256 twin layouts by 4.5.

The thesis, here
The value of one more hour of a source is the slope of a curve, and the curves bend: a pool covers the layouts geometrically, and the exchange rate of a poorer source falls as a power of the hours bought. A budget is spent well when every source bought has the same marginal value per dollar. The rule is exact for hours that add, and they do not add: an hour is worth what it adds to the pool it joins, so the best split of a budget is not monotone in the budget and every rule of thumb loses to it.
Linear position
Forced by: Prices and exchange rates together give a cost per useful hour for every source, the price of an hour divided by its rate, and the ranking they produce is not the ranking by price: footage, dearer to make than the other borrowed hours, becomes the dearest useful hour of any source that can teach the same task, a simulator is the cheapest useful hour only while its gap is small and sells no useful hour at any price once the gap is wide, and some columns cannot be bought from the sources that are cheap. But a cost per hour is the cost of the first hour. The thousandth hour of a source teaches the policy less than the first. How fast does the value of one more hour fall?
New idea: the value of one more hour is the slope of a curve that falls as hours are bought, and a budget is spent well when the marginal value per dollar is the same in every source bought; on the Bench a pool saturates geometrically, a poorer source's rate falls as a power of its hours, and an hour's value depends on the pool it joins.
Forces next: The value of every source falls as a power of the hours bought, so the rule is to buy from each source until its marginal value per dollar equals that of the next, and which source is cheapest changes as you buy. This treats the hours as an asset that keeps its value once bought. Demonstrations do, for the policy trained next year as well as for the one trained today. Corrections were made in the states an older policy visited, and the policy trained next year will not visit quite the same ones. How long does a correction keep its value?
The plan
Six moves. (1) Compare a source's first layouts with its last. (2) Fit how the value falls: two laws, and where each fit is poor. (3) Derive the rule, equal marginal value per dollar. (4) Let the policy score pools of two sources. (5) Test the rules of thumb at equal budgets. (6) Ask what the rule assumed about an hour once it is bought.

1 · The first layouts and the last

A layout is one demonstration on the Bench, 9.3 s of motion (lesson 16), and every success below is that of the policy that copies the stored layout nearest to a new one (lesson 7), on 1000 new layouts. Eight layouts of your own take it to 21.6 %. Add layouts of the twin arm, an arm of the same make labelled in task space that costs $15 an hour, $0.039 a layout (lesson 17), and success reaches 98.8 % at 256. The price of an hour never changes. What an hour returns does:

twin layouts boughtcost of the blocksuccess gainedpoints per dollar
0 to 4$0.15511.1 points71.6
64 to 128$2.4811.6 points4.7
128 to 256$4.965.2 points1.0

The first 4 layouts return 68 times the points per dollar of the last 128, from a source that charges the same for every hour. Lesson 17's cost per useful hour is a price over an average rate, so it is silent on this. A poorer source shows the fall in the rate itself: an hour of the older arm replaces 0.76 of an own hour after 16 layouts and 0.17 after 256, so its useful hour goes from $19.7 to $88.7, 4.5 times dearer. How fast does the value of one more hour fall, and in what shape?

2 · Two laws, and where each fit is poor

The pool covers the layouts geometrically. By lesson 7 the policy fails by the distance to the nearest stored layout, which falls as N−1/d for d numbers describing a layout, but a layout is lost only when that distance exceeds the margin, and N layouts serve a share 1 − e−N/Nc of the box: the failures that remain shrink by the same factor with every layout added. The ledger's own curve shows it. From 16 layouts to 64 each layout removes 2.4 % of the failures still left; from 8 to 16, 2.1 %; from 128 to 256, 1.5 %, where the curve bends away from the exponential with 1.0 % of the layouts still lost.

A power law with a floor, fitted with the ledger's fitPower to the failures of each pool (8 own layouts and the source named, from 4 attempted layouts up; the own curve from 8), looks acceptable every time, with the best floor at 0, though the exponent changes tenfold from one curve to the next:

poolexponent αR²
own layouts alone1.510.93
twin1.090.90
simulator, 1 % gap0.400.94
older arm0.270.99
footage0.150.94

Take the own curve. Its failures have a local exponent, −d ln(failures) / d ln(layouts), of 0.25 between 8 and 16 layouts and 2.74 between 128 and 256, 11 times larger, and a power law holds it fixed. Over one decade a power law with a floor imitates a geometric curve, and a good R² does not say which it is.

A poorer source's exchange rate falls as a power of its hours. The rate (lesson 16) is the number of own layouts that a bought layout replaces. On log axes it falls close to a straight line from 8 to 256 layouts: ρ(h) = k h−γ, so the own-equivalent hours bought, E = ρh = k hβ, grow as the power β = 1 − γ of the hours. Fits to the ledger's table, 8 own layouts as the base:

sourcerate at 16 layoutsat 256exponent γkR² of the line
twin arm1.280.930.111.680.94
simulator, 0.3 % gap1.050.670.161.610.91
simulator, 1 % gap0.710.250.311.510.97
older arm0.760.170.401.860.81
footage (too noisy)0.220.100.140.260.19

The twin's rate hardly falls and the older arm's falls as the 0.40 power. A source whose pool saturates below the own curve cannot keep its rate, and the poorer sources saturate lower: 256 layouts of the older arm lift 8 own layouts to 71.7 %, of the 1 % simulator to 82.3 %, of footage to 56.0 %, of the twin to 98.8 %. So the power describes the fall between 8 and 256 layouts and cannot hold beyond. Footage's rates never exceed 0.27 and its line explains 0.19 of their variance: no exponent to use. Each fit has six points, and these fits are the planner's model of §3, not the Bench's truth.

Published fits are the same kind of object. Lin et al. (2024) fit the optimality gap, one minus the normalised score, as a power law in the number of training environments, objects or pairs (six points a line), with exponents from −0.47 to −0.84. They call a floor as a third parameter not statistically justified at six points, in three of six lines the fitted gap at one environment exceeds the largest gap there can be, and they find no clear power law in the number of demonstrations, only a plateau near 800. How much carries to a real task? The Bench's whole own curve, 256 layouts, is 0.66 hours of motion and a layout is described by two numbers; lesson 7 found that 90 % takes 14 layouts for one number, 104 for two, and about 5.8 times as many for each further number, while π0's post-training sets run from 5 hours to over 100. Exponents, saturation points and totals belong to this task and policy; the shapes, and the rule built on them, are what carry.

3 · Equal marginal value per dollar

Let V be the policy's success after buying hs hours from each source s at ps dollars an hour, with a budget B = Σ pshs. The value of the next hour of s, per dollar, is λs = (∂V/∂hs) / ps.

  1. Take a plan with λa > λb and move one dollar from b to a: success changes by λa − λb > 0, so the plan was not the best. At the best plan every source bought has the same λ, and one not bought has a λ at zero hours no higher.
  2. With power-law rates the hours of s are worth Es = kshsβs own hours, and V = F(n0 + ΣEs), with F the own curve of lesson 16 and n0 the own layouts already held.
  3. So ∂V/∂hs = F′ ksβshsβs−1 with the same F′ for every source, and equal λ is an equal marginal cost of an own-equivalent hour: cs(h) = ps / (ksβshβs−1) = μ.
  4. Solving, hs = (μ ksβs / ps)1/(1−βs), with μ the value that makes Σ pshs = B.

Price enters through the exponent 1/(1 − β). At a fixed μ, halving a source's price multiplies the hours bought by 21/(1−β): 5.7 for β = 0.6, the older arm's, and 545 for β = 0.89, the twin's. A source whose rate hardly falls is a corner, bought whole or not at all. Your own robot has β = 1 and a flat marginal cost, its price, so $123.6 is the ceiling on μ: no source is bought past the point where its next useful hour costs what your own does. With the fits of §2 and lesson 17's prices (assumptions, not measurements), the next useful hour costs, after 8, 64 and 256 layouts:

sourceprice per hourafter 8 layoutsafter 64after 256
simulator, 1 % gap$0.19$0.34$0.66$1.01
twin arm$15$12.6$15.9$18.5
older arm$15$30.8$70.6$122.7
own robot$123.6 (kept hour)$123.6$123.6$123.6

Lesson 17's order at the first hours, simulator, twin, older arm, own, footage, holds, with one change. The older arm's next hour costs 4.0 times less than your own at 8 layouts and about the same at 256 (0.99 of it); the fit crosses at 261 layouts, and past that the older arm costs more than recording your own. The simulator's next hour triples and still costs 18 times less than the twin's. Footage, at $33.75 an hour, has a best rate at any point of 0.27, which never brings its useful hour below $125, about what your own costs. As a plan: the simulator first, the twin next, the older arm until it costs what your own does, exact if hours add. §4 asks the policy whether they do.

4 · A pool is not a sum

§3 gave every source a rate in own layouts and read one curve at the total. Test that on the policy. Each source supplies its own list of 256 layouts, so a pool never buys the same layout twice; the lookup takes the nearest stored layout of any source; and every number is the success on the same 1000 new layouts. Pools built on 8 own layouts, with a simulator whose arm model is off by 1 % and whose labels are joint angles (lesson 8):

poolsuccess
8 own layouts21.6 %
+ 256 twin layouts98.7 %
+ 256 simulated layouts82.2 %
+ 256 twin and 256 simulated layouts94.2 %

The same 256 simulated layouts add +60.6 points to the pool of 8 own layouts and −4.5 to the pool that holds the twin's. (The ledger's table reads the twin's 256 layouts at 98.8 % and the simulator's at 82.3 %, these pools 98.7 and 82.2: the table gives every foreign source one list of layouts, a pool a list of its own. These are one draw of the layouts, the one nearest the medians of eight; on all eight the simulator lowers the twin's pool, by 2.8 to 5.8 points.)

The mechanism is the lookup, which cannot see quality (lesson 8). In the pool with both, the nearest stored layout belongs to the simulator for 49.1 % of the test layouts, to the twin for 48.9 % and to you for 2.0 %. Following a simulated demonstration completes the course 88.2 % of the time, a twin's 100.0 %. Success is the sum over sources of share × reliability, and a source of lower reliability that wins half the shares lowers it. Sources enter a pool as competitors for the nearest slot, not as additions. The sum model of §3 misses this: it reads the own curve at 8 + ρtwinn + ρsimm with the ledger's rates, predicts 93.6 % for the best split of $3, and the policy measures 89.0 %, 4.6 points lower.

The widget

Spend a budget on two sources, and score the pool
Left: success on 1000 new layouts against the budget (log axis) for the best split (teal) and four rules of thumb: all on the cheaper hour (red), half each (grey), each source to 90 % (cyan, dotted), lesson 17's ranking (purple, dashed); amber is the budget. Right: success along the budget line as the other source takes more of it, measured (teal) and as the sum of the sources' rates predicts (dashed). Every pool is scored by running the policy (draw 7 of 8). The second slider moves the twin's price from a quarter of $15 to your own hour's.
spent
-
layouts of the other source
-
twin layouts
-
best split reaches
-
sum of rates predicts
-
all on the cheaper hour
-
half each
-
to 90 % each
-
lesson 17's ranking
-
other: nearest for
-
other: completes
-
twin: nearest for
-
twin: completes
-
next twin layout, per useful hour
-
next other layout, per useful hour
-
Show the core JS
RT.score = function (sp) {
...
  for (k = 0; k < NT; k++) {
    var bestD = Infinity, src = -1, idx = 0;
    for (s = 0; s < 3; s++) {
      if (use[s] === 0) continue;
      var d = Ls[s].bd[k * Ls[s].nd + use[s] - 1];
      if (d < bestD) { bestD = d; src = s; idx = Ls[s].bi[k * Ls[s].nd + use[s] - 1]; }
    }
    var r = RT.outcome(Ls[src].demos[idx], Ls[src].space, k);
    pick[src]++; if (r === 1) { ok++; good[src]++; }
  }
...
RT.split = function (B, sp0, cT, cO, m) {
  var spend = cO * m; if (spend > B + 1e-9) return null;
  var n = Math.min(RT.CAP, Math.floor((B - spend) / cT + 1e-9)), r = RT.score({ na: sp0.na, n: n, m: m, other: sp0.other });
  return { m: m, n: n, succ: r.succ, spent: spend + cT * n, r: r };
};
...
  for (m = 0; m <= top; m++) { c = RT.split(B, sp0, cT, cO, m); if (c && (!best || c.succ > best.succ)) best = c; }
...
RT.marginal = function (f, cost, h) { return cost / (f.k * f.beta * Math.pow(Math.min(RT.CAP, Math.max(8, h)), f.beta - 1)); };   // dollars of the next useful layout after h attempted ones (the fit holds from 8 to 256)

What to try. Leave the defaults: $3.00, the twin at $15 an hour, a simulator off by 1 %. The best split spends $2.98 on 159 simulated and 75 twin layouts and reaches 89.0 %, where the sum of rates predicts 93.6 %. Simulated layouts are the nearest for 62.3 % of the test layouts and complete 83.1 % of them, the twin's 33.7 % and 98.8 %. The four rules reach 82.2, 86.3, 88.4 and 88.4 %. At $6.00 the best split buys no simulator and 154 twin layouts, 96.5 %, and the ranking reaches 92.1 %; at $10.00 the best is 98.7 %. Choose the older arm and put the twin at your own hour's price: the best split at $3.00, $6.00 and $10.00 is 59 older-arm and 2 twin layouts, then 145 and 1, then 46 and 25, not monotone either. Choose the 3 % simulator with joint labels, the source that sells nothing (lesson 17): all the money on it reaches 5.1 % at $6.00, below the 21.6 % of your 8 layouts alone. With hand labels the sum predicts 99.0 % at $3.00, the policy measures 99.5 %, and the ranking is within 0.1 points of the best split: its demonstrations complete 99.5 % of the layouts they serve, about as many as the twin's.

5 · The rules of thumb, scored by the policy

Four rules for spending a budget, each scored by the policy, against the best split found by trying every number of simulated layouts the budget allows. All on the simulator spends everything on the cheaper hour; half each gives each source half the money; to 90 % each buys each source until it alone would reach 90 %; lesson 17's ranking buys the source with the lower cost per useful hour to the end of its catalogue, then the other.

budgetbest split: simulated + twin layoutsall on the simulatorhalf eachto 90 % eachlesson 17's ranking
$1245 + 22: 85.3 %82.2 %84.3 %85.2 %85.2 %
$3159 + 75: 89.0 %82.2 %86.3 %88.4 %88.4 %
$60 + 154: 96.5 %82.2 %88.4 %89.4 %92.1 %
$100 + 256: 98.7 %82.2 %90.7 %89.4 %94.1 %

At $1 the ranking and the 90 % rule are within 0.1 point of the best split, because they spend nearly everything on the simulator and so does the best split. Each fails differently later. All on the simulator stops at 82.2 % from $0.12, where its 256 layouts run out. Half each spends $5.12 of $10: the simulator's 256 layouts cost $0.12 and the rest of its half is left over. To 90 % each buys the simulator's 256 layouts, which alone reach 82.3 %, then the 92 twin layouts that alone reach 90 %; together they reach 89.4 %, and it stops at $3.69. Lesson 17's ranking is the best of the four and still 4.4 points behind at $6 and 4.6 at $10.

The best split is not monotone in the budget. It buys the simulator first, 245 layouts at $1 and 159 at $3, and none at $6 or $10; over eight draws of the layouts it falls below 32 simulated layouts at a budget between $3 and $5, and at $10 it buys at most 1. A block of 64 simulated layouts costs $0.031. On 8 own layouts it adds 38.1 points, 1,229 per dollar; in the pool that holds the twin's 256 layouts it adds −3.1 points, −100 per dollar. The twin's first 64 layouts cost $2.48 and add 61.6 points, 24.8 per dollar. At the start the simulator's block is worth 49 times the twin's per dollar, and after the twin it is worth less than nothing. That is how the cheapest source changes as you buy: the marginal value per dollar of each source depends on what the others have already put in the pool, and the λ's cross.

Road not taken · buy down lesson 17's ranking
Lesson 17 ends with a ranking by cost per useful hour, and the obvious way to spend with it is to buy the first source to the end of its catalogue, then the next. That is §3's rule for hours that add, and for a thin budget it is right. At $10 it loses 4.5 to 5.8 points to the best split over eight draws, because the 256 simulated layouts it bought first sit in the pool when the twin's arrive. Use the ranking to choose the first purchase and the policy to choose the last.
What this lesson did not do
The policy trusts every stored layout alike, which is w = 1 in the price on trust of lesson 8, so the dilution of §4 is that of an uncorrected lookup; a lower weight on the simulator would be expected to shrink it, and was not run. A source whose demonstrations are as reliable as the twin's, the 3 % simulator with hand labels, adds almost as the sum model says. The pools have 8 own layouts and one draw of the layouts (draw 7 of eight; ranges over the eight are quoted), and each catalogue stops at 256 attempted layouts, 0.66 hours. The prices are lesson 17's assumptions, and the best split moves with them (the second slider) while the pools do not. The task is one capability with no force or recovery column (lesson 17, §6) and is small: lesson 7's factor of about 5.8 per extra number is what a larger task does to the hours. Later lessons write it as a task size K: every number of hours multiplied by K, prices and rates unchanged.

6 · What the rule assumed

Every step of §3 read a curve of the hours bought, and bought is the word to check. The rule treats hours as an asset: once in the pool, their value is what the curve says, today and next year. §4 already broke one version of that: the same 256 simulated layouts are worth +60.6 points to a pool of 8 own layouts and −4.5 to a pool that holds the twin's, so an hour's value is not a property of the hour. That is about the pool at one time. The rule also assumes the value stays when the policy is retrained, and the two kinds of record the ledger buys behave differently there.

A demonstration is a record of what the expert did in a layout, the same record whichever policy reads it. The policy trained next year finds the same layouts and trajectories on file, and the coverage lesson 7 counts, the distance from a new layout to the nearest stored one, does not depend on which policy counts it. A demonstration keeps its value for next year's policy as for today's, so §3's rule applies to demonstrations as it stands.

A correction is a different record. It was made in a state that an older policy reached, and it labels that state with what the expert would do there (lesson 3). Its use depends on the policy that trains on it reaching that state, and the policy trained next year will not visit quite the same ones. How much that costs a correction is a measurement the rule does not contain, and until it is made the rule cannot be applied to a stock of corrections.

Common mistakes / failure modes

"buy the lowest cost per useful hour until it runs out, then the next"
At $10 that reaches 94.1 % where the best split reaches 98.7 %: simulated layouts are the nearest for 49.1 % of the layouts and complete 88.2 % of them (§4, §5).
"split the budget in half between the two cheapest sources"
At $10 half each spends $5.12 and reaches 90.7 %, 8.0 points short: the simulator's catalogue ended at $0.12 (§5).
"buy the cheapest hour"
The 3 % simulator with joint labels is the cheapest hour of the table; all the money on it reaches 5.1 %, below the 21.6 % of 8 own layouts alone (widget).
"the fit is good, so it is a power law"
R² 0.93 on the own curve, whose local exponent grows 11 times, from 0.25 to 2.74: the failures fall geometrically (§2).

Checkpoint exercise

Try it
Source A costs $12 an hour and hA hours of it are worth EA = 3√hA own hours; source B costs $3 an hour and is worth EB = 2√hB. The policy's success rises with EA + EB (hours add) and you have $48. (a) What ratio hA / hB makes the marginal value per dollar equal? (b) How many hours of each do you buy, and what is EA + EB? (c) What do all $48 on B, all on A, and $24 each give? Answer: (a) dE/dh = k / (2√h), so (3 / (2√hA)) / 12 = (2 / (2√hB)) / 3 gives √hA / √hB = 3/8 and hA / hB = 9/64 = 0.1406. (b) 12 · (9/64) hB + 3 hB = 48 gives hB = 10.24 and hA = 1.44: $17.28 on A and $30.72 on B, E = 3 · 1.2 + 2 · 3.2 = 10.0, with a marginal value per dollar of 0.1042 in both. (c) All on B is 16 hours, E = 8.0; all on A is 4 hours, E = 6.0; $24 each gives 9.90, close only because both exponents are ½. If A's hours competed with B's for the nearest slot, as in §4, this arithmetic would not say how to split the $48: the policy would have to score the pools.

Where this points next

A pool covers the layouts geometrically, each layout removing 2.4 % of the failures that remain, and the exchange rate of a poorer source falls as a power of its hours, with exponents from 0.11 for the twin to 0.40 for the older arm. A budget is spent well where the marginal value per dollar is equal across the sources bought, but hours do not add: at a budget of $3 to $5 the best split buys fewer than 32 simulated layouts, and lesson 17's ranking, bought down in order, loses 4.5 to 5.8 points at $10. All of it treats the hours as an asset that keeps its value once bought. Demonstrations do, for the policy trained next year as well as for the one trained today. Corrections were made in the states an older policy visited, and the policy trained next year will not visit quite the same ones. How long does a correction keep its value?

Takeaway
The first layouts of a source are worth far more than its last: the twin's first 4 return 68 times the points per dollar of its last 128. On the Bench a pool saturates geometrically and a poorer source's rate falls as a power of its hours (exponents 0.11 to 0.40, footage too noisy), six-point fits and not laws. A budget is spent well where the marginal value per dollar is equal across sources, so that every source's next useful hour costs the same, with your own hour's price as the ceiling. But an hour's value depends on the pool it joins: the same 256 simulated layouts add +60.6 points to 8 own layouts and −4.5 to a pool with the twin's, so the best split drops the simulator as the budget grows and every rule of thumb loses to it. The rule assumes an hour keeps its value once bought. That holds for demonstrations, and is the question for corrections.

Interview prompts

Companion reads: Lesson 7 · How much data: the coverage law, Lesson 8 · Other bodies and Lesson 3 · Labels on your own states.