The thousandth hour: diminishing returns
Lesson 17 priced every source per useful hour, and that is the price of its first hours. The hours after them teach the policy less, and this lesson measures how much less. On the Bench the failures a pool leaves shrink by about the same share with every layout added, and the exchange rate of a poorer source falls as a power of the hours bought, with exponents from 0.11 to 0.40. A budget is spent well when the next dollar buys the same success from every source bought. The policy then scores that rule, and hours do not add: the same 256 simulated layouts raise a pool of 8 own layouts by 60.6 points and lower a pool that holds 256 twin layouts by 4.5.
New idea: the value of one more hour is the slope of a curve that falls as hours are bought, and a budget is spent well when the marginal value per dollar is the same in every source bought; on the Bench a pool saturates geometrically, a poorer source's rate falls as a power of its hours, and an hour's value depends on the pool it joins.
Forces next: The value of every source falls as a power of the hours bought, so the rule is to buy from each source until its marginal value per dollar equals that of the next, and which source is cheapest changes as you buy. This treats the hours as an asset that keeps its value once bought. Demonstrations do, for the policy trained next year as well as for the one trained today. Corrections were made in the states an older policy visited, and the policy trained next year will not visit quite the same ones. How long does a correction keep its value?
1 · The first layouts and the last
A layout is one demonstration on the Bench, 9.3 s of motion (lesson 16), and every success below is that of the policy that copies the stored layout nearest to a new one (lesson 7), on 1000 new layouts. Eight layouts of your own take it to 21.6 %. Add layouts of the twin arm, an arm of the same make labelled in task space that costs $15 an hour, $0.039 a layout (lesson 17), and success reaches 98.8 % at 256. The price of an hour never changes. What an hour returns does:
| twin layouts bought | cost of the block | success gained | points per dollar |
|---|---|---|---|
| 0 to 4 | $0.155 | 11.1 points | 71.6 |
| 64 to 128 | $2.48 | 11.6 points | 4.7 |
| 128 to 256 | $4.96 | 5.2 points | 1.0 |
The first 4 layouts return 68 times the points per dollar of the last 128, from a source that charges the same for every hour. Lesson 17's cost per useful hour is a price over an average rate, so it is silent on this. A poorer source shows the fall in the rate itself: an hour of the older arm replaces 0.76 of an own hour after 16 layouts and 0.17 after 256, so its useful hour goes from $19.7 to $88.7, 4.5 times dearer. How fast does the value of one more hour fall, and in what shape?
2 · Two laws, and where each fit is poor
The pool covers the layouts geometrically. By lesson 7 the policy fails by the distance to the nearest stored layout, which falls as N−1/d for d numbers describing a layout, but a layout is lost only when that distance exceeds the margin, and N layouts serve a share 1 − e−N/Nc of the box: the failures that remain shrink by the same factor with every layout added. The ledger's own curve shows it. From 16 layouts to 64 each layout removes 2.4 % of the failures still left; from 8 to 16, 2.1 %; from 128 to 256, 1.5 %, where the curve bends away from the exponential with 1.0 % of the layouts still lost.
A power law with a floor, fitted with the ledger's fitPower to the failures of each pool (8 own layouts and the source named, from 4 attempted layouts up; the own curve from 8), looks acceptable every time, with the best floor at 0, though the exponent changes tenfold from one curve to the next:
| pool | exponent α | R² |
|---|---|---|
| own layouts alone | 1.51 | 0.93 |
| twin | 1.09 | 0.90 |
| simulator, 1 % gap | 0.40 | 0.94 |
| older arm | 0.27 | 0.99 |
| footage | 0.15 | 0.94 |
Take the own curve. Its failures have a local exponent, −d ln(failures) / d ln(layouts), of 0.25 between 8 and 16 layouts and 2.74 between 128 and 256, 11 times larger, and a power law holds it fixed. Over one decade a power law with a floor imitates a geometric curve, and a good R² does not say which it is.
A poorer source's exchange rate falls as a power of its hours. The rate (lesson 16) is the number of own layouts that a bought layout replaces. On log axes it falls close to a straight line from 8 to 256 layouts: ρ(h) = k h−γ, so the own-equivalent hours bought, E = ρh = k hβ, grow as the power β = 1 − γ of the hours. Fits to the ledger's table, 8 own layouts as the base:
| source | rate at 16 layouts | at 256 | exponent γ | k | R² of the line |
|---|---|---|---|---|---|
| twin arm | 1.28 | 0.93 | 0.11 | 1.68 | 0.94 |
| simulator, 0.3 % gap | 1.05 | 0.67 | 0.16 | 1.61 | 0.91 |
| simulator, 1 % gap | 0.71 | 0.25 | 0.31 | 1.51 | 0.97 |
| older arm | 0.76 | 0.17 | 0.40 | 1.86 | 0.81 |
| footage (too noisy) | 0.22 | 0.10 | 0.14 | 0.26 | 0.19 |
The twin's rate hardly falls and the older arm's falls as the 0.40 power. A source whose pool saturates below the own curve cannot keep its rate, and the poorer sources saturate lower: 256 layouts of the older arm lift 8 own layouts to 71.7 %, of the 1 % simulator to 82.3 %, of footage to 56.0 %, of the twin to 98.8 %. So the power describes the fall between 8 and 256 layouts and cannot hold beyond. Footage's rates never exceed 0.27 and its line explains 0.19 of their variance: no exponent to use. Each fit has six points, and these fits are the planner's model of §3, not the Bench's truth.
Published fits are the same kind of object. Lin et al. (2024) fit the optimality gap, one minus the normalised score, as a power law in the number of training environments, objects or pairs (six points a line), with exponents from −0.47 to −0.84. They call a floor as a third parameter not statistically justified at six points, in three of six lines the fitted gap at one environment exceeds the largest gap there can be, and they find no clear power law in the number of demonstrations, only a plateau near 800. How much carries to a real task? The Bench's whole own curve, 256 layouts, is 0.66 hours of motion and a layout is described by two numbers; lesson 7 found that 90 % takes 14 layouts for one number, 104 for two, and about 5.8 times as many for each further number, while π0's post-training sets run from 5 hours to over 100. Exponents, saturation points and totals belong to this task and policy; the shapes, and the rule built on them, are what carry.
3 · Equal marginal value per dollar
Let V be the policy's success after buying hs hours from each source s at ps dollars an hour, with a budget B = Σ pshs. The value of the next hour of s, per dollar, is λs = (∂V/∂hs) / ps.
- Take a plan with λa > λb and move one dollar from b to a: success changes by λa − λb > 0, so the plan was not the best. At the best plan every source bought has the same λ, and one not bought has a λ at zero hours no higher.
- With power-law rates the hours of s are worth Es = kshsβs own hours, and V = F(n0 + ΣEs), with F the own curve of lesson 16 and n0 the own layouts already held.
- So ∂V/∂hs = F′ ksβshsβs−1 with the same F′ for every source, and equal λ is an equal marginal cost of an own-equivalent hour: cs(h) = ps / (ksβshβs−1) = μ.
- Solving, hs = (μ ksβs / ps)1/(1−βs), with μ the value that makes Σ pshs = B.
Price enters through the exponent 1/(1 − β). At a fixed μ, halving a source's price multiplies the hours bought by 21/(1−β): 5.7 for β = 0.6, the older arm's, and 545 for β = 0.89, the twin's. A source whose rate hardly falls is a corner, bought whole or not at all. Your own robot has β = 1 and a flat marginal cost, its price, so $123.6 is the ceiling on μ: no source is bought past the point where its next useful hour costs what your own does. With the fits of §2 and lesson 17's prices (assumptions, not measurements), the next useful hour costs, after 8, 64 and 256 layouts:
| source | price per hour | after 8 layouts | after 64 | after 256 |
|---|---|---|---|---|
| simulator, 1 % gap | $0.19 | $0.34 | $0.66 | $1.01 |
| twin arm | $15 | $12.6 | $15.9 | $18.5 |
| older arm | $15 | $30.8 | $70.6 | $122.7 |
| own robot | $123.6 (kept hour) | $123.6 | $123.6 | $123.6 |
Lesson 17's order at the first hours, simulator, twin, older arm, own, footage, holds, with one change. The older arm's next hour costs 4.0 times less than your own at 8 layouts and about the same at 256 (0.99 of it); the fit crosses at 261 layouts, and past that the older arm costs more than recording your own. The simulator's next hour triples and still costs 18 times less than the twin's. Footage, at $33.75 an hour, has a best rate at any point of 0.27, which never brings its useful hour below $125, about what your own costs. As a plan: the simulator first, the twin next, the older arm until it costs what your own does, exact if hours add. §4 asks the policy whether they do.
4 · A pool is not a sum
§3 gave every source a rate in own layouts and read one curve at the total. Test that on the policy. Each source supplies its own list of 256 layouts, so a pool never buys the same layout twice; the lookup takes the nearest stored layout of any source; and every number is the success on the same 1000 new layouts. Pools built on 8 own layouts, with a simulator whose arm model is off by 1 % and whose labels are joint angles (lesson 8):
| pool | success |
|---|---|
| 8 own layouts | 21.6 % |
| + 256 twin layouts | 98.7 % |
| + 256 simulated layouts | 82.2 % |
| + 256 twin and 256 simulated layouts | 94.2 % |
The same 256 simulated layouts add +60.6 points to the pool of 8 own layouts and −4.5 to the pool that holds the twin's. (The ledger's table reads the twin's 256 layouts at 98.8 % and the simulator's at 82.3 %, these pools 98.7 and 82.2: the table gives every foreign source one list of layouts, a pool a list of its own. These are one draw of the layouts, the one nearest the medians of eight; on all eight the simulator lowers the twin's pool, by 2.8 to 5.8 points.)
The mechanism is the lookup, which cannot see quality (lesson 8). In the pool with both, the nearest stored layout belongs to the simulator for 49.1 % of the test layouts, to the twin for 48.9 % and to you for 2.0 %. Following a simulated demonstration completes the course 88.2 % of the time, a twin's 100.0 %. Success is the sum over sources of share × reliability, and a source of lower reliability that wins half the shares lowers it. Sources enter a pool as competitors for the nearest slot, not as additions. The sum model of §3 misses this: it reads the own curve at 8 + ρtwinn + ρsimm with the ledger's rates, predicts 93.6 % for the best split of $3, and the policy measures 89.0 %, 4.6 points lower.
The widget
What to try. Leave the defaults: $3.00, the twin at $15 an hour, a simulator off by 1 %. The best split spends $2.98 on 159 simulated and 75 twin layouts and reaches 89.0 %, where the sum of rates predicts 93.6 %. Simulated layouts are the nearest for 62.3 % of the test layouts and complete 83.1 % of them, the twin's 33.7 % and 98.8 %. The four rules reach 82.2, 86.3, 88.4 and 88.4 %. At $6.00 the best split buys no simulator and 154 twin layouts, 96.5 %, and the ranking reaches 92.1 %; at $10.00 the best is 98.7 %. Choose the older arm and put the twin at your own hour's price: the best split at $3.00, $6.00 and $10.00 is 59 older-arm and 2 twin layouts, then 145 and 1, then 46 and 25, not monotone either. Choose the 3 % simulator with joint labels, the source that sells nothing (lesson 17): all the money on it reaches 5.1 % at $6.00, below the 21.6 % of your 8 layouts alone. With hand labels the sum predicts 99.0 % at $3.00, the policy measures 99.5 %, and the ranking is within 0.1 points of the best split: its demonstrations complete 99.5 % of the layouts they serve, about as many as the twin's.
5 · The rules of thumb, scored by the policy
Four rules for spending a budget, each scored by the policy, against the best split found by trying every number of simulated layouts the budget allows. All on the simulator spends everything on the cheaper hour; half each gives each source half the money; to 90 % each buys each source until it alone would reach 90 %; lesson 17's ranking buys the source with the lower cost per useful hour to the end of its catalogue, then the other.
| budget | best split: simulated + twin layouts | all on the simulator | half each | to 90 % each | lesson 17's ranking |
|---|---|---|---|---|---|
| $1 | 245 + 22: 85.3 % | 82.2 % | 84.3 % | 85.2 % | 85.2 % |
| $3 | 159 + 75: 89.0 % | 82.2 % | 86.3 % | 88.4 % | 88.4 % |
| $6 | 0 + 154: 96.5 % | 82.2 % | 88.4 % | 89.4 % | 92.1 % |
| $10 | 0 + 256: 98.7 % | 82.2 % | 90.7 % | 89.4 % | 94.1 % |
At $1 the ranking and the 90 % rule are within 0.1 point of the best split, because they spend nearly everything on the simulator and so does the best split. Each fails differently later. All on the simulator stops at 82.2 % from $0.12, where its 256 layouts run out. Half each spends $5.12 of $10: the simulator's 256 layouts cost $0.12 and the rest of its half is left over. To 90 % each buys the simulator's 256 layouts, which alone reach 82.3 %, then the 92 twin layouts that alone reach 90 %; together they reach 89.4 %, and it stops at $3.69. Lesson 17's ranking is the best of the four and still 4.4 points behind at $6 and 4.6 at $10.
The best split is not monotone in the budget. It buys the simulator first, 245 layouts at $1 and 159 at $3, and none at $6 or $10; over eight draws of the layouts it falls below 32 simulated layouts at a budget between $3 and $5, and at $10 it buys at most 1. A block of 64 simulated layouts costs $0.031. On 8 own layouts it adds 38.1 points, 1,229 per dollar; in the pool that holds the twin's 256 layouts it adds −3.1 points, −100 per dollar. The twin's first 64 layouts cost $2.48 and add 61.6 points, 24.8 per dollar. At the start the simulator's block is worth 49 times the twin's per dollar, and after the twin it is worth less than nothing. That is how the cheapest source changes as you buy: the marginal value per dollar of each source depends on what the others have already put in the pool, and the λ's cross.
6 · What the rule assumed
Every step of §3 read a curve of the hours bought, and bought is the word to check. The rule treats hours as an asset: once in the pool, their value is what the curve says, today and next year. §4 already broke one version of that: the same 256 simulated layouts are worth +60.6 points to a pool of 8 own layouts and −4.5 to a pool that holds the twin's, so an hour's value is not a property of the hour. That is about the pool at one time. The rule also assumes the value stays when the policy is retrained, and the two kinds of record the ledger buys behave differently there.
A demonstration is a record of what the expert did in a layout, the same record whichever policy reads it. The policy trained next year finds the same layouts and trajectories on file, and the coverage lesson 7 counts, the distance from a new layout to the nearest stored one, does not depend on which policy counts it. A demonstration keeps its value for next year's policy as for today's, so §3's rule applies to demonstrations as it stands.
A correction is a different record. It was made in a state that an older policy reached, and it labels that state with what the expert would do there (lesson 3). Its use depends on the policy that trains on it reaching that state, and the policy trained next year will not visit quite the same ones. How much that costs a correction is a measurement the rule does not contain, and until it is made the rule cannot be applied to a stock of corrections.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
A pool covers the layouts geometrically, each layout removing 2.4 % of the failures that remain, and the exchange rate of a poorer source falls as a power of its hours, with exponents from 0.11 for the twin to 0.40 for the older arm. A budget is spent well where the marginal value per dollar is equal across the sources bought, but hours do not add: at a budget of $3 to $5 the best split buys fewer than 32 simulated layouts, and lesson 17's ranking, bought down in order, loses 4.5 to 5.8 points at $10. All of it treats the hours as an asset that keeps its value once bought. Demonstrations do, for the policy trained next year as well as for the one trained today. Corrections were made in the states an older policy visited, and the policy trained next year will not visit quite the same ones. How long does a correction keep its value?
Interview prompts
- Why is a cost per useful hour not what the next hour costs? (§1, §3 — a price over an average rate: the older arm's next useful hour at 256 layouts costs $122.7, its average $88.7.)
- State the rule for splitting a budget across sources and derive it. (§3 — move a dollar from the lower marginal value per dollar to the higher until equal; with power-law rates, an equal marginal cost of an own-equivalent hour.)
- How does a geometric fall of failures differ from a power law, and why does R² not tell them apart? (§2 — a power law has a fixed local exponent, here 0.25 growing to 2.74, yet it fits at R² 0.93.)
- Why can adding a cheap source lower the success of a pool? (§4 — the lookup cannot see quality: simulated layouts are the nearest for 49.1 % of the layouts and complete 88.2 % of them against 100.0 % for the twin's.)
- Why does the best split drop the cheapest source as the budget grows? (§5 — 64 simulated layouts are worth 1,229 points per dollar on 8 own layouts and −100 on a pool with the twin's 256.)
- What does the rule assume about an hour once it is bought? (§6 — that it keeps its value: true of demonstrations, the open question for corrections, made in states the older policy visited.)
Companion reads: Lesson 7 · How much data: the coverage law, Lesson 8 · Other bodies and Lesson 3 · Labels on your own states.