all_lessons/Robot Model Training/20 · The flywheellesson 20 / 24

The flywheel: a fleet as a source

Lesson 19 left a stock of corrections that keeps its value, a supply of failures that runs out, and one cheap continuing source: the deployed robots, with a person stepping in when they fail. This lesson prices that fleet on the Bench. At an assumed 20 s of supervisor time a takeover costs 4.8 times what an attempt earns, so a fleet of first-clone robots needs 1.57 supervisors per robot and loses $27 an hour for each, and a fleet earns its supervisors only above a success rate of 79.4 %. Below it the programme pays for every point, whatever the size of the fleet, which only changes the days. Above it the work pays and the failures that supply the corrections grow rarer: the climb from 90 % to 98 % takes 6.1 times the days of the climb to 90 %. How much of each source to buy, it cannot say.

The thesis, here
A fleet's money and its data depend on one number, the share of attempts that fail. A failure costs a supervisor, supplies a correction, and teaches less each time as failures thin out. So there is a success rate s*, set by what a takeover costs against what an attempt earns, below which the programme funds every point and above which the work pays; and because the corrections come from the failures, the climb above s* slows as the policy improves.
Linear position
Forced by: A correction does not lose its value when the policy it was made for is retrained: a better policy stays among the states of a worse one, and on the Bench a set made for the first clone is worth as much per frame to the fourth clone as a fresh set, losing value only when the world moves outward, by about a fifth when the gust doubles. Demonstrations and force recordings, which the expert made and no policy did, keep theirs whoever is trained. What runs out is the supply: a correction teaches what its policy got wrong, the first set adds about eighteen points of success and the fourth about one, and a policy that fails one run in twenty has little left to show. The cheapest continuing supply of corrections is the deployed robots themselves, working, with a person stepping in when they fail. Does a fleet like that improve the policy fast enough to pay for the people who supervise it?
New idea: a fleet is a source whose supply, cost and earnings all depend on the policy's own failure rate p, so it breaks even at p* = (v − r) / c, where an attempt earns v, the robot costs r and a takeover costs c; below that rate it has to be funded until it takes off, and above it the supply of failures falls with p. The deficit is set by c and by the corrections needed, not by the number of robots, which only buys days.
Forces next: A fleet pays for its own supervisors only above the success rate at which a takeover costs less than the task it saves; below that rate every point the fleet gains is paid for by the programme, and above it the failures that supply the corrections grow rarer as the policy improves. We now have sources with falling returns, a stock of corrections that keeps its value while its supply runs out, and a fleet that has to be funded until it takes off, and each has a price, a value curve and a shelf life. Given a budget, how much of each do we buy, and in what order?
The plan
Six moves. (1) Run a fleet's hour at the first clone and count the supervisors. (2) Price one attempt and find the success rate at which it pays. (3) Measure what a takeover buys and how fast failures thin out. (4) Run the fleet over generations: the deficit, the days, the size of the fleet. (5) Rule out the alternatives by computation. (6) Ask what is left once the fleet pays.

1 · A fleet's hour at the first clone

Put a robot on the five-post course (lessons 1 to 3, lesson 19) and let it attempt the course again and again. An attempt is one run, 9.3 s of motion (lesson 16): 387.1 attempts an hour. The policy is lesson 1's nearest-demo clone, trained on 20 calm demonstrations; its success rate s is the share of 1000 unassisted runs under the gust of 0.05 that finish (lesson 19), and p = 1 − s is the share of attempts that fail, that is, would end in a post. A person stands by. As in lesson 3 the person needs 0.3 s (6 steps) to reach the controls, then drives with the expert's command until the cup is within 1 cm of the path and hands back, and only the frames they drove are labelled. This Bench person has a gift a real one lacks: they know which attempts fail. We find out by running the attempt twice from the same start under the same gust, once to see where the cup strays into a post and once with the person called at that stray. A failure that leaves them less than their reaction time cannot be prevented (0.28 % of the attempts measured below). §5 prices the gift.

Prices are lesson 17's assumptions plus two of this lesson's own; no source prints what a supervisor costs.

quantityvaluewhere it comes from
wage of a supervisor$40 an hourlesson 17
an arm$18,000 over 4,000 hours: $4.50 an hourlesson 17
share of a supervisor's time spent correcting, d0.50lesson 17: their utilisation, so waiting is paid for
supervisor time one failing attempt takes, τ20 sthis lesson: the Bench measures 1.5 s at the controls; finding the robot, taking the controls and handing them back is not on the Bench
hours a robot works in a day8this lesson

A success saves a person doing the 9.3-s task by hand: v = $40 × 9.3 / 3600 = $0.1033. The robot costs r = $4.50 / 387.1 = $0.0116 an attempt. A failing attempt costs a supervisor c = $40 × τ / (3600 d) = $0.4444 at τ = 20 s, $0.0222 a second, which is lesson 17's price of an hour of corrections without the arm: the work pays for the arm. Every attempt delivers its task, alone or with the person's help, so a robot-hour at failure share p earns 387.1 (v − r − p c) dollars and needs p × 387.1 × τ / (3600 d) supervisors: the failures an hour, times the seconds each takes, over the share of a supervisor's time that can be spent on them. Three policies from the measurement of §3 (n is the number of failing attempts the fleet has taken over so far):

policysuccessfailures an hoursupervisors per robotnet per robot-hour
first clone, n = 063.5 %141.31.57−$27.3
after n = 2083.1 %65.40.73+$6.4
after n = 10098.5 %6.00.07+$32.8

Ten first-clone robots need 15.7 supervisors and lose $273 an hour: at the first clone the fleet is its supervisors. By the third row each robot earns $32.8 an hour of the $40 its work is worth. Between the first two rows the policy crossed a rate. §2 finds it, and §4 asks whether the policy gets there fast enough.

2 · What one attempt earns

An attempt at failure share p earns e(p) = v − r − p c = 0.0917 − 0.4444 p dollars: it earns v, pays the robot r, and with probability p pays a supervisor c. It is worth running when e > 0, that is, below p* = (v − r) / c = 0.2063: a success rate of s* = 79.4 %.

The break-even exists because a takeover costs $0.444, 4.85 times the $0.0917 an attempt earns net of the robot. A takeover that costs more than the task it saves can still be paid for, but only if fewer than one attempt in 4.85 needs one. Were c below v − r, which is a takeover shorter than 4.13 s, every attempt would pay at every success rate and there would be nothing to fund. The 1.5 s the Bench measures at the controls is such a case. What makes a fleet a cost is the part of a takeover the Bench does not measure.

Supervisors follow p. At p* a robot needs 0.89 of them. One supervisor per robot at all times would cost $40 an hour, the whole output of a robot that never fails, so no fixed staff can be right: the cost has to fall as the failures do. And s* depends on τ and d only through τ / d, the paid seconds a failure costs:

τd = 0.25d = 0.5d = 1
10 s79.4 %58.7 %17.5 %
20 s89.7 %79.4 %58.7 %
40 s94.8 %89.7 %79.4 %

Equal values run along each diagonal: a supervisor who is busy half the time and takes 20 s a failure costs what one who is always busy and takes 40 s costs. The first clone, which fails 36.5 % of its attempts, is below s* at 20 s and above it at 10 s and less. Whether a fleet gets from one to the other depends on how quickly takeovers improve the policy.

3 · What a takeover buys

A takeover is a correction: on the Bench 23.1 labelled frames, 1.15 s of motion, for 1.5 s at the controls. The clone is retrained on the 20 demonstrations and every frame labelled so far (lesson 19: old corrections keep their value, so retraining is storing). A generation is ten failing attempts taken over, after which the clone is retrained and the fleet carries on with it. Four lineages, each with its own gusts and starts, each clone scored on 1000 unassisted runs, so a row is 4000 runs except the first, one clone, which scores 63.5 % here and 62.0 % in lesson 19 on other runs.

failures taken over, nfailure share pattempts per failure, 1 / pattempts a lineage made to get there
036.5 %2.70
1025.5 %3.921
2016.9 %5.964
406.65 %15208
1001.55 %64.52,303
4000.33 %307.765,106

Read it twice. As a curve: the failure share falls by 30 % over the first ten failures and then, from 20 on, as a power of n, a line of slope −1.29 on log–log axes (R² 0.965) from 16.9 % to 0.33 %. It is a measurement of this draw of four lineages, not a law, and its tail levels off near 0.44 %. As a supply: an attempt fails with probability p, so a robot makes 1 / p attempts for each failure, 2.7 at the first clone and 64.5 at n = 100, and the attempts a lineage actually made (last column) agree with the sum of failures / p to within 9 % at n = 100. Corrections arrive at p per attempt, and each is worth less than the last: the first failure taken over removes 1.31 points of failure, the 40th 0.13.

What a correction costs does not change: c per failing attempt, whatever p is. Per labelled frame that is $0.0193, 16 times lesson 17's $0.00124 for a dedicated labeller, because 20 s of supervisor are spread over 23.1 frames. The fleet does not win on the price of a frame. If it wins, it wins on what the work pays.

4 · The fleet over generations

A generation needs ten failures taken over. At task size K (K times the failures, prices and rates unchanged, as lesson 18 said of a larger task) it needs 10 K, which a policy with failure share pg collects in 10 K / pg attempts. So with F robots working 8 hours a day, generation g runs 10 K / (pg · 387.1 · 8 F) days, earns (v − r) for every attempt, pays c for each of its 10 K failures, and ends with the clone retrained, ng+1 = ng + 10:

netg = 10 K [ (v − r) / pg − c ]     daysg = 10 K / (pg · 387.1 · 8 · F)

A generation pays when pg < p*, the same break-even. The deficit is the sum of the negative nets before the first paying generation, and F appears in no term of it: ten times the robots make each generation a tenth as long and lose ten times as much a day. Fleet size buys days, not dollars. At the default (first clone, F = 10, K = 1000, τ = 20 s) generation 0 makes 27,397 attempts in 0.88 days and loses $1,932; generation 1 makes 39,177 in 1.27 days and loses $852; generation 2 runs the policy after n = 20, whose failure share 16.9 % is below p* = 0.2063, and pays. The deficit is $2,783, spent on 19.6 points (63.5 % to 83.1 %), $142 a point at K = 1000 and $0.14 at the Bench's own scale. That answers lesson 19's question at these prices: a first-clone fleet meets a failure every 2.7 attempts, so the policy improves fast enough, but not fast enough to pay for itself on the way, and someone funds the first 19.6 points. The Bench's climb to 98 % takes 2,346 labelled frames, so K = 1000 stands for a task needing a thousand times the corrections. Prices and rates carry to such a task, and s*, the shapes and the order of the alternatives with them; days and dollars grow with K.

The widget

Run a fleet from deployment to 99 %
Left: what one attempt earns after the robot (teal) and what its failures cost (red) against success; amber is s*, the dots the policy at the start and at 98 %. Right: the policy's success (top) and the fleet's cumulative cash (bottom, linear within plus or minus the deficit, logarithmic beyond) against days on a log axis; a step is a generation. The curve is §3's, prices are lesson 17's, a robot works 8 hours a day.
break-even success s*
-
supervisors per robot at the start
-
net per robot-hour at the start
-
first paying generation starts
-
deficit before it
-
programme cost per point
-
days to 90 %
-
days to 98 %
-
cash at 98 %
-
days to 99 %
-
Show the core JS
  var pStar = Math.min(1, (v - r) / c);
FL.net = function (E, p) { return E.v - E.r - p * E.c; };
FL.crew = function (E, p) { return p * E.perHour * E.tau / (3600 * E.duty); };
...
    var p = FL.p(n), att = o.K * FL.B / p, net = att * (E.v - E.r) - o.K * FL.B * E.c;
    if (take === null && net > 0) take = g;
    days += att / per; cash += net; n += FL.B; g++;

What to try. Leave the defaults: first clone, 10 robots, 20 s, K = 1000. s* is 79.4 %; each robot needs 1.57 supervisors and loses $27.3 an hour; the first paying generation starts on day 2.1 after a deficit of $2,783; 90 % comes on day 6.8, 98 % on day 48.9 with $98,762 in hand, and 99 % on day 221. Slide the robots to 1: the deficit is still $2,783 and every day count is ten times as long, take-off on day 21.5; at 100 robots it is on day 0.21. Slide the starting success to 70 % and the deficit falls to $1,388; to 80 %, above s*, there is none. Slide τ to 10 s: s* falls to 58.7 %, below the first clone, and there is no deficit; at 40 s s* is 89.7 %, the deficit $16,135, 5.8 times as large, and the first 4 generations lose; at 80 s it is $56,628. At 1.5 s, the Bench's own driving, there is no s*. Set K to 1: dollars and days divide by 1000, to $2.78 and 0.05 days to 98 %, and nothing else moves.

5 · The alternatives, ruled out by computation

Label every frame of every run. A person labelling at 0.05 s a frame costs $0.207 an attempt, 2.0 times what an attempt is worth, whatever s is, so the work can never pay for it. And the frames are worth less. Four hundred labelled frames raise the first clone's success by 15.4 points when they are takeover frames of failing attempts (eight draws, 11.0 to 18.4) and by 3.8 when they come from successful runs (−0.7 to 8.5), 4.1 times less per frame. For a clone that already succeeds 95.8 % of the time they add 0.8 and 0.2. The value of a frame is where the failures are.

Buy the points as demonstrations recorded under a gust. Ten demonstrations under the gust of 0.05 reach 83.5 % for $3.19, $0.16 a point above the 63.5 % of 20 calm ones; 80 reach 84.0 % for $25.55, $1.25 a point, and none of the eight sizes passes 84.0 %. At the Bench's scale the fleet's deficit buys the same ground for $2.78 and does not stop there.

Road not taken · supervise by a distance rule
Lesson 3's person takes over when the cup is more than 2.5 cm from the path, which needs no foresight. The Bench shows the price. The expert itself crosses 2.5 cm 4.6 times a run, so the person is called 4.6 times a run for a policy that never fails, 4.1 times for the first clone and 4.0 times for a clone that fails 1.8 % of its attempts: the calls do not fall as the policy improves. At the controls 4.4 s of every 9.3-s attempt costs $0.097 before any overhead, more than the $0.0917 the attempt earns, so such a fleet never pays. A rule at 4 cm calls the person 1.05 and 0.41 times a run for those two clones, 2.9 and 22.5 calls per failure. A call costs what a failure costs, so calls per failure multiply c: at 22.5 the break-even is 99.1 %, above where that clone stands. Luo et al. (2024) report the same contrast: the intervention rate of HIL-SERL falls during training, that of its HG-DAgger baseline does not.
What this lesson did not do
The supervisor knows which attempts fail, found by running the attempt twice; nothing here says how a person or a monitor would know. Labels are exact, though a person's drift (lesson 4), and the 0.3 s reaction is an assumption, not a measurement of people. The robot is not charged for waiting for the person, an attempt needs no reset, and retraining, evaluation and the arms' purchase price are free or inside r. The curve is one task and one clone that stores frames, measured on four lineages: pooling any three of them gives deficits from $2,638 to $2,967 and days to 98 % from 48 to 51, so the shapes and orders of magnitude are the measurement and the third digit is not. The 0.28 % of failures too fast for a person teach nothing and sit inside p. The Bench is small, its climb to 98 % takes 2,346 labelled frames, so K stands in for a real task's scale, and τ, d and the wage are assumptions. How the fleet combines with the other sources under a budget is lesson 21.

6 · What the fleet leaves

Above s* the work pays for the corrections, so the points cost no dollars. They cost time. The loop the title calls a flywheel, failures supplying corrections that remove failures, feeds on its own supply, and each turn leaves less to feed it. A generation takes 10 K / pg attempts: at the default the first generation takes 0.88 days, the one that brings the policy to 90 % 2.78 days, and the one that brings it to 98 % 13 days. The policy reaches 90 % on day 6.8, 98 % on day 48.9 and 99 % on day 221: the stretch from 90 % to 98 % takes 6.1 times the days of the climb to 90 %, and the stretch to 99 % takes 31 times. More robots buy the days back, and the work pays for them, since every robot earns above s*: 100 of them take 0.7, 4.9 and 22 days. A fleet of cars shows the same thinning of supply: Waymo's safety-related disengagements in California fell from 0.8 to 0.2 per 1,000 miles between 2015 and 2016 (Waymo, 2017), and each one is a correction.

The price of a point shows what is left to decide. A dedicated round of corrections (lesson 19: 800 labelled frames at lesson 17's price, $0.99) buys a point at the first clone for $0.053. The fleet buys it, net of what its attempts earn, for its c per failure over the points a failure removes:

clone (lesson 19)failure sharededicated round, per pointfleet at the same failure share, per point
138.1 %$0.053$0.156
222.6 %$0.101$0.041
313.1 %$0.182paid $0.52
45.6 %$1.158paid $10.59

A dedicated round buys the first points for 2.9 times less than the fleet and the later ones for more and more, 22 times more by the fourth clone, because its supply of failures thins as the fleet's does and it pays for the robot's time as well as the labels. The fleet's price falls through zero at s* and then pays. Neither source is cheapest everywhere, and a third, demonstrations recorded under a gust, has a price and a ceiling of its own.

Common mistakes / failure modes

"the fleet improves the policy, so it pays for itself"
Below s* the programme pays for every point: $2,783 for the first 19.6 at K = 1000, and $273 an hour for ten first-clone robots until then (§1, §4).
"a takeover is a few seconds of driving, so supervision is cheap"
At the Bench's 1.5 s there is no s*. The cost is everything around the driving, and at 20 s s* is 79.4 % (§2).
"label everything the fleet records"
$0.207 an attempt, 2.0 times what it earns, for frames worth 4.1 times less (§5).
"the supervisor can step in when the cup strays"
At 2.5 cm the expert itself is called 4.6 times a run, and the calls do not fall as the policy improves (§5).
"once it pays, the fleet takes off"
It pays and it slows: 98 % takes 7.1 times the days of 90 % (§6).

Checkpoint exercise

Try it
A supervisor costs $40 an hour and is busy half the time; a takeover takes 30 s of their time; an attempt earns v − r = $0.0917, as here. (a) What does a failing attempt cost? (b) At what success rate does the fleet break even? (c) A fleet of 20 robots has a policy that fails 10 % of its attempts: how many supervisors does it need? (d) What does one robot-hour earn? Answer: (a) c = 40 × 30 / (3600 × 0.5) = $0.667. (b) p* = 0.0917 / 0.667 = 0.1376, so s* = 86.2 %. (c) Per robot 0.10 × 387.1 × 30 / (3600 × 0.5) = 0.645, so 12.9 for 20 robots. (d) 387.1 × (0.0917 − 0.1 × 0.667) = $9.7: positive, because the policy's 90 % is above s* = 86.2 %.

Where this points next

A fleet pays for its supervisors above s* = 79.4 %, where a takeover, $0.444, costs what an attempt earns net of the robot, $0.0917, once in 4.8 attempts. Below it the first 19.6 points cost $2,783 at K = 1000 whatever the number of robots; above it the work pays while the failures that supply the corrections thin out, and 98 % takes 7.1 times the days of 90 %. A dedicated round buys a point for 2.9 times less than the fleet at the first clone, the fleet buys it for less from the second, and above s* the fleet is paid. We now have sources with falling returns, a stock of corrections that keeps its value while its supply runs out, and a fleet that has to be funded until it takes off, and each has a price, a value curve and a shelf life. Given a budget, how much of each do we buy, and in what order?

Takeaway
A fleet is a source whose supply, cost and earnings all depend on the policy's failure share p. An attempt earns v − r − p c, so it pays below p* = (v − r) / c, a success rate of 79.4 % at the assumed prices, and the break-even exists only because a takeover costs more than the task it saves. Below s* the programme funds every point, $2,783 for the first 19.6 at K = 1000, and that sum does not depend on the number of robots, which only change the days. The failure share falls like a power of the failures taken over and the supply is 1 / p attempts per failure, so above s* the climb slows: 98 % takes 7.1 times the days of 90 %. The ledger now has sources with a price, a value curve and a shelf life each, and none is cheapest at every success rate.

Interview prompts

Companion reads: Lesson 3 · Labels on your own states (the takeover) and Lesson 4 · People are not functions (what a person's labels are worth).