all_lessons/Robot Model Training/24 · What it eatslesson 24 / 24

What does embodied training eat the most?

Lessons 16 to 20, 22 and 23 each measured one thing about an hour of data, and lesson 21 spent a budget with what lessons 16 to 19 had found. Its plan, $34,938 for P = 94.9 %, is wrong three ways once lessons 20, 22 and 23 are billed: moving the hours adds $4,816, a fleet buys the recovery for $924 where gust demonstrations cost $6,727, and a quarter-new corpus delivers 61.6 %; a planner that trusted the cheapest hours would reach 33.5 %. This lesson finds where each factor enters a purchase, which fixes the order of a seven-step procedure, runs it on one programme, scores four planners that each leave a step out, and reads what the plan eats: by volume borrowed demonstrations, by value force-bearing hours. The prices are this year’s, on a small Bench.

The thesis, here
A purchase is a point: dollars out, success in. Every factor of the ledger enters it in one of two places, the dollars (the price, the cost of moving, fixed costs) or the level (the column’s curve, the zeros, the share that is new). A wrong dollar shows on the bill; a wrong level shows only when the policy is measured. The procedure is the order in which the factors must be read to compute the points.
Linear position
Forced by: Moving the data costs a fixed share of the budget and decides whether the processors are ever busy; it is the last bill. We can now price every source, its value, its shelf life, its duplicates and its transport. Put together, which kind of data does embodied training eat the most, by volume and by value, and what should be bought next?
New idea: read each factor where it enters a purchase, in that order: columns and zeros, prices with moving and the new share, the fleet’s take-off, the lumps priced with their hours, the sequence by marginal value per dollar of ln P, and a measurement after each block. Leaving one step out costs a number that §4 computes.
Forces next: The ledger says what to buy next for one policy on one robot, at this year's prices and exchange rates. Prices move, simulators improve and models get larger, and some columns, force and recovery and intent, are absent from every cheap source. Which of these prices will move first, and which of the zeros will not?
The plan
Six moves. (1) Bill lesson 21’s plan again with lessons 20, 22 and 23. (2) Find where each factor enters a purchase and what kind of error follows. (3) Run the seven steps on a stated programme. (4) Score four planners that each leave one step out. (5) Read what the plan eats, by volume and by value. (6) Move the prices and see what moves.

1 · Lesson 21’s plan, billed again

Lesson 21’s best plan for the programme, which must generalise, recover and make contact, at task size K = 1,000 (lesson 18’s device: hours multiplied, prices unchanged), buys the twin arm’s 256,000 layouts, 20,000 demonstrations under a gust, and the force rig with 10,000 insertions: $34,938 for P = 94.9 %. It counts the dollars of hours and nothing else. Bill it with lessons 20 and 23:

purchaselesson 21’s billbilled again
twin arm, 256,000 layouts$9,920$14,274: moving them costs $4,354, 44 % of the twin’s $15 hour
20,000 gust demonstrations$6,387$6,727 (97.5 %); 2,000, then the fleet: $924 (98.4 %)
the rig and 10,000 insertions$18,631$18,753
the plan$34,938, P = 94.9 %$39,754, 13.8 % more; with the fleet $33,951, P = 95.7 %

Two faults do not show on a bill. If a quarter of what is recorded is new (lesson 22), the same plan delivers 61.6 % (generalise 82.0, recover 91.0, contact 82.5) and cannot tell. And a planner that trusted the cheapest hours would buy none of the short columns: the simulator’s layouts, calm demonstrations and force-less insertions end at 33.5 %, for calm hours stay at 63.5 % and force-less ones at 64.2 % however many are bought (lessons 16, 17 and 21).

The repairs interact. Moving adds 44 % to the cost of the twin’s hours, which puts its first 128,000 layouts behind the rig in the order of purchases: 0.0183 against 0.0229 of ln P per $1,000, where lesson 21 had 0.0266 against 0.0230. The fleet makes recovery the second purchase and nearly free. Mending the plan one lesson at a time is not a method. Where does each factor enter a purchase, and in what order must they be read?

2 · Where each factor enters a purchase

Buy n units of a source s for a column c (a unit is a layout, a demonstration or an insertion, counted at scale 1 and multiplied by K). A unit costs its price ps and the cost ts of moving it to the accelerators (lesson 23), both in dollars a unit, and the purchase carries a fixed cost F once (the rig, the correction pipeline, a fleet’s deficit). Only a share qs of the units are new (lesson 22’s effective fraction; 1 for rendered units), so they act on the column’s curve sc,s, a cell of the table, as qs n units would, and give the column its success, the level:

x = F + K n (ps + ts)     level = sc,s(qs n)

The pair (x, level) is a point, a plan is one point per column, and the programme succeeds with P, the product of the levels. Every factor of the ledger has a place in these two lines:

factorlessonenters asa mistake shows on
exchange rate, curve16, 18the function sc,sthe level
structural zero16, 17sc,s flat: calm hours stay at 63.5 %, force-less ones at 64.2 %the level
price ps, moving ts17, 23dollars a unit (t = 0 for rendered hours)the bill
new share qs22the argument qs nthe level
fixed cost F20, 21a lump once; a fleet’s deficit depends on the level it starts fromthe bill
shelf life19what the next policy finds still worth its price, and what is left to buythe next budget

Two kinds of error follow, and they are not alike. A wrong dollar shows on the bill as soon as the first block is paid. A wrong level shows nowhere in the plan, which keeps promising its P; only a measurement of the policy can say that it is short. The order of the procedure is the order in which the places can be filled. The columns and the zeros decide which points exist, so they come first. Price, moving and the new share fix each source’s points before sources are compared, and its class (lesson 19) says what survives to the next policy. The fleet’s deficit depends on the level reached before it starts ($2,783 from the first clone, $77 after 2,000 gust demonstrations, none after 3,000, §3), so it is read after the cheaper levels. The rig, the pipeline and the deficit are lumps and are priced with the hours they unlock (lesson 21). Only then can the points be sorted by marginal value per dollar of ln P, and the last step is the measurement.

3 · The procedure, run on a programme

The programme is lesson 21’s at K = 1,000, on the lab’s own arm with three 1080p cameras, a budget of $30,000, a fleet of 10 robots (lesson 20’s default, a takeover of 20 s) with 60 days to finish, lesson 23’s ten epochs and twelve months of storage, and the ledger’s prices. The widget below moves each of these. The seven steps, with what each reads:

stepreadson the programme
1 columnswhat the task needsgeneralise, recover, contact, from 21.6, 63.5 and 64.2 %: P = 8.8 %
2 zeroswhich source has a point in which columnrecovery from gust demonstrations, rounds of corrections or the fleet; contact only from force-bearing hours behind the rig; the 3 % simulator sells nothing
3 sourcesrate, price, moving, new share, curvethe table below
4 classescapital, inventory, correctionsdemonstrations and force recordings are capital, kept for the next policy; the simulator is inventory and stops at 82.3 %; corrections keep their value, but a fresh set adds 18.5 points to the first clone and 0.85 to the fourth
5 take-offs* and the level the fleet starts froms* = 79.4 %: $2,783 of deficit from the first clone, $77 after 2,000 gust demonstrations (79.0 %), none after 3,000 (88.5 %)
6 sequenced ln s per dollar, lumps with their hoursthe five blocks below
7 measurethe bill, then the level, after each block§4

Step 3 gives each source one row, priced per motion hour (a layout is 9.3 s; footage is moved as one of the three cameras, an assumption). The rate ρ is lesson 16’s at 8 own layouts and 64 attempted ones, lesson 17’s operating point; the last two columns are lesson 17’s cost per useful hour and the same with moving in:

sourceclassprice p, $/hrate ρmoving t, $/hp / ρ(p + t) / ρ
simulator, 1 % gapinventory0.190.4200.440.44
twin armcapital15.001.016.5814.821.3
older armcapital15.000.380.3539.640.6
own robotcapital123.6116.58123.6130.2
footageinventory33.750.162.19206.4219.8
a supervised hourcorrections89.00recovery only6.58no rate on the layouts
a force-bearing hourcapital197.78 and the rigcontact only6.62no rate on the layouts

Moving does not reorder lesson 17’s ranking of the five sources that have a layouts rate; it makes the twin’s useful hour 44 % dearer. With a quarter of the borrowed layouts new it costs $85.2 and the older arm’s $162.2, above the $130.2 of an own hour whose layouts are all new: the older arm leaves the ranking, the twin stays. Step 6 sorts the blocks of the programme’s plan. The twin’s 128,000 layouts cost $7,137 and the sequence charges $7,013, because they replace the simulator’s $124 (a plan holds one of the two, lesson 18):

blockreachescostln P per $1,000total
simulator, 256,000 layoutsgeneralise 82.3 %$124120.9 to 2.09 over six steps$124
2,000 gust demonstrations, then the fleetrecover 98.4 %$9240.4735$1,048
the rig and 10,000 insertionscontact 98.5 %$18,7530.0229$19,801
twin arm, 128,000 layoutsgeneralise 93.6 %$7,0130.0183$26,814
twin arm, 256,000 layoutsgeneralise 98.8 %$7,1370.0076$33,951

Recovery is nearly free because step 5 changes the route: 2,000 gust demonstrations ($673) stop just short of s*, the fleet starts with a deficit of $77 and moving the 26.4 hours of frames it labels costs $174, so the corrections come at $924 in all, in 55.5 days; from the first clone the fleet costs $2,981, and gust demonstrations alone $6,727. The widget runs the seven steps and lets a planner skip one.

The widget

The whole ledger as one plan
Left: success P against the budget (log axis) for the best plan (teal) and, when a step is left out, what that planner gets by buying down its own list at the world’s prices (red) and promises (dashed); amber is the budget. Right: the best plan’s bill, hours (teal), moving them (purple), fixed costs (amber), and the sequence of §3, filled where the plan holds the block.
the plan costs
-
P of the plan
-
generalise
-
recover
-
contact
-
hours bought, dollars
-
moving them, dollars
-
fixed costs, dollars
-
motion hours bought
-
stored
-
next block costs
-
the planner promises
-
the planner gets
-
behind the procedure
-
Show the core JS
    q = recorded(d.src) ? V.q : 1;
    r.hrs = K * d.u * pu; r.mov = V.t ? K * d.u * mv[kind] * H1 : 0; r.hours = K * d.u * H1; r.gb = r.hours * gb[kind];
    r.s = lerp(cv.gen[key].u, cv.gen[key].s, q * d.u);
      f = fleet(from.s, o);
      r.hrs = from.hrs; r.mov = from.mov + (V.t ? f.hours * mv.corr : 0); r.inst = from.inst + (V.k ? f.deficit : 0); r.hours = from.hours + f.hours; r.gb = from.gb + f.hours * gb.corr;
  r.x = r.hrs + r.mov + r.inst; return r;
...
    b = seq[i]; c = b.col; if (!open[c] || b.k !== at[c] + 1) continue;
    cost = b.to[X] - m.cols[c].hull[at[c]][X];
    if (cost <= left + 1e-9) { left -= cost; at[c] = b.k; } else open[c] = false;

What to try. Leave the defaults: $30,000, three columns, the whole procedure. The best plan buys the twin’s 128,000 layouts, 2,000 gust demonstrations then the fleet, and the rig with 10,000 insertions, for $26,814 (hours $9,230, moving $2,507, fixed costs $15,077) and reaches P = 90.7 % (93.6, 98.4, 98.5); the next block, the twin’s other 128,000 layouts at $7,137, is $3,951 out of reach. At $20,000 the plan drops the twin and keeps the simulator: 79.7 %. Then let the planner leave a step out, as in §4: without moving the data, at $20,000, it gets 62.4 %, 17.4 points behind; without the take-off, at $3,000, 33.5 % against 51.9 %; without the zeros, at $30,000, it promises 95.7 % and gets 62.3 %. Set the new share to 1/4: the procedure buys no twin and reaches 66.8 %, while the planner that ignores the share promises 90.7 % and gets 55.3 %. A 0.3 % simulator at $20,000: 94.2 % for $19,801, generalise at 97.2 %.

4 · Leave one step out

The procedure and four planners follow the rule of step 6: buy down your own list, a block that does not fit closing its column. They differ in one belief each, pay the world’s prices and are paid the world’s levels, so a gap is the cost of the belief. The procedure’s list is within 4.4 points of the exhaustive best plan at all of 101 budgets from $1,000 to $50,000, 0.30 on average. Success in %, at four budgets:

planner$3,000$12,000$20,000$30,000promises at $30,000
the procedure51.959.179.790.790.7
leaves out moving the data51.959.162.490.795.7
leaves out the take-off33.559.062.390.690.6
leaves out the zeros33.551.959.062.395.7
a quarter new: the procedure41.751.966.866.866.8
a quarter new: leaves out the share33.543.051.755.390.7

Moving the data is a wrong dollar: at $36,000 the plan it believes in costs $29,267 and the world sends $33,951, 16.0 % more, and because moving is 44 % of a twin hour and little of the rest, the list is reordered. At $20,000 the planner buys the twin’s 256,000 layouts and the recovery for $15,198; the $4,802 left cannot pay for the rig. The take-off is a wrong dollar too: the fleet from the first clone costs $2,981, $2,783 of it the deficit, not the $197 believed, and the step that lifts the policy above s* with demonstrations first is never taken. The zeros and the new share are wrong levels: the plans cost what they were said to cost, no bill warns, and the gap is between a promise and a measurement. With a quarter new the procedure keeps the simulator, takes the fleet and buys the rig with 4,000 insertions; the planner that leaves the share out buys the twin.

Each omission is invisible at some budgets and costly at others. Over the 101 budgets the planner that leaves out moving is more than a point behind at 8 of them, worst 17.4 points; the take-off at 42, worst 18.4; the zeros at 83, worst 33.4; the share, with a quarter new, at 37, worst 24.4. A check at one budget proves nothing, so step 7 measures after every block. The level tells the share: 128,000 twin layouts of which a quarter are new lift generalise to 68.2 % where the table says 93.6 %, and the twin’s curve read backwards at 68.2 % returns 32 new layouts in 128, a share of 0.25.

Road not taken · one number per source
Put the factors into one number, the all-in cost per useful hour (p + t) / (q ρ), rank the sources by it and buy down the ranking, each to the end of its catalogue. It is tempting because it is lesson 17’s ranking with lessons 22 and 23 inside, and it keeps lesson 17’s order. At $30,000 it buys the twin’s 256,000 layouts and the 20,000 gust demonstrations, has $8,999 left and cannot pay for the rig ($15,375 with 1,000 insertions): P = 61.8 %, against 90.7 %. A rate is defined per column and most sources have none in two of the three, so one number cannot compare a layout with a force-bearing hour; a lump and a take-off have no price per unit; and the number is the price of the first hour (lesson 18).
What this lesson did not do
The plan is for one policy, one robot, one draw of each table (cells 200 or 1,000 runs wide, intervals ±0.02 to ±0.07) and this year’s assumed prices, at a task a thousand times the Bench’s: orders, ratios and thresholds carry, dollars do not. The fleet entered above the first clone follows lesson 20’s starting-success slider, not a measured lineage, and the share of new layouts is one number for every recorded source. The accelerator time that trains (lesson 23’s learning line) is paid whatever the kind of data and is left out. The Bench has no intent column, so the plan cannot price it.

5 · What it eats

The plan at $30,000, by kind of data (the simulator’s row is the block the twin replaces):

kindhoursstored GBdollarsof the billln P gainedper $1,000
the twin’s 128,000 layouts330.73,571$7,13726.6 %1.470.205
recovery: gust demonstrations, then the fleet31.6341$9243.4 %0.440.473
force-bearing insertions and the rig18.4200$18,75369.9 %0.430.023
total380.64,113$26,814100 %
(simulator, 256,000 layouts)661.30$1240.5 %1.3410.8

By volume the programme eats the twin’s demonstrations: 86.9 % of its hours and 86.8 % of its bytes, 18 times the hours of the force-bearing kind. By value, meaning the bill, it eats force-bearing hours: 4.8 % of the hours and 69.9 % of the bill. The rankings do not agree: the kind with the most hours has 26.6 % of the bill, the kind with the fewest 69.9 %. The whole plan (661.3 twin hours, 7.1 TB) has the same shape: 55.2 % of the bill for the force kind, 42.0 % for the twin, 2.7 % for recovery. Per dollar the order of the bill is exactly reversed, in ln P per $1,000 from the empty start: the simulator 10.8, recovery 0.47, the twin 0.21, the rig 0.023; and the twin adds 62.9 % of the ln P that the plan gains, the force kind 18.4 %. The programme gets most of its success from hours that many sources can make and pays most of its bill for the one column only a person at the arm can fill.

What should be bought next? Past this plan the next block is the twin’s other 128,000 layouts: $7,137 for +5.2 points of generalise, P = 95.7 %, $3,951 beyond the budget. What would serve better is not for sale: a simulator with a 0.3 % gap (§6).

6 · What would move

The best plan at $30,000, planned again when one assumption moves:

changebillchange in the billP
none$26,81490.7 %
accelerators $0.25 an hour$26,289−2.0 %90.7 %
accelerators $25 an hour$20,639−23.0 %79.7 %
a simulator with a 0.3 % gap$19,801−26.2 %94.2 %
a force rig at $1,500$20,451−23.7 %95.7 %

A tenth of today’s compute price takes 2.0 % off the bill: idle accelerators are 23 % of the moving bill and moving 9.3 % of the plan (ten times dearer, they push the twin out and the bill falls with P). A free simulator saves $124, 0.6 % of the $19,801 plan that holds it. What moves the plan is the simulator’s gap, not its price: at 0.3 % it reaches 97.2 % of generalise for the same $124 and replaces the twin’s $7,137, so the plan at $20,000 goes from 79.7 to 94.2 %. The one price on the force line that can fall is the rig’s: at $1,500 the bill falls 23.7 % and P rises to 95.7 %.

What is left when the cheap prices fall is the scarce column. With the 0.3 % simulator the plan costs $19,801 and $18,753 of it, 94.7 %, is the force-bearing line: moving its 18.4 hours costs $122, so no compute price reaches it, and a free rig would still leave $3,753 for the hours themselves. Recovery costs $924 and 55.5 days of ten robots, with a supply of failures that thins as the policy improves; intent has no column on the Bench at all.

Common mistakes / failure modes

"rank the sources by what a useful hour costs and buy down the ranking"
At $30,000 that reaches 61.8 % where the sequence reaches 90.7 %: a rate exists in one column, and the rig is a lump (§4).
"moving the data is a rounding error"
It is 13.8 % of lesson 21’s bill and 44 % of a twin hour; left out, the planner is 17.4 points behind at its worst budget (§1, §4).
"a fleet pays for itself"
Above s* = 79.4 % it does; from the first clone the programme pays $2,783, and 2,000 demonstrations first cut that to $77 (§3, §4).
"the cheapest hour is the one to buy"
Cheap hours stop at 33.5 % and the leaves-out-the-zeros planner at 62.3 % while promising 95.7 %; the force-bearing hour is the dearest and 69.9 % of the bill (§4, §5).

Checkpoint exercise

Try it
A borrowed source costs $10 an hour and $5 to move, has a rate of 0.8 on the layouts, and only a quarter of its hours are new. Your own hour, all of whose layouts are new, costs $123.61 and $6.58 to move. (a) What does a useful hour of the borrowed source cost? (b) With 1/8 new? (c) At what new share does it cost as much as your own? Answer: (a) (10 + 5) / (0.25 · 0.8) = $75.0, against $130.19 for your own, so it is still the cheaper buy. (b) At 1/8, $150.0, dearer. (c) (10 + 5) / (q · 0.8) = 130.19 gives q = 0.144, one hour in 6.9; below it the borrowed hours are worth less than recording your own (§2, §3).

Where this points next

The procedure turns the measurements of eight lessons into one plan: at $30,000 it spends $26,814 for P = 90.7 %, and what it eats depends on whether you count hours or dollars, the twin’s demonstrations by volume, force-bearing hours by value. A plan is for one policy on one robot at this year’s prices, and its numbers say what to watch: a tenth of the compute price takes 2.0 % off the bill, a simulator at a 0.3 % gap takes 26.2 % off and adds 3.5 points, and what remains is the scarce columns: force and recovery are 73.4 % of the bill at $30,000, and with that simulator the force line alone is 94.7 % of the plan, which no compute price reaches. Which of these prices will move first, and which of the zeros will not?

Takeaway
A purchase is a point (dollars, success) and every factor of the ledger enters its dollars (price, moving, fixed costs, the fleet’s deficit) or its level (the curve, the zeros, the share that is new). Dollar errors show on the bill and level errors only in a measurement, so the procedure reads columns and zeros, then prices and shares, then the take-off, then the lumps with their hours, buys by marginal value per dollar of ln P and measures after each block. At K = 1,000 and $30,000 that is the twin’s 128,000 layouts, 2,000 gust demonstrations then the fleet, and the rig with 10,000 insertions: $26,814 for 90.7 %. Leaving a step out costs up to 33.4 points (the zeros), 18.4 (the take-off), 17.4 (moving) or 24.4 (the share), at different budgets. By volume the programme eats borrowed demonstrations, 86.9 % of its hours; by value, force-bearing hours, 69.9 % of its bill from 4.8 % of the hours. Prices are assumptions; the structure carries.

Interview prompts

Companion reads: Lesson 14 · The training recipe (what the recipe asks of data), Lesson 15 · Evaluation you can afford (the other bill) and Valuation · 09 Sensitivity and margin of safety (one assumption at a time; in Chinese).