all_lessons/Robot Model Training/21 · Spending the budget: the allocation problemlesson 21 / 24

Spending the budget: the allocation problem

Lesson 20 left a stock of corrections, a fleet and the sources of lessons 16 to 19, each with a price, a curve and a shelf life, and one question: given a budget, how much of each, and in what order? This lesson answers it for a programme that must generalise, recover and make contact, so that its success is the product of three. Spending by the marginal value per dollar of that product gives a sequence and not a mixture: the simulator's layouts first, gust demonstrations next, the twin's layouts after. A force rig, a fixed cost that buys no success by itself, is what the rule cannot see: on a task a thousand times the Bench's the best plan buys it at $18,680, lesson 17's ranking by price at $31,794, and the rule item by item never. It still counts hours as hours.

The thesis, here
A programme that must generalise, recover and make contact succeeds when all three do, so the log of its success is a sum of three logs, and a budget is spent well when the next dollar adds the same to that sum wherever it goes: the weakest capability first, the cheapest step first. Run on the ledger this is a sequence whose order does not depend on the budget, so no fixed mixture is right across budgets. A fixed cost, the force rig, is what the rule cannot see, because it buys nothing alone: priced with the hours it unlocks, it is bought before lesson 17's ranking by price would buy it.
Linear position
Forced by: A fleet pays for its own supervisors only above the success rate at which a takeover costs less than the task it saves; below that rate every point the fleet gains is paid for by the programme, and above it the failures that supply the corrections grow rarer as the policy improves. We now have sources with falling returns, a stock of corrections that keeps its value while its supply runs out, and a fleet that has to be funded until it takes off, and each has a price, a value curve and a shelf life. Given a budget, how much of each do we buy, and in what order?
New idea: spend by the marginal value per dollar of the log of the product of the capabilities, a step at a time, and price an instrument together with the hours it unlocks. The plan is a sequence, cheap breadth first, and the force rig enters it ahead of the dearer breadth that its price rank would buy first.
Forces next: Spending by marginal value per dollar gives a sequence and not a mixture: cheap breadth first, the scarce column when the cheap source saturates, and the instrument for that column before its price rank says so. The plan counts hours as if an hour were an hour of information. A hundred thousand recorded episodes can hold eight hundred distinct situations, a policy trained on its own samples loses its tail, and a bad batch looks like a good one. What is the effective size of a dataset?
The plan
Six moves. (1) Put the three capabilities, their sources and a task size K on one ledger. (2) Choose the objective, the log of a product, and derive the rule it implies. (3) Run the rule: the order of purchases. (4) Reject the fixed mixture, the ranking by price, the sum and the minimum by computing each. (5) Add the instrument: a fixed cost defeats a rule that looks one step ahead, and a block restores it. (6) Ask what the plan assumed about an hour.

1 · Three columns on one ledger

The programme of lessons 16 to 20 needs three kinds of success. It must generalise to layouts it has not seen (lessons 7 and 8), recover when a gust of 0.05 displaces it (the five-post course of lessons 1 to 3) and make contact (a peg with a millimetre of clearance, lesson 10), and it succeeds when all three do. Let sg, sr and sc be the three success rates, each measured on the Bench task of that name, and assume them independent (a gust demonstration may also teach layouts, and the plan does not credit it). The programme then succeeds with probability P = sg sr sc.

It starts with what the Bench's courses start with: 8 own layouts (sg = 21.6 %), 20 calm demonstrations (sr = 63.5 %) and force-less insertions, which plateau at 64.2 % however many there are, so P0 = 8.8 %. Each column has sources, each with a price per unit from lesson 17 (assumptions, not measurements; a unit is a layout, a demonstration or an insertion attempt):

columnsourceprice of a unitplans: units, dollars, success
generalise, from 21.6 %simulator off by 1 %, joint labels$0.48 a thousand layouts256 layouts: $0.12, 82.3 %
twin arm$0.039 a layout128: $4.96, 93.6 %; 256: $9.92, 98.8 %
recover, from 63.5 %own demonstrations recorded under a gust of 0.10$0.319 a demonstration3: $0.96, 88.5 %; 10: $3.19, 95.5 %; 20: $6.39, 97.5 %
contact, from 64.2 %force-bearing insertions$0.363 an insertion, and the rig, $15,000 once1 insertion: $0.36, 82.5 %; 10: $3.63, 98.5 % (plus the rig)

Each cell is a plan: a point (x dollars, s success), the cheapest way the table has to reach s with that source, and a column is the list of its plans. A plan of the programme is one plan from each column, and with three short lists the best one at a budget can be found by trying every combination: that exhaustive search is the referee below. A plan counts as better than a cheaper one only if it adds a point of success, since the table's cells are 200 or 1000 runs wide (without this rule the best plan below would gain at most 0.2 % in P).

Other sources appear on no column's list, each rejected by the cost to reach a level: the older arm, footage and more own layouts never reach a level for less than the twin or the simulator, and recording under the test's own gust, 0.05, never for less than under 0.10. Three rounds of corrections on the 20 calm demonstrations reach 96.0 % for $3.08 of supervision where ten gust demonstrations reach 95.5 % for $3.19, with 95 % intervals (92.3 to 98.0, 91.7 to 97.6) that almost coincide; corrections also need a pipeline, assumed here to cost two engineer-weeks at lesson 17's rate, $8,000. An instrument is worth its fixed cost only for a column that no cheaper source supplies.

Scale. The Bench is small: everything its columns can buy, apart from the rig, costs $19.94. Let K multiply every number of hours, prices and rates unchanged (lesson 18's device): a plan that buys n units at the Bench buys Kn at scale K and costs K times as much, except for a fixed cost. Lesson 7 measured about 5.8 times as many layouts for each extra number a layout carries, so four extra numbers give K = 5.84 = 1,132; this lesson uses K = 1,000 and shows K = 1. K is an assumption: the order of purchases and the logic carry to larger tasks, the dollars do not.

The question. How much of each? The columns saturate at scales that differ by a factor of 121,015: the simulator's catalogue costs $0.12 at K = 1 and the contact column $15,004, rig included. Lesson 17's answer, the cheapest hour first, reaches P = 61.8 % with $23,714 at K = 1,000, where the best plan reaches 77.4 %; lesson 18's is derived for one success curve, and these are three, multiplied.

2 · The objective, and the rule it implies

Maximising P is maximising ln P = ln sg + ln sr + ln sc. With xi the dollars spent on column i and Σ xi = B, let λi = (dsi/dxi) / si be the gain in ln si per dollar: lesson 18's marginal value per dollar, for the log. Moving a dollar from column a to column b changes ln P by λb − λa, so a plan is best when no move helps: every column bought has the same λ, and one not bought has a λ at zero dollars no higher. Since si divides the slope, a weak column has a larger λ than a strong one with the same slope. That is what "weakest first" means, a statement about λ and not about s. The sum of the levels would count a point in a column at 21.6 % the same as a point in one at 98 %, and the minimum ignores how cheaply a strong column could still improve; §4 computes both.

A column is a list of plans, not a curve, so λ is a slope between plans: a step. If a step buys less per dollar than the step after it, as when the first step pays a lump and the second does not, the two are one block, priced at their average; merge steps until λ falls along the column (the upper concave envelope of ln s against dollars). The rule: list the blocks of the three columns in order of λ and buy down the list; a block that does not fit closes its column, and what is left buys the best single step that fits. The list is the sequence, and a bigger budget cuts the same list further down. At a budget that ends on a block boundary the rule gives the best plan, because the version with divisible blocks is optimal for its budget and a boundary budget buys whole blocks; between boundaries it can lose, and §3 measures how much.

3 · The sequence at K = 1,000

The table lists the blocks of the three columns at K = 1,000 in the order the rule buys them, with the cost of each, its λ per $1,000 and the total so far. The simulator's six steps are one row.

blockreachescostλ per $1,000total
simulator, 256,000 layoutsgeneralise 82.3 %$124120.9 to 2.09$124
gust demonstrations, to 3,000recover 88.5 %$9580.346$1,082
to 5,000recover 91.0 %$6390.0436$1,721
to 10,000recover 95.5 %$1,5970.0302$3,317
twin, 128,000 layoutsgeneralise 93.6 %$4,8360.0266$8,153
the rig and 10,000 insertionscontact 98.5 %$18,6310.0230$26,785
twin, 256,000 layoutsgeneralise 98.8 %$4,9600.0109$31,745
gust demonstrations, to 20,000recover 97.5 %$3,1930.0065$34,938

The twin's 128,000 layouts cost $4,960; the sequence charges $4,836 because they replace the simulator's $124: a plan holds one of the two (pools of two sources do not add, lesson 18). Against the exhaustive search the rule is exact at all 13 block boundaries; between them it loses 0.84 % of P on average over the widget's 41 budgets and at most 20.2 %, at $23,714, where the budget cuts through the rig's lump (at most 5.1 % without it).

Cheap breadth first. The simulator fills generalise from 21.6 % to 82.3 % for $124, 0.35 % of the $34,938 the whole list costs, and the first block and the last differ by a factor of 18,626. Weakest first, by λ. Generalise starts weakest, yet recover is bought before the twin, because the simulator has lifted generalise and λ divides by s. One list, cut by the budget. $3,317 buys the first nine blocks and $34,938 all 13. Contact, the second weakest column once the simulator has run, comes eleventh, because its first dollar is a lump of $15,000: weakest first fails for a column that starts with a lump (§5).

What K moves. Multiply K by ten and every block that buys only hours costs ten times as much and has a tenth of the λ, so their order is the same at every K (checked at K = 1, 10, 100, 1,000 and 10,000). The rig's $15,000 does not scale, so its block moves: last of 13 at K = 1 (λ 0.0286 per $1,000 against 6.49 for the last block of hours), 11th at K = 1,000, 8th of 14 at K = 10,000.

The widget

Spend a budget on three capabilities, and see who buys the rig
Left: success P against the budget (log axis) for the best plan, an exhaustive search over the plans of the three columns (teal), the block order of §3 (purple, dashed), the rule of §2 with no rig on offer, "item by item" (red), lesson 17's price rank (green, dotted) and a mixture keeping the shares of the best plan at the amber budget (cyan, dotted). Right: what the best plan buys and the sequence of §3, a square filled where the plan holds the block. The last three readouts: where each rule first buys force.
the best plan spends
-
generalise
-
recover
-
contact
-
success of the plan, P
-
block order (§3)
-
item by item, no rig on offer
-
lesson 17's price rank
-
maximising the sum
-
this plan's shares, half the budget
-
this plan's shares, twice the budget
-
rig bought from: best plan
-
rig bought from: block order
-
rig bought from: price rank
-
Show the core JS
AL.best = function (m, B, f) {
  var A = m.cols[0].pts, Bp = m.cols[1].pts, Cp = m.cols[2].pts, bestV = -Infinity, bi = [0, 0, 0], i, j, k, v, sum = f === 'sum';
  for (i = 0; i < A.length && A[i].x <= B + 1e-9; i++) for (j = 0; j < Bp.length && A[i].x + Bp[j].x <= B + 1e-9; j++) for (k = 0; k < Cp.length && A[i].x + Bp[j].x + Cp[k].x <= B + 1e-9; k++) {
    v = sum ? A[i].s + Bp[j].s + Cp[k].s : Math.log(A[i].s * Bp[j].s * Cp[k].s);
    if (v > bestV + 1e-12) { bestV = v; bi = [i, j, k]; }
  }
  return AL.planOf(m, bi);
};
...
function hullOf(pts) {
  var h = [], i, g = function (p) { return Math.log(p.s); };
  for (i = 0; i < pts.length; i++) {
    while (h.length > 1) {
      var o = h[h.length - 2], a = h[h.length - 1], p = pts[i];
      if ((a.x - o.x) * (g(p) - g(o)) - (g(a) - g(o)) * (p.x - o.x) >= 0) h.pop(); else break;
    }
    h.push(pts[i]);
  }
  return h;
}

What to try. Leave the defaults: $23,714, K = 1,000, a rig at $15,000, a 1 % simulator. The best plan spends $21,949: generalise 82.3 % (the simulator's 256,000 layouts), recover 95.5 %, contact 98.5 % (the rig and 10,000 insertions), P = 77.4 %; the block order, the item-by-item rule and the price rank stop at 61.8 %. Keep this plan's shares and halve the budget: 42.2 % against the best plan's 58.6 %. At $17,783 the best plan has no rig (61.8 %); at $31,623 the block order is the best plan again (89.9 %) while the item-by-item rule is still 61.8 %. At K = 1, $23,714 buys everything (94.9 %) and the three thresholds are $15,004, $15,017 and $15,017.

4 · Four rules, rejected by computation

budget, K = 1,000best planlesson 17's price rankthe shares right at $10,000
$1,33446.7 %46.7 %33.5 %
$10,00057.4 %57.4 %57.4 %
$23,71477.4 %61.8 %61.8 %
$42,17094.9 %94.9 %61.8 %

A fixed mixture gives each column a fixed share of whatever the budget is. The best plan at $10,000 spends 61 % on generalise, 39 % on recover and nothing on contact, and a mixture with those shares is that plan there. It is wrong on both sides: at $1,334 its recover share, $522, is below the $639 of the first gust plan and buys nothing; at $42,170 its contact share is nothing. The shares of the best plan itself fall from 100 % on generalise at $100 to 0.6 % a little below $20,000, and the best of 231 fixed share vectors on a 5-point grid, chosen to lose least at its worst of 41 budgets from $10 to $1,000,000, still keeps only 66.5 % of the best plan's P there. A mixture is the sequence read at one budget.

Lesson 17's ranking buys the cheapest hour first, each source to the end of its catalogue: $124 of simulator, $9,920 of twin, $6,387 of gust demonstrations, $16,431 in all, which fills every column but contact. At $23,714 it has $7,283 left, short of the $15,363 that the cheapest plan with a rig costs. It is right at $1,334 and $10,000, where the cheapest hours are the best ones, and wrong where a column needs a lump.

The sum and the minimum. Maximising sg + sr + sc reaches the same P as the product's plan at all 41 budgets with the 1 % simulator, because here the weakest column is also the cheapest to lift. It loses where the weakest column has no cheap source: with the 3 % simulator, which sells nothing (lesson 17), $1,000 buys 21.0 % by the product and 19.2 % by the sum. The minimum fails more plainly. At $10,000 its largest value is 64.2 %, the contact plateau, which nothing below the rig can raise, so 14 plans tie on it with P from 36.7 % to 57.4 %; the cheapest costs $701 and stops at 36.7 %, where the product's plan spends $8,153 and reaches 57.4 %.

5 · The instrument: a fixed cost

The contact column's list begins with a jump. Lesson 17 priced a force hour at $214.4 by writing the rig off over the arm's 4,000 hours, $3.75 of every hour and $16.67 of every hour of force data, which is right only for a programme that keeps the rig busy that long. At K = 1,000 the best plan records 10,000 insertions, 18.4 hours, and cannot write off what it will not use: it pays $15,000 once, and the hour costs $197.8. Contact's first plan is the rig and 1,000 insertions, $15,363 for 82.5 %; its second is 10,000 insertions, $18,631 for 98.5 %. The first step adds 0.251 to ln sc for $15,363 (λ = 0.0164 per $1,000), the second 0.177 for $3,268 (λ = 0.0542): λ rises, so they are one block.

The rule one step ahead cannot buy it. A step of force-bearing hours is worth nothing without the rig, and the rig nothing without hours: valued item by item, as they stand, both have λ = 0 and the contact column is never started. This is the rule of §2 applied to what can be bought without an instrument. Once the other columns are full, at $16,307 and 61.8 %, more money buys nothing: at $100,000 the plan is still 61.8 %.

Priced as a block, it enters the list. The rig and 10,000 insertions cost $18,631 and raise ln sc by 0.429: λ = 0.0230 per $1,000. That is below the twin's first block (0.0266) and above its second (0.0109) and the last gust demonstrations (0.0065), the eleventh block of the sequence, and the rule that buys blocks reaches it at $26,785. What decides it is the average value per dollar of the whole block against the marginal value per dollar of the best alternative left.

The best plan buys it sooner. Without the rig the best plan runs out of things to buy at $16,307, with 61.8 %, and every dollar above that is idle. The best plan with the rig is the cheapest one that beats it: the simulator's layouts, 10,000 gust demonstrations, the rig and 1,000 insertions, $18,680 for 64.8 %. It holds no twin. The twin's layouts, which the best plan has held since $5,918, pay for the rig, and come back at $24,549. A plan for a bigger budget does not contain the plan for a smaller one, here because of a lump (the simulator is likewise dropped when the twin arrives): the sequence is the order in which a budget is spent, not a history of spending.

task sizebest planblock orderlesson 17's price rankitem by item
K = 1$15,004$15,017$15,017never
K = 1,000$18,680$26,785$31,794never
K = 10,000$26,258$29,451$182,937never

The table gives the budget at which each rule first buys force-bearing hours. At K = 1 the first three start within $13 of each other: everything else costs $19.94 and the rig is 99.9 % of the bill, so the Bench cannot show the effect. At K = 1,000 the best plan starts 1.7 times earlier than the price rank, at K = 10,000 7.0 times. The price rank orders by price per hour, the force hour is the dearest ($197.8 against $123.6, $15 and $0.19), and it buys the twin's $9,920 first; the best plan compares whole plans and takes the breadth that is nearly free, the simulator, in place of the breadth that is not, the twin, until the rig is paid for.

Road not taken · amortise the rig into the hour
Lesson 17 already priced the rig, $16.67 of every hour of force data, so the instrument could stay inside the price and the rule of §2 apply unchanged. That is right for a programme that uses the rig for the arm's whole life, which at this Bench is K = 49,017. At K = 1,000 the programme records 18.4 hours and the amortised charge is $306, 49 times too small. Planned with that price the full plan costs $20,244 and promises P = 94.9 %; the bill is $34,938, and $20,244 actually buys 71.7 %.
What this lesson did not do
Each capability is one Bench task and the table holds one draw of each (95 % intervals from ±0.02 to ±0.07), so thresholds move with the draw; what carries is the order and the logic, not the dollars. K scales hours and not fixed costs, and a larger task would change the curves' shape too. The capabilities are independent and fed only by their own sources, and a column's plan uses one source, so the simulator's layouts are dropped when the twin's arrive (lesson 18). The rig and the pipeline are this lesson's assumptions. The plan is for one retraining: demonstrations and force recordings are capital and carry over, the stock of corrections keeps its value but its supply runs out (lesson 19), and a fleet's corrections are not a source here (lesson 20). What an hour is worth when many repeat one another is §6 and lesson 22; what moving it costs is lesson 23.

6 · What the plan assumed

Every plan above reads a curve of success against hours at the count of hours bought, and each curve was measured on hours that were all different: the layouts of the table are drawn afresh and a demonstration is a new run. A purchase is worth what the curve says only if each unit of it is new information. Suppose only a share q of what the plan buys is new, and read each curve at q times the count. The plan of $34,938 that promises P = 94.9 % (the twin's 256,000 layouts, 20,000 gust demonstrations, the rig and 10,000 insertions) delivers 73.7 % at q = ½ and 61.6 % at q = ¼ (generalise 82.0 %, recover 91.0 %, contact 82.5 %): 33.3 points of its promise are gone, and the plan cannot know, because its axis is dollars and hours.

Common mistakes / failure modes

"buy the cheapest hour first"
At $23,714 the price rank reaches 61.8 % where the best plan reaches 77.4 %: it fills the cheap columns and has $7,283 left, short of a rig (§4).
"split the budget in fixed proportions"
The shares exactly right at $10,000 reach 61.8 % at $42,170 against 94.9 %; the best fixed shares keep 66.5 % at their worst budget (§4).
"the Bench's dollars say when to buy the rig"
At K = 1 the rules start the rig within $13 of each other; at K = 1,000 the best plan is 1.7 times earlier than the price rank (§3, §5).
"the marginal rule will find the instrument"
Item by item the rig and its hours both have λ = 0, so P stays at 61.8 % at $100,000, where the best plan reaches 94.9 % (§5).

Checkpoint exercise

Try it
Two columns multiply. Column A starts at s = 0.40 and has two plans: $20 reaches 0.60, $50 reaches 0.80. Column B starts at 0.50 and has one: a fixed cost of $60 and then $20 of hours, $80 in all, reaches 0.90. (a) What is λ for each step of A and for B's block? (b) What does the rule reach at $100? (c) What is the best plan at $100? Answer: (a) λA1 = ln(0.60/0.40)/20 = 0.0203, λA2 = ln(0.80/0.60)/30 = 0.0096, and B's block adds ln(0.90/0.50) = 0.588 for $80: 0.0073 a dollar, so B is last. (b) The rule buys A1 and A2 ($50) and B's $80 does not fit: P = 0.80 × 0.50 = 0.40, which is also what the rule item by item reaches. (c) A1 and B cost exactly $100: P = 0.60 × 0.90 = 0.54. The best plan buys the instrument and drops A2, whose λ is higher, because B's block adds 0.588 to ln P where A2 adds 0.288 and $100 cannot hold both.

Where this points next

A budget is spent as a sequence: cheap breadth first (the simulator's 256,000 layouts for $124), the scarce column when the cheap source has run out of catalogue (10,000 gust demonstrations, the rig and 1,000 insertions at $18,680), and the instrument for that column before lesson 17's price rank would buy it ($31,794). The plan counts hours as if an hour were an hour of information: if only a quarter of what it bought is new, it delivers 61.6 % where it promised 94.9 %, and it cannot tell. A hundred thousand recorded episodes can hold eight hundred distinct situations, a policy trained on its own samples loses its tail, and a bad batch looks like a good one. What is the effective size of a dataset?

Takeaway
The programme succeeds when generalise, recover and contact all do, so the objective is ln P, a sum, and a budget is spent well where the next dollar adds the same to it in every column bought: the weakest column first, by λ = (ds/dx)/s. On the ledger at K = 1,000 that is a sequence, the simulator's layouts for $124, then gust demonstrations, then the twin's, and the order of the blocks that buy only hours does not depend on K; a fixed mixture is right at one budget, and the ranking by price spends on the cheapest hours. A fixed cost defeats a rule that looks one step ahead: the rig buys nothing alone, so item by item it is never bought and P stays at 61.8 %. Priced with the hours it unlocks, the best plan buys it at $18,680, before the block order ($26,785) and the price rank ($31,794), and drops the twin's layouts to pay for it; at K = 1 the rules start it within $13 of each other, so the Bench cannot show this and the scale does. The plan counts hours as hours: if a quarter of what it bought is new, $34,938 delivers 61.6 % instead of 94.9 %.

Interview prompts

Companion reads: Lesson 10 · What cameras cannot see (the force column), Lesson 7 · How much data: the coverage law (the factor of 5.8 behind K) and Financials · 22 Cyclicals and heavy assets (fixed costs and utilisation; in Chinese).