Spending the budget: the allocation problem
Lesson 20 left a stock of corrections, a fleet and the sources of lessons 16 to 19, each with a price, a curve and a shelf life, and one question: given a budget, how much of each, and in what order? This lesson answers it for a programme that must generalise, recover and make contact, so that its success is the product of three. Spending by the marginal value per dollar of that product gives a sequence and not a mixture: the simulator's layouts first, gust demonstrations next, the twin's layouts after. A force rig, a fixed cost that buys no success by itself, is what the rule cannot see: on a task a thousand times the Bench's the best plan buys it at $18,680, lesson 17's ranking by price at $31,794, and the rule item by item never. It still counts hours as hours.
New idea: spend by the marginal value per dollar of the log of the product of the capabilities, a step at a time, and price an instrument together with the hours it unlocks. The plan is a sequence, cheap breadth first, and the force rig enters it ahead of the dearer breadth that its price rank would buy first.
Forces next: Spending by marginal value per dollar gives a sequence and not a mixture: cheap breadth first, the scarce column when the cheap source saturates, and the instrument for that column before its price rank says so. The plan counts hours as if an hour were an hour of information. A hundred thousand recorded episodes can hold eight hundred distinct situations, a policy trained on its own samples loses its tail, and a bad batch looks like a good one. What is the effective size of a dataset?
1 · Three columns on one ledger
The programme of lessons 16 to 20 needs three kinds of success. It must generalise to layouts it has not seen (lessons 7 and 8), recover when a gust of 0.05 displaces it (the five-post course of lessons 1 to 3) and make contact (a peg with a millimetre of clearance, lesson 10), and it succeeds when all three do. Let sg, sr and sc be the three success rates, each measured on the Bench task of that name, and assume them independent (a gust demonstration may also teach layouts, and the plan does not credit it). The programme then succeeds with probability P = sg sr sc.
It starts with what the Bench's courses start with: 8 own layouts (sg = 21.6 %), 20 calm demonstrations (sr = 63.5 %) and force-less insertions, which plateau at 64.2 % however many there are, so P0 = 8.8 %. Each column has sources, each with a price per unit from lesson 17 (assumptions, not measurements; a unit is a layout, a demonstration or an insertion attempt):
| column | source | price of a unit | plans: units, dollars, success |
|---|---|---|---|
| generalise, from 21.6 % | simulator off by 1 %, joint labels | $0.48 a thousand layouts | 256 layouts: $0.12, 82.3 % |
| twin arm | $0.039 a layout | 128: $4.96, 93.6 %; 256: $9.92, 98.8 % | |
| recover, from 63.5 % | own demonstrations recorded under a gust of 0.10 | $0.319 a demonstration | 3: $0.96, 88.5 %; 10: $3.19, 95.5 %; 20: $6.39, 97.5 % |
| contact, from 64.2 % | force-bearing insertions | $0.363 an insertion, and the rig, $15,000 once | 1 insertion: $0.36, 82.5 %; 10: $3.63, 98.5 % (plus the rig) |
Each cell is a plan: a point (x dollars, s success), the cheapest way the table has to reach s with that source, and a column is the list of its plans. A plan of the programme is one plan from each column, and with three short lists the best one at a budget can be found by trying every combination: that exhaustive search is the referee below. A plan counts as better than a cheaper one only if it adds a point of success, since the table's cells are 200 or 1000 runs wide (without this rule the best plan below would gain at most 0.2 % in P).
Other sources appear on no column's list, each rejected by the cost to reach a level: the older arm, footage and more own layouts never reach a level for less than the twin or the simulator, and recording under the test's own gust, 0.05, never for less than under 0.10. Three rounds of corrections on the 20 calm demonstrations reach 96.0 % for $3.08 of supervision where ten gust demonstrations reach 95.5 % for $3.19, with 95 % intervals (92.3 to 98.0, 91.7 to 97.6) that almost coincide; corrections also need a pipeline, assumed here to cost two engineer-weeks at lesson 17's rate, $8,000. An instrument is worth its fixed cost only for a column that no cheaper source supplies.
Scale. The Bench is small: everything its columns can buy, apart from the rig, costs $19.94. Let K multiply every number of hours, prices and rates unchanged (lesson 18's device): a plan that buys n units at the Bench buys Kn at scale K and costs K times as much, except for a fixed cost. Lesson 7 measured about 5.8 times as many layouts for each extra number a layout carries, so four extra numbers give K = 5.84 = 1,132; this lesson uses K = 1,000 and shows K = 1. K is an assumption: the order of purchases and the logic carry to larger tasks, the dollars do not.
The question. How much of each? The columns saturate at scales that differ by a factor of 121,015: the simulator's catalogue costs $0.12 at K = 1 and the contact column $15,004, rig included. Lesson 17's answer, the cheapest hour first, reaches P = 61.8 % with $23,714 at K = 1,000, where the best plan reaches 77.4 %; lesson 18's is derived for one success curve, and these are three, multiplied.
2 · The objective, and the rule it implies
Maximising P is maximising ln P = ln sg + ln sr + ln sc. With xi the dollars spent on column i and Σ xi = B, let λi = (dsi/dxi) / si be the gain in ln si per dollar: lesson 18's marginal value per dollar, for the log. Moving a dollar from column a to column b changes ln P by λb − λa, so a plan is best when no move helps: every column bought has the same λ, and one not bought has a λ at zero dollars no higher. Since si divides the slope, a weak column has a larger λ than a strong one with the same slope. That is what "weakest first" means, a statement about λ and not about s. The sum of the levels would count a point in a column at 21.6 % the same as a point in one at 98 %, and the minimum ignores how cheaply a strong column could still improve; §4 computes both.
A column is a list of plans, not a curve, so λ is a slope between plans: a step. If a step buys less per dollar than the step after it, as when the first step pays a lump and the second does not, the two are one block, priced at their average; merge steps until λ falls along the column (the upper concave envelope of ln s against dollars). The rule: list the blocks of the three columns in order of λ and buy down the list; a block that does not fit closes its column, and what is left buys the best single step that fits. The list is the sequence, and a bigger budget cuts the same list further down. At a budget that ends on a block boundary the rule gives the best plan, because the version with divisible blocks is optimal for its budget and a boundary budget buys whole blocks; between boundaries it can lose, and §3 measures how much.
3 · The sequence at K = 1,000
The table lists the blocks of the three columns at K = 1,000 in the order the rule buys them, with the cost of each, its λ per $1,000 and the total so far. The simulator's six steps are one row.
| block | reaches | cost | λ per $1,000 | total |
|---|---|---|---|---|
| simulator, 256,000 layouts | generalise 82.3 % | $124 | 120.9 to 2.09 | $124 |
| gust demonstrations, to 3,000 | recover 88.5 % | $958 | 0.346 | $1,082 |
| to 5,000 | recover 91.0 % | $639 | 0.0436 | $1,721 |
| to 10,000 | recover 95.5 % | $1,597 | 0.0302 | $3,317 |
| twin, 128,000 layouts | generalise 93.6 % | $4,836 | 0.0266 | $8,153 |
| the rig and 10,000 insertions | contact 98.5 % | $18,631 | 0.0230 | $26,785 |
| twin, 256,000 layouts | generalise 98.8 % | $4,960 | 0.0109 | $31,745 |
| gust demonstrations, to 20,000 | recover 97.5 % | $3,193 | 0.0065 | $34,938 |
The twin's 128,000 layouts cost $4,960; the sequence charges $4,836 because they replace the simulator's $124: a plan holds one of the two (pools of two sources do not add, lesson 18). Against the exhaustive search the rule is exact at all 13 block boundaries; between them it loses 0.84 % of P on average over the widget's 41 budgets and at most 20.2 %, at $23,714, where the budget cuts through the rig's lump (at most 5.1 % without it).
Cheap breadth first. The simulator fills generalise from 21.6 % to 82.3 % for $124, 0.35 % of the $34,938 the whole list costs, and the first block and the last differ by a factor of 18,626. Weakest first, by λ. Generalise starts weakest, yet recover is bought before the twin, because the simulator has lifted generalise and λ divides by s. One list, cut by the budget. $3,317 buys the first nine blocks and $34,938 all 13. Contact, the second weakest column once the simulator has run, comes eleventh, because its first dollar is a lump of $15,000: weakest first fails for a column that starts with a lump (§5).
What K moves. Multiply K by ten and every block that buys only hours costs ten times as much and has a tenth of the λ, so their order is the same at every K (checked at K = 1, 10, 100, 1,000 and 10,000). The rig's $15,000 does not scale, so its block moves: last of 13 at K = 1 (λ 0.0286 per $1,000 against 6.49 for the last block of hours), 11th at K = 1,000, 8th of 14 at K = 10,000.
The widget
What to try. Leave the defaults: $23,714, K = 1,000, a rig at $15,000, a 1 % simulator. The best plan spends $21,949: generalise 82.3 % (the simulator's 256,000 layouts), recover 95.5 %, contact 98.5 % (the rig and 10,000 insertions), P = 77.4 %; the block order, the item-by-item rule and the price rank stop at 61.8 %. Keep this plan's shares and halve the budget: 42.2 % against the best plan's 58.6 %. At $17,783 the best plan has no rig (61.8 %); at $31,623 the block order is the best plan again (89.9 %) while the item-by-item rule is still 61.8 %. At K = 1, $23,714 buys everything (94.9 %) and the three thresholds are $15,004, $15,017 and $15,017.
4 · Four rules, rejected by computation
| budget, K = 1,000 | best plan | lesson 17's price rank | the shares right at $10,000 |
|---|---|---|---|
| $1,334 | 46.7 % | 46.7 % | 33.5 % |
| $10,000 | 57.4 % | 57.4 % | 57.4 % |
| $23,714 | 77.4 % | 61.8 % | 61.8 % |
| $42,170 | 94.9 % | 94.9 % | 61.8 % |
A fixed mixture gives each column a fixed share of whatever the budget is. The best plan at $10,000 spends 61 % on generalise, 39 % on recover and nothing on contact, and a mixture with those shares is that plan there. It is wrong on both sides: at $1,334 its recover share, $522, is below the $639 of the first gust plan and buys nothing; at $42,170 its contact share is nothing. The shares of the best plan itself fall from 100 % on generalise at $100 to 0.6 % a little below $20,000, and the best of 231 fixed share vectors on a 5-point grid, chosen to lose least at its worst of 41 budgets from $10 to $1,000,000, still keeps only 66.5 % of the best plan's P there. A mixture is the sequence read at one budget.
Lesson 17's ranking buys the cheapest hour first, each source to the end of its catalogue: $124 of simulator, $9,920 of twin, $6,387 of gust demonstrations, $16,431 in all, which fills every column but contact. At $23,714 it has $7,283 left, short of the $15,363 that the cheapest plan with a rig costs. It is right at $1,334 and $10,000, where the cheapest hours are the best ones, and wrong where a column needs a lump.
The sum and the minimum. Maximising sg + sr + sc reaches the same P as the product's plan at all 41 budgets with the 1 % simulator, because here the weakest column is also the cheapest to lift. It loses where the weakest column has no cheap source: with the 3 % simulator, which sells nothing (lesson 17), $1,000 buys 21.0 % by the product and 19.2 % by the sum. The minimum fails more plainly. At $10,000 its largest value is 64.2 %, the contact plateau, which nothing below the rig can raise, so 14 plans tie on it with P from 36.7 % to 57.4 %; the cheapest costs $701 and stops at 36.7 %, where the product's plan spends $8,153 and reaches 57.4 %.
5 · The instrument: a fixed cost
The contact column's list begins with a jump. Lesson 17 priced a force hour at $214.4 by writing the rig off over the arm's 4,000 hours, $3.75 of every hour and $16.67 of every hour of force data, which is right only for a programme that keeps the rig busy that long. At K = 1,000 the best plan records 10,000 insertions, 18.4 hours, and cannot write off what it will not use: it pays $15,000 once, and the hour costs $197.8. Contact's first plan is the rig and 1,000 insertions, $15,363 for 82.5 %; its second is 10,000 insertions, $18,631 for 98.5 %. The first step adds 0.251 to ln sc for $15,363 (λ = 0.0164 per $1,000), the second 0.177 for $3,268 (λ = 0.0542): λ rises, so they are one block.
The rule one step ahead cannot buy it. A step of force-bearing hours is worth nothing without the rig, and the rig nothing without hours: valued item by item, as they stand, both have λ = 0 and the contact column is never started. This is the rule of §2 applied to what can be bought without an instrument. Once the other columns are full, at $16,307 and 61.8 %, more money buys nothing: at $100,000 the plan is still 61.8 %.
Priced as a block, it enters the list. The rig and 10,000 insertions cost $18,631 and raise ln sc by 0.429: λ = 0.0230 per $1,000. That is below the twin's first block (0.0266) and above its second (0.0109) and the last gust demonstrations (0.0065), the eleventh block of the sequence, and the rule that buys blocks reaches it at $26,785. What decides it is the average value per dollar of the whole block against the marginal value per dollar of the best alternative left.
The best plan buys it sooner. Without the rig the best plan runs out of things to buy at $16,307, with 61.8 %, and every dollar above that is idle. The best plan with the rig is the cheapest one that beats it: the simulator's layouts, 10,000 gust demonstrations, the rig and 1,000 insertions, $18,680 for 64.8 %. It holds no twin. The twin's layouts, which the best plan has held since $5,918, pay for the rig, and come back at $24,549. A plan for a bigger budget does not contain the plan for a smaller one, here because of a lump (the simulator is likewise dropped when the twin arrives): the sequence is the order in which a budget is spent, not a history of spending.
| task size | best plan | block order | lesson 17's price rank | item by item |
|---|---|---|---|---|
| K = 1 | $15,004 | $15,017 | $15,017 | never |
| K = 1,000 | $18,680 | $26,785 | $31,794 | never |
| K = 10,000 | $26,258 | $29,451 | $182,937 | never |
The table gives the budget at which each rule first buys force-bearing hours. At K = 1 the first three start within $13 of each other: everything else costs $19.94 and the rig is 99.9 % of the bill, so the Bench cannot show the effect. At K = 1,000 the best plan starts 1.7 times earlier than the price rank, at K = 10,000 7.0 times. The price rank orders by price per hour, the force hour is the dearest ($197.8 against $123.6, $15 and $0.19), and it buys the twin's $9,920 first; the best plan compares whole plans and takes the breadth that is nearly free, the simulator, in place of the breadth that is not, the twin, until the rig is paid for.
6 · What the plan assumed
Every plan above reads a curve of success against hours at the count of hours bought, and each curve was measured on hours that were all different: the layouts of the table are drawn afresh and a demonstration is a new run. A purchase is worth what the curve says only if each unit of it is new information. Suppose only a share q of what the plan buys is new, and read each curve at q times the count. The plan of $34,938 that promises P = 94.9 % (the twin's 256,000 layouts, 20,000 gust demonstrations, the rig and 10,000 insertions) delivers 73.7 % at q = ½ and 61.6 % at q = ¼ (generalise 82.0 %, recover 91.0 %, contact 82.5 %): 33.3 points of its promise are gone, and the plan cannot know, because its axis is dollars and hours.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
A budget is spent as a sequence: cheap breadth first (the simulator's 256,000 layouts for $124), the scarce column when the cheap source has run out of catalogue (10,000 gust demonstrations, the rig and 1,000 insertions at $18,680), and the instrument for that column before lesson 17's price rank would buy it ($31,794). The plan counts hours as if an hour were an hour of information: if only a quarter of what it bought is new, it delivers 61.6 % where it promised 94.9 %, and it cannot tell. A hundred thousand recorded episodes can hold eight hundred distinct situations, a policy trained on its own samples loses its tail, and a bad batch looks like a good one. What is the effective size of a dataset?
Interview prompts
- Why maximise the log of the product, and what does "weakest first" mean? (§2 — ln P is a sum, and λ = (ds/dx)/s is larger for a weak column with the same slope; it is a statement about λ, not about s.)
- Why is the plan a sequence and not a mixture? (§3, §4 — the order of blocks by λ does not depend on the budget, while the shares of the best plan go from 100 % on generalise to 0.6 %.)
- Why does a rule that values each step as it stands never buy an instrument? (§5 — the rig and the hours behind it both have λ = 0 until the other is held, so P stays at 61.8 %.)
- What does ranking sources by price get wrong? (§4 — it buys the cheap columns whole: at $23,714 it stops at 61.8 % with $7,283 left, short of a rig.)
- How do you decide whether to buy an instrument? (§5 — compare whole plans: the cheapest plan with the rig that beats the best plan without it, $18,680 for 64.8 % against 61.8 %.)
- What does the plan assume about an hour? (§6 — that each is new information: at q = ¼ the plan that promises 94.9 % delivers 61.6 %.)
Companion reads: Lesson 10 · What cameras cannot see (the force column), Lesson 7 · How much data: the coverage law (the factor of 5.8 behind K) and Financials · 22 Cyclicals and heavy assets (fixed costs and utilisation; in Chinese).