all_lessons/Robot Model Training/index24 lessons · ~10h read

Training a Robot Model, from first principles

A world model says what would happen if the robot did something. A robot has to choose, about twenty times a second, and the cheapest teacher for a chooser is a person who already is one. This track builds the chooser one forced step at a time, on a bench small enough to run in your browser: a two-link arm carrying a cup past a row of vases. It copies a person's demonstrations, finds why a copy that is accurate where the person went is lost where it goes, and repairs that with labels on its own states, distributions, chunks and feedback inside the action. It then asks how much data a policy needs and of which kind (layouts, other bodies, video, force, simulation), improves a policy from a score alone, builds the model whose two clocks fit a robot, trains it in stages without erasing what it knew, and prices the evaluation that says whether any of it worked. The last nine lessons then put a unit and a price on everything the first fifteen asked for. They measure every source of data in one unit, the hours of your own robot's demonstrations that an hour of it replaces; price each; let the value of an hour fall as you buy; ask how long an hour keeps its value and whether a fleet of working robots pays for its supervision; spend a budget as a sequence of purchases; count effective hours instead of recorded ones; price moving the bits; and read the ledger to say what training a robot model eats the most.

The seed question
A person carries a cup past a row of vases without touching one. You record twenty of their runs and train a controller on them. On frames the person never produced it copies them to within a few percent. Run it on a course of five vases and it knocks one over in more than a third of its runs. What did the controller learn, what did it not, and what data would fix it?
The question the second half answers
You have a budget, one robot, and a choice among another laboratory's demonstrations on a different arm, a hundred thousand hours of video, a simulator, and a person watching your robot work. Which of them should the next dollar go to, and what would you have to know to say?
Who this is for
An engineer or researcher who has met the words (behaviour cloning, DAgger, action chunking, diffusion policy, cross-embodiment, vision-language-action models, residual reinforcement learning, sim-to-real) and wants to see why each exists, and who may then have to decide what data a robot programme buys. You need linear algebra, probability and the idea of training a model by gradient descent; the World Models, Reinforcement Learning and (for the last nine lessons) Financials tracks are good company and none is required. By the end of lesson 15 you can say why cloning is regression and why its score misleads; why a policy's own mistakes make its test set and cost the square of the horizon; what a label on the policy's own states buys and what it costs; why a regression returns the average of two routes and what to output instead; why a policy should commit to a chunk and what the chunk should be made of; how many layouts a policy needs and what other bodies, video, force and simulators can and cannot supply; what a score alone can teach; how a model with two clocks is built and trained; and how many trials it takes to know. By the end of lesson 24 you can say what an exchange rate between two data sources is and how to measure it; why the ranking by cost per recorded hour and the ranking by cost per useful hour disagree by orders of magnitude; how to buy under falling returns; which data keeps its value (capital, inventory) and what runs out (the supply of failures to correct); when a fleet becomes a flywheel; why a budget is spent as a sequence; how to count a dataset's effective size; and where the machine stalls on data it cannot read fast enough.
How this track is built
An original first-principles track, not a survey of papers. (1) A relay, not a list. Every lesson opens with the failure the previous one left standing and closes by stating the next; that hand-off is one sentence, the baton, which appears word for word at the end of one lesson, at the start of the next and in the table below, and a script diffs them. (2) One world. Everything happens on the Bench: a planar arm of two 0.5 m links, commanded in joint velocities twenty times a second through a noisy plant, a row of up to six vases, and an expert you can read. Because the world is small and exactly known, every claim is a computation you can watch, and a policy can be compared at any state with the expert. The value of data in the second half is measured on the same Bench, so an exchange rate is a computation and not an opinion. (3) The widgets compute. Each lesson has one widget that runs the mechanism the lesson derives, with real rollouts, real learners and real statistics, and a "What to try" paragraph that says which control to move first. (4) Every number is checked. Each quoted number is produced by an independent script, and the Bench is allowed to disagree with folklore: where it does, the lesson reports what it measured and says what it cannot show. Outside facts (what a paper reported, how big a dataset is) come from primary sources and are cited by name and year. (5) Prices are assumptions, and labelled as such. From lesson 17 on, each price enters through a visible table of assumptions that the widget lets you move. The Bench's exchange rates and curves teach the method and the shape of the answer (which way, which order of magnitude, where curves cross); real values depend on the task, the policy and the year, and the lessons say so.

The exam

The first fifteen lessons are graded by one question asked in different ways: when the robot runs the policy, does the task get done, and how would we know? A copy error, a training loss and a validation curve each measure something else, and the first lessons are about the gap between them. The later ones are the moves forced by what a policy needs in order to close it. The last nine are graded by one formula: value per dollar = marginal value × exchange rate ÷ price. Each of those lessons measures or derives one factor of it, or one reason the factors change as you buy.

Part I · Copy a person (lessons 01–03)

A controller from demonstrations, why it fails in the loop, and the labels that repair it.

01
Copy a person who knows: behaviour cloning
Why a cached policy beats planning at every step, what a demonstration is, a nearest-demo regressor, and the two numbers that disagree: copy error on held-out frames and success when the policy runs.
02
The policy writes its own test set
The deviation recursion, the funnel the data draw around the demonstrations, the chance of escaping it each step, the quadratic cost and its measured exponent, and what more data of the same kind does not change.
03
Labels on your own states: DAgger and corrections
DAgger on the slalom, why it beats more demonstrations and noise injection at the same label budget, and what happens when the labeller is a person who has to take over to be accurate.

Part II · What the data really is (lessons 04–06)

People are not functions, so policies output distributions; committing to sequences; and feedback built into the action itself.

04
People are not functions: distributions over actions
The average of two routes, why squared error returns it, nearest-demo sampling, a mixture head and a codebook of prototypes, and the dithering that sampling at every step still allows.
05
Commit: action chunks
Mode-hopping measured, chunk length against the number of decisions, and the open-loop cost: noise that velocity commands integrate over the length of the chunk.
06
Targets, not velocities: feedback in the action
Why a velocity output is an integrator and a tracked target is not, the loop gain of each, the units and resolution of an action, what stiffness costs, and why targets copy the demonstrations' world.

Part III · Where more data comes from (lessons 07–11)

How much is enough, other bodies, video without actions, the signal a camera lacks, and simulation.

07
How much data: the coverage law
Error as distance to the nearest demonstrated layout, power laws in the number of layouts, why repeating a layout buys almost nothing, and what ninety percent costs in operator-hours.
08
Other bodies: units, pooling, interference
Why joint-angle actions from another arm are wrong labels, end-effector targets with per-body inverse kinematics, the benefit of pooling and the interference when bodies differ.
09
Watching: learning from video without actions
An inverse-dynamics model from labelled motion, pseudo-labels at scale, how label error propagates, and what a change of state cannot reveal: force, effort, intent.
10
What cameras cannot see: contact and force
Insertion with a millimetre of clearance, the success ceiling of a position-only policy, compliance and force feedback, and why the force has to be in the data.
11
Simulation: unlimited data, wrong physics
A plant mismatch measured as lost success, randomisation width against caution, system identification from a few real rollouts, and the interior optimum between them.

Part IV · Beyond the demonstrator (lessons 12–15)

A score as the only teacher, the anatomy of the model, the training recipe, and the evaluation that pays for it all.

12
A score as the only teacher: reinforcement learning
Policy-gradient and evolution-strategy updates on the Bench, attempts needed from scratch against from a cloned prior, the residual formulation, and the sample arithmetic in robot-hours.
13
Anatomy of a robot model: two clocks
The latency budget of a control loop, what a vision-language backbone adds and costs, an action expert that emits chunks, and the staleness between the two clocks.
14
The recipe: stages, mixtures and forgetting
Pretrain, post-train, correct; the forgetting curve of fine-tuning on a few tasks; the co-training ratio that prevents it; how to notice the loss.
15
Evaluation you can afford
Wilson intervals for twenty trials, the minimum detectable difference, trials per arm, rank reversals between metrics, and the robot-weeks an honest evaluation costs.

Part V · Value and price (lessons 16–17)

Every source of data in one unit, then in dollars, and the two rankings that disagree.

16
An hour is not an hour: exchange rates
The Bench as a laboratory: an hour of another arm's data, of labelled video, of simulation or of corrections weighed against hours of your own, and the sources whose rate is zero.
17
What an hour costs: prices and the ledger
Operator-hours, resets and discards, compute and authoring cost; the cost per effective hour of every source; two rankings that disagree; the columns no cheap source supplies.

Part VI · How value moves (lessons 18–20)

Diminishing returns, a stock of corrections that keeps its value while its supply runs out, and a fleet that pays for its supervisors only above a success rate.

18
The thousandth hour: diminishing returns
Learning curves as power laws, marginal value, where each source saturates, and why the cheapest source changes as you buy.
19
Is data perishable? Shelf life and the supply of corrections
The relevance of old corrections to a retrained policy, measured across generations (it does not fall); where it does fall (the world moves outward); why the supply of corrections, not the stock, runs out; and what demonstrations and force recordings keep.
20
The flywheel: a fleet as a source
A fleet with supervisors on the Bench: data per day, value per correction, supervision cost, the take-off condition and the cost of each percentage point.

Part VII · Spending it (lessons 21–24)

The allocation problem, the effective size of a dataset, the bill for moving it, and the honest answer.

21
Spending the budget: the allocation problem
Capabilities that compose like a product, saturation scales that differ by orders of magnitude, the greedy sequence, and why the instrument for a column comes before its price rank.
22
Effective hours: duplicates, quality and collapse
Effective sample size, clusters of duplicates that are up-weighted and not merely wasted, tail coverage that decays as a dataset repeats, and training on your own samples.
23
Moving the bits: the infrastructure bill
Raw and stored bytes per training token, decode-bound pipelines, the cores needed to feed one accelerator, and utilisation against resolution and frame rate.
24
What does embodied training eat the most?
Both rankings, the structural zeros, the capital, inventory and corrections classification, the flywheel threshold, and a worked purchase plan.

The whole derivation on one page

Read the second column of any row and the fourth column of the row above it: they are the same sentence. That is the linear thinking made visible. Each lesson's problem it inherits is exactly the previous lesson's problem it creates, and the middle column is the single move that resolves the first and manufactures the second. The first row inherits its problem from the last lesson of World Models (lesson 31, the whole bill for training one), where a model that can say what would happen has never chosen anything; the last row asks which prices will move and which gaps will not.

#The problem it inheritsThe one moveThe problem it creates
Part I · Copy a person
01A trained world model answers what would happen if the robot did something, and it has never chosen anything: it has no goal, and the data it is tested on do not change when it is used. A robot has to choose, about twenty times a second, from what it has just sensed, and the cheapest teacher for a chooser is a person who already is one. What does a robot learn by copying a person doing the task, and how would we know it had learned it?Copy a person who knows: behaviour cloning: Turn demonstrations into a controller by regression, and measure it two ways: how well it copies, and whether the task gets done.A policy cloned from twenty demonstrations copies the expert to within a few percent on frames it has never seen, makes about four times that error on the frames it produces itself, and fails more than a third of the time on a course of five posts and more than half the time on six. The two kinds of frame come from different places: the first from where the expert went, the second from where the policy goes. Why is a policy that is accurate where the expert went lost where it goes, and how fast does the loss grow with the length of the task?
02A policy cloned from twenty demonstrations copies the expert to within a few percent on frames it has never seen, makes about four times that error on the frames it produces itself, and fails more than a third of the time on a course of five posts and more than half the time on six. The two kinds of frame come from different places: the first from where the expert went, the second from where the policy goes. Why is a policy that is accurate where the expert went lost where it goes, and how fast does the loss grow with the length of the task?The policy writes its own test set: Derive how a policy's own mistakes become its next input: a loop gain and an extrapolation error that make cost grow as the square of the horizon.A policy's mistakes land it in states its training data never showed, where its next action is a guess, so one slip is charged for every step that follows: the cost grows as the square of the horizon unless the loop pulls the policy back or the training data include the states the policy reaches. More demonstrations recorded the same way leave the curve where it is. The only one who knows what to do in a state the policy has just wandered into is the expert. Can the expert label the states the policy visits, and what does that cost?
03A policy's mistakes land it in states its training data never showed, where its next action is a guess, so one slip is charged for every step that follows: the cost grows as the square of the horizon unless the loop pulls the policy back or the training data include the states the policy reaches. More demonstrations recorded the same way leave the curve where it is. The only one who knows what to do in a state the policy has just wandered into is the expert. Can the expert label the states the policy visits, and what does that cost?Labels on your own states: DAgger and corrections: Put the states the policy reaches into the training set, labelled by the expert: aggregate, retrain, repeat.Training on the policy's own states, with the expert's actions as labels, takes a policy that finishes five posts about six times in ten to one that finishes them about nineteen times in twenty after four rounds of five rollouts, and a person supplies those labels best by taking over when the policy goes wrong. But people are not functions: asked from the same pose, or asked by two operators, one goes left of the post and another right, and a regression asked to fit both returns their average, a route that runs into the post. What should a policy output when the right action is not one point?
Part II · What the data really is
04Training on the policy's own states, with the expert's actions as labels, takes a policy that finishes five posts about six times in ten to one that finishes them about nineteen times in twenty after four rounds of five rollouts, and a person supplies those labels best by taking over when the policy goes wrong. But people are not functions: asked from the same pose, or asked by two operators, one goes left of the post and another right, and a regression asked to fit both returns their average, a route that runs into the post. What should a policy output when the right action is not one point?People are not functions: distributions over actions: Replace the single answer by a distribution over actions and sample it, so that two good routes stay two routes.A policy that outputs a distribution over actions, and acts by sampling it, recovers both routes instead of their collision. But it draws a new sample at every step, and near the fork each draw can pick a different route: the arm dithers between them and a quarter of the rollouts still end in a post. Committing to one route for a while would stop the dithering, at the price of not looking while committed. How long should a policy commit, and what should it commit to?
05A policy that outputs a distribution over actions, and acts by sampling it, recovers both routes instead of their collision. But it draws a new sample at every step, and near the fork each draw can pick a different route: the arm dithers between them and a quarter of the rollouts still end in a post. Committing to one route for a while would stop the dithering, at the price of not looking while committed. How long should a policy commit, and what should it commit to?Commit: action chunks: Sample a whole sequence of actions at once and execute it, so that the choice of route is made once.Sampling a whole sequence of actions at once and executing it commits the policy to one route, and the dithering stops. But the sequence is a list of velocities, a velocity command is added to the joint angles at every step, and every bit of noise in the plant stays added for the whole chunk: the longer the chunk, the further the arm drifts from where the sequence meant it to be, and the gain from committing is paid back as drift. What should the policy output, so that the plant's noise does not accumulate while it commits?
06Sampling a whole sequence of actions at once and executing it commits the policy to one route, and the dithering stops. But the sequence is a list of velocities, a velocity command is added to the joint angles at every step, and every bit of noise in the plant stays added for the whole chunk: the longer the chunk, the further the arm drifts from where the sequence meant it to be, and the gain from committing is paid back as drift. What should the policy output, so that the plant's noise does not accumulate while it commits?Targets, not velocities: feedback in the action: Make the output a target position tracked by a stiff controller, so that the plant's noise is corrected every step instead of accumulating.A chunk of target positions, tracked by a stiff controller that pulls the arm back to the target at every step, removes the accumulation: on the demonstrated layout the policy completes the course nearly every time at any chunk length. But targets are anchored to where the demonstrations were. Shift the vases by a few centimetres and the policy walks faithfully into them. The policy has to read where the vases are and generalise across layouts, and nobody has said how many layouts that takes. How does success on a new layout grow with the data, and which data buy it?
Part III · Where more data comes from
07A chunk of target positions, tracked by a stiff controller that pulls the arm back to the target at every step, removes the accumulation: on the demonstrated layout the policy completes the course nearly every time at any chunk length. But targets are anchored to where the demonstrations were. Shift the vases by a few centimetres and the policy walks faithfully into them. The policy has to read where the vases are and generalise across layouts, and nobody has said how many layouts that takes. How does success on a new layout grow with the data, and which data buy it?How much data: the coverage law: Measure how success on a new layout grows with the number of layouts and demonstrations, and derive the exponent from what a local learner can reach.For a policy that generalises only to what resembles its data, the error on a new layout is the distance to the nearest demonstrated layout, so success grows with the number of different layouts, with a power set by how many numbers describe a layout, and hardly at all with more demonstrations of a layout already seen. Ninety percent costs about a hundred layouts when two numbers describe a layout, and every further number multiplies that by the range it varies over divided by the width one layout covers; a real scene is described by many more than two numbers, and one arm cannot pay for that. Other laboratories have already collected such data, on other arms. What can a robot use from demonstrations made on a body that is not its own?
08For a policy that generalises only to what resembles its data, the error on a new layout is the distance to the nearest demonstrated layout, so success grows with the number of different layouts, with a power set by how many numbers describe a layout, and hardly at all with more demonstrations of a layout already seen. Ninety percent costs about a hundred layouts when two numbers describe a layout, and every further number multiplies that by the range it varies over divided by the width one layout covers; a real scene is described by many more than two numbers, and one arm cannot pay for that. Other laboratories have already collected such data, on other arms. What can a robot use from demonstrations made on a body that is not its own?Other bodies: units, pooling, interference: Pool demonstrations from other arms by writing the action in a space every arm shares, and let each body supply its own joint-level controller.Demonstrations from other bodies help when the action is written in a space every body shares, the place the hand should go, and each arm turns that into its own joint motions; written in joint angles they are worse than no data. Pooled this way, all the robot demonstrations ever recorded are a small fraction of what people have filmed themselves doing, and the video has no actions in it. What can be learned from watching, and what is missing from the video that no amount of it can supply?
09Demonstrations from other bodies help when the action is written in a space every body shares, the place the hand should go, and each arm turns that into its own joint motions; written in joint angles they are worse than no data. Pooled this way, all the robot demonstrations ever recorded are a small fraction of what people have filmed themselves doing, and the video has no actions in it. What can be learned from watching, and what is missing from the video that no amount of it can supply?Watching: learning from video without actions: Infer the missing actions from motion with a model trained on a few labelled hours, then train on the pseudo-labelled video.A model that infers the missing actions from a few hours of labelled motion can turn a large pile of video into demonstrations, and what it learns is what to do next, not how hard to push. Video records where things moved, never the force that moved them: an insertion with a millimetre of clearance, watched through a camera that cannot resolve a millimetre, looks the same when it succeeds and when it jams. What must a policy sense, and how must it act, when the task is decided by something the camera cannot see?
10A model that infers the missing actions from a few hours of labelled motion can turn a large pile of video into demonstrations, and what it learns is what to do next, not how hard to push. Video records where things moved, never the force that moved them: an insertion with a millimetre of clearance, watched through a camera that cannot resolve a millimetre, looks the same when it succeeds and when it jams. What must a policy sense, and how must it act, when the task is decided by something the camera cannot see?What cameras cannot see: contact and force: Give the policy the signal the task depends on, force, and an action that yields to it, instead of position alone.A policy that feels force and yields to it, instead of holding a position stiffly, inserts the peg where position control jams, and the force it needs is in the data only if someone recorded it with a force sensor in the loop, on a real robot, slowly, with wear. A simulator produces force, actions and unlimited demonstrations, at the price of a physics that is not quite the real one. How wrong can a simulator be before a policy trained in it fails on the robot, and what buys back the difference?
11A policy that feels force and yields to it, instead of holding a position stiffly, inserts the peg where position control jams, and the force it needs is in the data only if someone recorded it with a force sensor in the loop, on a real robot, slowly, with wear. A simulator produces force, actions and unlimited demonstrations, at the price of a physics that is not quite the real one. How wrong can a simulator be before a policy trained in it fails on the robot, and what buys back the difference?Simulation: unlimited data, wrong physics: Train in a simulator and close the gap with randomised physics and a few real rollouts for calibration.A simulator is worth using when its error is narrower than the policy can tolerate: randomising the physics widens what a policy survives at the price of a more cautious motion, and calibrating the simulator on a few real rollouts narrows the gap itself. Both assume there is someone in the simulator to copy, a scripted expert or a planner. For most tasks that matter there is no such expert, only a way to tell whether the task was done. What can be learned from a score alone, and how many attempts does it take?
Part IV · Beyond the demonstrator
12A simulator is worth using when its error is narrower than the policy can tolerate: randomising the physics widens what a policy survives at the price of a more cautious motion, and calibrating the simulator on a few real rollouts narrows the gap itself. Both assume there is someone in the simulator to copy, a scripted expert or a planner. For most tasks that matter there is no such expert, only a way to tell whether the task was done. What can be learned from a score alone, and how many attempts does it take?A score as the only teacher: reinforcement learning: Improve a policy from a reward alone: sample the policy, keep what scored, and start from a good prior so that few attempts are needed.Learning from a score alone works, slowly, on a small task, and works in an afternoon of robot time only when it starts from a policy that is already nearly right and learns a small correction. Every source so far feeds that starting point: demonstrations, other bodies, video, force, simulation. It is one model that must read language and images, feel force, emit chunks, and still answer within one control step. How is such a model built, and what does it cost to run it at the speed of the robot?
13Learning from a score alone works, slowly, on a small task, and works in an afternoon of robot time only when it starts from a policy that is already nearly right and learns a small correction. Every source so far feeds that starting point: demonstrations, other bodies, video, force, simulation. It is one model that must read language and images, feel force, emit chunks, and still answer within one control step. How is such a model built, and what does it cost to run it at the speed of the robot?Anatomy of a robot model: two clocks: Split the model into a slow reasoner and a fast action network, and decide what each must do within its own clock.A large model that reads and reasons is too slow to steer an arm directly, so it runs slowly and hands a plan to a small fast network that runs at the control rate, and the split decides what each part must be trained on. Training the parts in order, on data of very different size and quality, has its own failure: fine-tuning on a few dozen tasks raises every number we track while erasing what the language model knew. In what order, on what mixture, should the pieces be trained, and how do we notice the loss?
14A large model that reads and reasons is too slow to steer an arm directly, so it runs slowly and hands a plan to a small fast network that runs at the control rate, and the split decides what each part must be trained on. Training the parts in order, on data of very different size and quality, has its own failure: fine-tuning on a few dozen tasks raises every number we track while erasing what the language model knew. In what order, on what mixture, should the pieces be trained, and how do we notice the loss?The recipe: stages, mixtures and forgetting: Train in stages and keep old data in the mixture, so that what the model knew survives what it is taught.Staging the training and mixing the old data into the new keeps the language the model started with, at a measurable price in speed and in the tasks it learns. We now have a recipe and a checkpoint, and need to say whether the checkpoint is better than the last one. An evaluation is a count of trials, and each trial costs a robot and a person for minutes. How many trials does it take to tell one policy from another, and what does that do to the cost of the whole programme?
15Staging the training and mixing the old data into the new keeps the language the model started with, at a measurable price in speed and in the tasks it learns. We now have a recipe and a checkpoint, and need to say whether the checkpoint is better than the last one. An evaluation is a count of trials, and each trial costs a robot and a person for minutes. How many trials does it take to tell one policy from another, and what does that do to the cost of the whole programme?Evaluation you can afford: Treat the success rate as a count of trials: put an interval on it, size the experiment for the difference you care about, and price it.Twenty trials cannot tell seventy-six percent from sixty-eight, telling a ten-point improvement apart takes hundreds of trials per policy, and a programme that evaluates honestly spends more robot time on evaluation than on training. The programme also needs data of several kinds, each from a source with a different price, a different usefulness and a different shelf life. How much of each kind do we need, in what unit can they even be compared, and what does each cost?
Part V · Value and price
16Twenty trials cannot tell seventy-six percent from sixty-eight, telling a ten-point improvement apart takes hundreds of trials per policy, and a programme that evaluates honestly spends more robot time on evaluation than on training. The programme also needs data of several kinds, each from a source with a different price, a different usefulness and a different shelf life. How much of each kind do we need, in what unit can they even be compared, and what does each cost?An hour is not an hour: exchange rates: Measure every source in the same unit: how many hours of your own robot's demonstrations an hour of it replaces.Measured on the Bench, an hour of another arm's demonstrations replaces a fraction of an hour of your own, an hour of labelled video a smaller fraction, an hour of simulation a fraction that depends on the gap, and a source that lacks a column the task needs replaces none of it, however many hours there are. Exchange rates put every source in one unit, but not in dollars. What does an hour of each source cost?
17Measured on the Bench, an hour of another arm's demonstrations replaces a fraction of an hour of your own, an hour of labelled video a smaller fraction, an hour of simulation a fraction that depends on the gap, and a source that lacks a column the task needs replaces none of it, however many hours there are. Exchange rates put every source in one unit, but not in dollars. What does an hour of each source cost?What an hour costs: prices and the ledger: Build each source's price per hour from labour, throughput, compute and authoring, and divide it by the source's exchange rate.Prices and exchange rates together give a cost per useful hour for every source, the price of an hour divided by its rate, and the ranking they produce is not the ranking by price: footage, dearer to make than the other borrowed hours, becomes the dearest useful hour of any source that can teach the same task, a simulator is the cheapest useful hour only while its gap is small and sells no useful hour at any price once the gap is wide, and some columns cannot be bought from the sources that are cheap. But a cost per hour is the cost of the first hour. The thousandth hour of a source teaches the policy less than the first. How fast does the value of one more hour fall?
Part VI · How value moves
18Prices and exchange rates together give a cost per useful hour for every source, the price of an hour divided by its rate, and the ranking they produce is not the ranking by price: footage, dearer to make than the other borrowed hours, becomes the dearest useful hour of any source that can teach the same task, a simulator is the cheapest useful hour only while its gap is small and sells no useful hour at any price once the gap is wide, and some columns cannot be bought from the sources that are cheap. But a cost per hour is the cost of the first hour. The thousandth hour of a source teaches the policy less than the first. How fast does the value of one more hour fall?The thousandth hour: diminishing returns: Measure the value curve of each source and buy until the marginal value per dollar is equal across sources.The value of every source falls as a power of the hours bought, so the rule is to buy from each source until its marginal value per dollar equals that of the next, and which source is cheapest changes as you buy. This treats the hours as an asset that keeps its value once bought. Demonstrations do, for the policy trained next year as well as for the one trained today. Corrections were made in the states an older policy visited, and the policy trained next year will not visit quite the same ones. How long does a correction keep its value?
19The value of every source falls as a power of the hours bought, so the rule is to buy from each source until its marginal value per dollar equals that of the next, and which source is cheapest changes as you buy. This treats the hours as an asset that keeps its value once bought. Demonstrations do, for the policy trained next year as well as for the one trained today. Corrections were made in the states an older policy visited, and the policy trained next year will not visit quite the same ones. How long does a correction keep its value?Is data perishable? Shelf life and the supply of corrections: Classify data by how long it stays useful: capital that keeps its value, inventory that saturates, and corrections, whose value lasts as long as the policy stays among the states they were made in.A correction does not lose its value when the policy it was made for is retrained: a better policy stays among the states of a worse one, and on the Bench a set made for the first clone is worth as much per frame to the fourth clone as a fresh set, losing value only when the world moves outward, by about a fifth when the gust doubles. Demonstrations and force recordings, which the expert made and no policy did, keep theirs whoever is trained. What runs out is the supply: a correction teaches what its policy got wrong, the first set adds about eighteen points of success and the fourth about one, and a policy that fails one run in twenty has little left to show. The cheapest continuing supply of corrections is the deployed robots themselves, working, with a person stepping in when they fail. Does a fleet like that improve the policy fast enough to pay for the people who supervise it?
20A correction does not lose its value when the policy it was made for is retrained: a better policy stays among the states of a worse one, and on the Bench a set made for the first clone is worth as much per frame to the fourth clone as a fresh set, losing value only when the world moves outward, by about a fifth when the gust doubles. Demonstrations and force recordings, which the expert made and no policy did, keep theirs whoever is trained. What runs out is the supply: a correction teaches what its policy got wrong, the first set adds about eighteen points of success and the fourth about one, and a policy that fails one run in twenty has little left to show. The cheapest continuing supply of corrections is the deployed robots themselves, working, with a person stepping in when they fail. Does a fleet like that improve the policy fast enough to pay for the people who supervise it?The flywheel: a fleet as a source: Model deployment as a data source whose supply depends on the policy's own success, and find the success rate at which it pays for itself.A fleet pays for its own supervisors only above the success rate at which a takeover costs less than the task it saves; below that rate every point the fleet gains is paid for by the programme, and above it the failures that supply the corrections grow rarer as the policy improves. We now have sources with falling returns, a stock of corrections that keeps its value while its supply runs out, and a fleet that has to be funded until it takes off, and each has a price, a value curve and a shelf life. Given a budget, how much of each do we buy, and in what order?
Part VII · Spending it
21A fleet pays for its own supervisors only above the success rate at which a takeover costs less than the task it saves; below that rate every point the fleet gains is paid for by the programme, and above it the failures that supply the corrections grow rarer as the policy improves. We now have sources with falling returns, a stock of corrections that keeps its value while its supply runs out, and a fleet that has to be funded until it takes off, and each has a price, a value curve and a shelf life. Given a budget, how much of each do we buy, and in what order?Spending the budget: the allocation problem: Maximise the weakest capability under a budget by marginal value per dollar: a sequence of purchases, not a mixture.Spending by marginal value per dollar gives a sequence and not a mixture: cheap breadth first, the scarce column when the cheap source saturates, and the instrument for that column before its price rank says so. The plan counts hours as if an hour were an hour of information. A hundred thousand recorded episodes can hold eight hundred distinct situations, a policy trained on its own samples loses its tail, and a bad batch looks like a good one. What is the effective size of a dataset?
22Spending by marginal value per dollar gives a sequence and not a mixture: cheap breadth first, the scarce column when the cheap source saturates, and the instrument for that column before its price rank says so. The plan counts hours as if an hour were an hour of information. A hundred thousand recorded episodes can hold eight hundred distinct situations, a policy trained on its own samples loses its tail, and a bad batch looks like a good one. What is the effective size of a dataset?Effective hours: duplicates, quality and collapse: Count effective rather than recorded hours: discount duplicates, weight by quality, and watch for loss of the tail in loops.Counting effective hours instead of recorded ones shrinks a corpus, sometimes by orders of magnitude, and the hours that survive still have to reach the accelerators: a model that trains on video spends its time decoding it. If the processors wait for the data, the data budget is not the only bill. What does it cost to move the effective hours, and where does the machine stall?
23Counting effective hours instead of recorded ones shrinks a corpus, sometimes by orders of magnitude, and the hours that survive still have to reach the accelerators: a model that trains on video spends its time decoding it. If the processors wait for the data, the data budget is not the only bill. What does it cost to move the effective hours, and where does the machine stall?Moving the bits: the infrastructure bill: Price the pipeline: bytes per hour, decode throughput, and the time an accelerator spends waiting for data.Moving the data costs a fixed share of the budget and decides whether the processors are ever busy; it is the last bill. We can now price every source, its value, its shelf life, its duplicates and its transport. Put together, which kind of data does embodied training eat the most, by volume and by value, and what should be bought next?
24Moving the data costs a fixed share of the budget and decides whether the processors are ever busy; it is the last bill. We can now price every source, its value, its shelf life, its duplicates and its transport. Put together, which kind of data does embodied training eat the most, by volume and by value, and what should be bought next?What does embodied training eat the most?: Combine value, price, shelf life, effective size and transport into one procedure for what to buy next.The ledger says what to buy next for one policy on one robot, at this year's prices and exchange rates. Prices move, simulators improve and models get larger, and some columns, force and recovery and intent, are absent from every cheap source. Which of these prices will move first, and which of the zeros will not?
How to read this
Straight through, in order: the track is one argument, and each lesson's "Linear position" box names the problem it inherits, the single idea it adds and the problem it hands on. If you have time only for the spine, read 01 → 02 → 03 → 04 → 06 → 07 → 12 → 15 → 16 → 17 → 18 → 21 → 24. Each lesson has one widget that runs its mechanism; from lesson 16 on, each computes one line of the ledger. To re-check every quoted number yourself, run python3 tools/chain/validate_chain.py --series rob --oracles from the repository root: it re-derives each one with an independent script.

Where this track sits

Before it: World Models builds what a robot would plan with and, in its last fifteen lessons, asks how such a model is trained; it ends on the question this track begins with. Beside it: World Models · 27 Data mixture asks the data question for a world model, Reinforcement Learning · 17 Imitation and IRL, 19 Offline RL and 56 Robotics cover the same problems from the decision-process side, Synthetic Vision covers the simulated data that lesson 11 asks for, Data Engineering and Data-Intensive Systems cover the pipelines of lesson 23, and Financials covers unit economics and discounting. After it: the index of all lessons.