Training a Robot Model, from first principles
A world model says what would happen if the robot did something. A robot has to choose, about twenty times a second, and the cheapest teacher for a chooser is a person who already is one. This track builds the chooser one forced step at a time, on a bench small enough to run in your browser: a two-link arm carrying a cup past a row of vases. It copies a person's demonstrations, finds why a copy that is accurate where the person went is lost where it goes, and repairs that with labels on its own states, distributions, chunks and feedback inside the action. It then asks how much data a policy needs and of which kind (layouts, other bodies, video, force, simulation), improves a policy from a score alone, builds the model whose two clocks fit a robot, trains it in stages without erasing what it knew, and prices the evaluation that says whether any of it worked. The last nine lessons then put a unit and a price on everything the first fifteen asked for. They measure every source of data in one unit, the hours of your own robot's demonstrations that an hour of it replaces; price each; let the value of an hour fall as you buy; ask how long an hour keeps its value and whether a fleet of working robots pays for its supervision; spend a budget as a sequence of purchases; count effective hours instead of recorded ones; price moving the bits; and read the ledger to say what training a robot model eats the most.
The exam
The first fifteen lessons are graded by one question asked in different ways: when the robot runs the policy, does the task get done, and how would we know? A copy error, a training loss and a validation curve each measure something else, and the first lessons are about the gap between them. The later ones are the moves forced by what a policy needs in order to close it. The last nine are graded by one formula: value per dollar = marginal value × exchange rate ÷ price. Each of those lessons measures or derives one factor of it, or one reason the factors change as you buy.
- Copy a person. A controller from demonstrations, why it fails in the loop it creates, and the labels that repair it (lessons 1–3).
- What the data really is. People are not functions, so a policy outputs a distribution; a policy commits to a sequence; and the action carries its own feedback (lessons 4–6).
- Where more data comes from. How much is enough and of which kind: layouts, other bodies, video without actions, the force a camera does not see, and simulation (lessons 7–11).
- Beyond the demonstrator. A score as the only teacher, the anatomy of the model, the training recipe, and the evaluation that pays for it all (lessons 12–15).
- Value and price. Every source in one unit, then in dollars, and the two rankings that disagree (lessons 16–17).
- How value moves. Diminishing returns, a stock of corrections that keeps its value while its supply runs out, and a fleet that pays for its supervisors only above a success rate (lessons 18–20).
- Spending it. The allocation problem, the effective size of a dataset, the bill for moving the bits, and the honest answer (lessons 21–24).
Part I · Copy a person (lessons 01–03)
A controller from demonstrations, why it fails in the loop, and the labels that repair it.
Part II · What the data really is (lessons 04–06)
People are not functions, so policies output distributions; committing to sequences; and feedback built into the action itself.
Part III · Where more data comes from (lessons 07–11)
How much is enough, other bodies, video without actions, the signal a camera lacks, and simulation.
Part IV · Beyond the demonstrator (lessons 12–15)
A score as the only teacher, the anatomy of the model, the training recipe, and the evaluation that pays for it all.
Part V · Value and price (lessons 16–17)
Every source of data in one unit, then in dollars, and the two rankings that disagree.
Part VI · How value moves (lessons 18–20)
Diminishing returns, a stock of corrections that keeps its value while its supply runs out, and a fleet that pays for its supervisors only above a success rate.
Part VII · Spending it (lessons 21–24)
The allocation problem, the effective size of a dataset, the bill for moving it, and the honest answer.
The whole derivation on one page
Read the second column of any row and the fourth column of the row above it: they are the same sentence. That is the linear thinking made visible. Each lesson's problem it inherits is exactly the previous lesson's problem it creates, and the middle column is the single move that resolves the first and manufactures the second. The first row inherits its problem from the last lesson of World Models (lesson 31, the whole bill for training one), where a model that can say what would happen has never chosen anything; the last row asks which prices will move and which gaps will not.
| # | The problem it inherits | The one move | The problem it creates |
|---|---|---|---|
| Part I · Copy a person | |||
| 01 | A trained world model answers what would happen if the robot did something, and it has never chosen anything: it has no goal, and the data it is tested on do not change when it is used. A robot has to choose, about twenty times a second, from what it has just sensed, and the cheapest teacher for a chooser is a person who already is one. What does a robot learn by copying a person doing the task, and how would we know it had learned it? | Copy a person who knows: behaviour cloning: Turn demonstrations into a controller by regression, and measure it two ways: how well it copies, and whether the task gets done. | A policy cloned from twenty demonstrations copies the expert to within a few percent on frames it has never seen, makes about four times that error on the frames it produces itself, and fails more than a third of the time on a course of five posts and more than half the time on six. The two kinds of frame come from different places: the first from where the expert went, the second from where the policy goes. Why is a policy that is accurate where the expert went lost where it goes, and how fast does the loss grow with the length of the task? |
| 02 | A policy cloned from twenty demonstrations copies the expert to within a few percent on frames it has never seen, makes about four times that error on the frames it produces itself, and fails more than a third of the time on a course of five posts and more than half the time on six. The two kinds of frame come from different places: the first from where the expert went, the second from where the policy goes. Why is a policy that is accurate where the expert went lost where it goes, and how fast does the loss grow with the length of the task? | The policy writes its own test set: Derive how a policy's own mistakes become its next input: a loop gain and an extrapolation error that make cost grow as the square of the horizon. | A policy's mistakes land it in states its training data never showed, where its next action is a guess, so one slip is charged for every step that follows: the cost grows as the square of the horizon unless the loop pulls the policy back or the training data include the states the policy reaches. More demonstrations recorded the same way leave the curve where it is. The only one who knows what to do in a state the policy has just wandered into is the expert. Can the expert label the states the policy visits, and what does that cost? |
| 03 | A policy's mistakes land it in states its training data never showed, where its next action is a guess, so one slip is charged for every step that follows: the cost grows as the square of the horizon unless the loop pulls the policy back or the training data include the states the policy reaches. More demonstrations recorded the same way leave the curve where it is. The only one who knows what to do in a state the policy has just wandered into is the expert. Can the expert label the states the policy visits, and what does that cost? | Labels on your own states: DAgger and corrections: Put the states the policy reaches into the training set, labelled by the expert: aggregate, retrain, repeat. | Training on the policy's own states, with the expert's actions as labels, takes a policy that finishes five posts about six times in ten to one that finishes them about nineteen times in twenty after four rounds of five rollouts, and a person supplies those labels best by taking over when the policy goes wrong. But people are not functions: asked from the same pose, or asked by two operators, one goes left of the post and another right, and a regression asked to fit both returns their average, a route that runs into the post. What should a policy output when the right action is not one point? |
| Part II · What the data really is | |||
| 04 | Training on the policy's own states, with the expert's actions as labels, takes a policy that finishes five posts about six times in ten to one that finishes them about nineteen times in twenty after four rounds of five rollouts, and a person supplies those labels best by taking over when the policy goes wrong. But people are not functions: asked from the same pose, or asked by two operators, one goes left of the post and another right, and a regression asked to fit both returns their average, a route that runs into the post. What should a policy output when the right action is not one point? | People are not functions: distributions over actions: Replace the single answer by a distribution over actions and sample it, so that two good routes stay two routes. | A policy that outputs a distribution over actions, and acts by sampling it, recovers both routes instead of their collision. But it draws a new sample at every step, and near the fork each draw can pick a different route: the arm dithers between them and a quarter of the rollouts still end in a post. Committing to one route for a while would stop the dithering, at the price of not looking while committed. How long should a policy commit, and what should it commit to? |
| 05 | A policy that outputs a distribution over actions, and acts by sampling it, recovers both routes instead of their collision. But it draws a new sample at every step, and near the fork each draw can pick a different route: the arm dithers between them and a quarter of the rollouts still end in a post. Committing to one route for a while would stop the dithering, at the price of not looking while committed. How long should a policy commit, and what should it commit to? | Commit: action chunks: Sample a whole sequence of actions at once and execute it, so that the choice of route is made once. | Sampling a whole sequence of actions at once and executing it commits the policy to one route, and the dithering stops. But the sequence is a list of velocities, a velocity command is added to the joint angles at every step, and every bit of noise in the plant stays added for the whole chunk: the longer the chunk, the further the arm drifts from where the sequence meant it to be, and the gain from committing is paid back as drift. What should the policy output, so that the plant's noise does not accumulate while it commits? |
| 06 | Sampling a whole sequence of actions at once and executing it commits the policy to one route, and the dithering stops. But the sequence is a list of velocities, a velocity command is added to the joint angles at every step, and every bit of noise in the plant stays added for the whole chunk: the longer the chunk, the further the arm drifts from where the sequence meant it to be, and the gain from committing is paid back as drift. What should the policy output, so that the plant's noise does not accumulate while it commits? | Targets, not velocities: feedback in the action: Make the output a target position tracked by a stiff controller, so that the plant's noise is corrected every step instead of accumulating. | A chunk of target positions, tracked by a stiff controller that pulls the arm back to the target at every step, removes the accumulation: on the demonstrated layout the policy completes the course nearly every time at any chunk length. But targets are anchored to where the demonstrations were. Shift the vases by a few centimetres and the policy walks faithfully into them. The policy has to read where the vases are and generalise across layouts, and nobody has said how many layouts that takes. How does success on a new layout grow with the data, and which data buy it? |
| Part III · Where more data comes from | |||
| 07 | A chunk of target positions, tracked by a stiff controller that pulls the arm back to the target at every step, removes the accumulation: on the demonstrated layout the policy completes the course nearly every time at any chunk length. But targets are anchored to where the demonstrations were. Shift the vases by a few centimetres and the policy walks faithfully into them. The policy has to read where the vases are and generalise across layouts, and nobody has said how many layouts that takes. How does success on a new layout grow with the data, and which data buy it? | How much data: the coverage law: Measure how success on a new layout grows with the number of layouts and demonstrations, and derive the exponent from what a local learner can reach. | For a policy that generalises only to what resembles its data, the error on a new layout is the distance to the nearest demonstrated layout, so success grows with the number of different layouts, with a power set by how many numbers describe a layout, and hardly at all with more demonstrations of a layout already seen. Ninety percent costs about a hundred layouts when two numbers describe a layout, and every further number multiplies that by the range it varies over divided by the width one layout covers; a real scene is described by many more than two numbers, and one arm cannot pay for that. Other laboratories have already collected such data, on other arms. What can a robot use from demonstrations made on a body that is not its own? |
| 08 | For a policy that generalises only to what resembles its data, the error on a new layout is the distance to the nearest demonstrated layout, so success grows with the number of different layouts, with a power set by how many numbers describe a layout, and hardly at all with more demonstrations of a layout already seen. Ninety percent costs about a hundred layouts when two numbers describe a layout, and every further number multiplies that by the range it varies over divided by the width one layout covers; a real scene is described by many more than two numbers, and one arm cannot pay for that. Other laboratories have already collected such data, on other arms. What can a robot use from demonstrations made on a body that is not its own? | Other bodies: units, pooling, interference: Pool demonstrations from other arms by writing the action in a space every arm shares, and let each body supply its own joint-level controller. | Demonstrations from other bodies help when the action is written in a space every body shares, the place the hand should go, and each arm turns that into its own joint motions; written in joint angles they are worse than no data. Pooled this way, all the robot demonstrations ever recorded are a small fraction of what people have filmed themselves doing, and the video has no actions in it. What can be learned from watching, and what is missing from the video that no amount of it can supply? |
| 09 | Demonstrations from other bodies help when the action is written in a space every body shares, the place the hand should go, and each arm turns that into its own joint motions; written in joint angles they are worse than no data. Pooled this way, all the robot demonstrations ever recorded are a small fraction of what people have filmed themselves doing, and the video has no actions in it. What can be learned from watching, and what is missing from the video that no amount of it can supply? | Watching: learning from video without actions: Infer the missing actions from motion with a model trained on a few labelled hours, then train on the pseudo-labelled video. | A model that infers the missing actions from a few hours of labelled motion can turn a large pile of video into demonstrations, and what it learns is what to do next, not how hard to push. Video records where things moved, never the force that moved them: an insertion with a millimetre of clearance, watched through a camera that cannot resolve a millimetre, looks the same when it succeeds and when it jams. What must a policy sense, and how must it act, when the task is decided by something the camera cannot see? |
| 10 | A model that infers the missing actions from a few hours of labelled motion can turn a large pile of video into demonstrations, and what it learns is what to do next, not how hard to push. Video records where things moved, never the force that moved them: an insertion with a millimetre of clearance, watched through a camera that cannot resolve a millimetre, looks the same when it succeeds and when it jams. What must a policy sense, and how must it act, when the task is decided by something the camera cannot see? | What cameras cannot see: contact and force: Give the policy the signal the task depends on, force, and an action that yields to it, instead of position alone. | A policy that feels force and yields to it, instead of holding a position stiffly, inserts the peg where position control jams, and the force it needs is in the data only if someone recorded it with a force sensor in the loop, on a real robot, slowly, with wear. A simulator produces force, actions and unlimited demonstrations, at the price of a physics that is not quite the real one. How wrong can a simulator be before a policy trained in it fails on the robot, and what buys back the difference? |
| 11 | A policy that feels force and yields to it, instead of holding a position stiffly, inserts the peg where position control jams, and the force it needs is in the data only if someone recorded it with a force sensor in the loop, on a real robot, slowly, with wear. A simulator produces force, actions and unlimited demonstrations, at the price of a physics that is not quite the real one. How wrong can a simulator be before a policy trained in it fails on the robot, and what buys back the difference? | Simulation: unlimited data, wrong physics: Train in a simulator and close the gap with randomised physics and a few real rollouts for calibration. | A simulator is worth using when its error is narrower than the policy can tolerate: randomising the physics widens what a policy survives at the price of a more cautious motion, and calibrating the simulator on a few real rollouts narrows the gap itself. Both assume there is someone in the simulator to copy, a scripted expert or a planner. For most tasks that matter there is no such expert, only a way to tell whether the task was done. What can be learned from a score alone, and how many attempts does it take? |
| Part IV · Beyond the demonstrator | |||
| 12 | A simulator is worth using when its error is narrower than the policy can tolerate: randomising the physics widens what a policy survives at the price of a more cautious motion, and calibrating the simulator on a few real rollouts narrows the gap itself. Both assume there is someone in the simulator to copy, a scripted expert or a planner. For most tasks that matter there is no such expert, only a way to tell whether the task was done. What can be learned from a score alone, and how many attempts does it take? | A score as the only teacher: reinforcement learning: Improve a policy from a reward alone: sample the policy, keep what scored, and start from a good prior so that few attempts are needed. | Learning from a score alone works, slowly, on a small task, and works in an afternoon of robot time only when it starts from a policy that is already nearly right and learns a small correction. Every source so far feeds that starting point: demonstrations, other bodies, video, force, simulation. It is one model that must read language and images, feel force, emit chunks, and still answer within one control step. How is such a model built, and what does it cost to run it at the speed of the robot? |
| 13 | Learning from a score alone works, slowly, on a small task, and works in an afternoon of robot time only when it starts from a policy that is already nearly right and learns a small correction. Every source so far feeds that starting point: demonstrations, other bodies, video, force, simulation. It is one model that must read language and images, feel force, emit chunks, and still answer within one control step. How is such a model built, and what does it cost to run it at the speed of the robot? | Anatomy of a robot model: two clocks: Split the model into a slow reasoner and a fast action network, and decide what each must do within its own clock. | A large model that reads and reasons is too slow to steer an arm directly, so it runs slowly and hands a plan to a small fast network that runs at the control rate, and the split decides what each part must be trained on. Training the parts in order, on data of very different size and quality, has its own failure: fine-tuning on a few dozen tasks raises every number we track while erasing what the language model knew. In what order, on what mixture, should the pieces be trained, and how do we notice the loss? |
| 14 | A large model that reads and reasons is too slow to steer an arm directly, so it runs slowly and hands a plan to a small fast network that runs at the control rate, and the split decides what each part must be trained on. Training the parts in order, on data of very different size and quality, has its own failure: fine-tuning on a few dozen tasks raises every number we track while erasing what the language model knew. In what order, on what mixture, should the pieces be trained, and how do we notice the loss? | The recipe: stages, mixtures and forgetting: Train in stages and keep old data in the mixture, so that what the model knew survives what it is taught. | Staging the training and mixing the old data into the new keeps the language the model started with, at a measurable price in speed and in the tasks it learns. We now have a recipe and a checkpoint, and need to say whether the checkpoint is better than the last one. An evaluation is a count of trials, and each trial costs a robot and a person for minutes. How many trials does it take to tell one policy from another, and what does that do to the cost of the whole programme? |
| 15 | Staging the training and mixing the old data into the new keeps the language the model started with, at a measurable price in speed and in the tasks it learns. We now have a recipe and a checkpoint, and need to say whether the checkpoint is better than the last one. An evaluation is a count of trials, and each trial costs a robot and a person for minutes. How many trials does it take to tell one policy from another, and what does that do to the cost of the whole programme? | Evaluation you can afford: Treat the success rate as a count of trials: put an interval on it, size the experiment for the difference you care about, and price it. | Twenty trials cannot tell seventy-six percent from sixty-eight, telling a ten-point improvement apart takes hundreds of trials per policy, and a programme that evaluates honestly spends more robot time on evaluation than on training. The programme also needs data of several kinds, each from a source with a different price, a different usefulness and a different shelf life. How much of each kind do we need, in what unit can they even be compared, and what does each cost? |
| Part V · Value and price | |||
| 16 | Twenty trials cannot tell seventy-six percent from sixty-eight, telling a ten-point improvement apart takes hundreds of trials per policy, and a programme that evaluates honestly spends more robot time on evaluation than on training. The programme also needs data of several kinds, each from a source with a different price, a different usefulness and a different shelf life. How much of each kind do we need, in what unit can they even be compared, and what does each cost? | An hour is not an hour: exchange rates: Measure every source in the same unit: how many hours of your own robot's demonstrations an hour of it replaces. | Measured on the Bench, an hour of another arm's demonstrations replaces a fraction of an hour of your own, an hour of labelled video a smaller fraction, an hour of simulation a fraction that depends on the gap, and a source that lacks a column the task needs replaces none of it, however many hours there are. Exchange rates put every source in one unit, but not in dollars. What does an hour of each source cost? |
| 17 | Measured on the Bench, an hour of another arm's demonstrations replaces a fraction of an hour of your own, an hour of labelled video a smaller fraction, an hour of simulation a fraction that depends on the gap, and a source that lacks a column the task needs replaces none of it, however many hours there are. Exchange rates put every source in one unit, but not in dollars. What does an hour of each source cost? | What an hour costs: prices and the ledger: Build each source's price per hour from labour, throughput, compute and authoring, and divide it by the source's exchange rate. | Prices and exchange rates together give a cost per useful hour for every source, the price of an hour divided by its rate, and the ranking they produce is not the ranking by price: footage, dearer to make than the other borrowed hours, becomes the dearest useful hour of any source that can teach the same task, a simulator is the cheapest useful hour only while its gap is small and sells no useful hour at any price once the gap is wide, and some columns cannot be bought from the sources that are cheap. But a cost per hour is the cost of the first hour. The thousandth hour of a source teaches the policy less than the first. How fast does the value of one more hour fall? |
| Part VI · How value moves | |||
| 18 | Prices and exchange rates together give a cost per useful hour for every source, the price of an hour divided by its rate, and the ranking they produce is not the ranking by price: footage, dearer to make than the other borrowed hours, becomes the dearest useful hour of any source that can teach the same task, a simulator is the cheapest useful hour only while its gap is small and sells no useful hour at any price once the gap is wide, and some columns cannot be bought from the sources that are cheap. But a cost per hour is the cost of the first hour. The thousandth hour of a source teaches the policy less than the first. How fast does the value of one more hour fall? | The thousandth hour: diminishing returns: Measure the value curve of each source and buy until the marginal value per dollar is equal across sources. | The value of every source falls as a power of the hours bought, so the rule is to buy from each source until its marginal value per dollar equals that of the next, and which source is cheapest changes as you buy. This treats the hours as an asset that keeps its value once bought. Demonstrations do, for the policy trained next year as well as for the one trained today. Corrections were made in the states an older policy visited, and the policy trained next year will not visit quite the same ones. How long does a correction keep its value? |
| 19 | The value of every source falls as a power of the hours bought, so the rule is to buy from each source until its marginal value per dollar equals that of the next, and which source is cheapest changes as you buy. This treats the hours as an asset that keeps its value once bought. Demonstrations do, for the policy trained next year as well as for the one trained today. Corrections were made in the states an older policy visited, and the policy trained next year will not visit quite the same ones. How long does a correction keep its value? | Is data perishable? Shelf life and the supply of corrections: Classify data by how long it stays useful: capital that keeps its value, inventory that saturates, and corrections, whose value lasts as long as the policy stays among the states they were made in. | A correction does not lose its value when the policy it was made for is retrained: a better policy stays among the states of a worse one, and on the Bench a set made for the first clone is worth as much per frame to the fourth clone as a fresh set, losing value only when the world moves outward, by about a fifth when the gust doubles. Demonstrations and force recordings, which the expert made and no policy did, keep theirs whoever is trained. What runs out is the supply: a correction teaches what its policy got wrong, the first set adds about eighteen points of success and the fourth about one, and a policy that fails one run in twenty has little left to show. The cheapest continuing supply of corrections is the deployed robots themselves, working, with a person stepping in when they fail. Does a fleet like that improve the policy fast enough to pay for the people who supervise it? |
| 20 | A correction does not lose its value when the policy it was made for is retrained: a better policy stays among the states of a worse one, and on the Bench a set made for the first clone is worth as much per frame to the fourth clone as a fresh set, losing value only when the world moves outward, by about a fifth when the gust doubles. Demonstrations and force recordings, which the expert made and no policy did, keep theirs whoever is trained. What runs out is the supply: a correction teaches what its policy got wrong, the first set adds about eighteen points of success and the fourth about one, and a policy that fails one run in twenty has little left to show. The cheapest continuing supply of corrections is the deployed robots themselves, working, with a person stepping in when they fail. Does a fleet like that improve the policy fast enough to pay for the people who supervise it? | The flywheel: a fleet as a source: Model deployment as a data source whose supply depends on the policy's own success, and find the success rate at which it pays for itself. | A fleet pays for its own supervisors only above the success rate at which a takeover costs less than the task it saves; below that rate every point the fleet gains is paid for by the programme, and above it the failures that supply the corrections grow rarer as the policy improves. We now have sources with falling returns, a stock of corrections that keeps its value while its supply runs out, and a fleet that has to be funded until it takes off, and each has a price, a value curve and a shelf life. Given a budget, how much of each do we buy, and in what order? |
| Part VII · Spending it | |||
| 21 | A fleet pays for its own supervisors only above the success rate at which a takeover costs less than the task it saves; below that rate every point the fleet gains is paid for by the programme, and above it the failures that supply the corrections grow rarer as the policy improves. We now have sources with falling returns, a stock of corrections that keeps its value while its supply runs out, and a fleet that has to be funded until it takes off, and each has a price, a value curve and a shelf life. Given a budget, how much of each do we buy, and in what order? | Spending the budget: the allocation problem: Maximise the weakest capability under a budget by marginal value per dollar: a sequence of purchases, not a mixture. | Spending by marginal value per dollar gives a sequence and not a mixture: cheap breadth first, the scarce column when the cheap source saturates, and the instrument for that column before its price rank says so. The plan counts hours as if an hour were an hour of information. A hundred thousand recorded episodes can hold eight hundred distinct situations, a policy trained on its own samples loses its tail, and a bad batch looks like a good one. What is the effective size of a dataset? |
| 22 | Spending by marginal value per dollar gives a sequence and not a mixture: cheap breadth first, the scarce column when the cheap source saturates, and the instrument for that column before its price rank says so. The plan counts hours as if an hour were an hour of information. A hundred thousand recorded episodes can hold eight hundred distinct situations, a policy trained on its own samples loses its tail, and a bad batch looks like a good one. What is the effective size of a dataset? | Effective hours: duplicates, quality and collapse: Count effective rather than recorded hours: discount duplicates, weight by quality, and watch for loss of the tail in loops. | Counting effective hours instead of recorded ones shrinks a corpus, sometimes by orders of magnitude, and the hours that survive still have to reach the accelerators: a model that trains on video spends its time decoding it. If the processors wait for the data, the data budget is not the only bill. What does it cost to move the effective hours, and where does the machine stall? |
| 23 | Counting effective hours instead of recorded ones shrinks a corpus, sometimes by orders of magnitude, and the hours that survive still have to reach the accelerators: a model that trains on video spends its time decoding it. If the processors wait for the data, the data budget is not the only bill. What does it cost to move the effective hours, and where does the machine stall? | Moving the bits: the infrastructure bill: Price the pipeline: bytes per hour, decode throughput, and the time an accelerator spends waiting for data. | Moving the data costs a fixed share of the budget and decides whether the processors are ever busy; it is the last bill. We can now price every source, its value, its shelf life, its duplicates and its transport. Put together, which kind of data does embodied training eat the most, by volume and by value, and what should be bought next? |
| 24 | Moving the data costs a fixed share of the budget and decides whether the processors are ever busy; it is the last bill. We can now price every source, its value, its shelf life, its duplicates and its transport. Put together, which kind of data does embodied training eat the most, by volume and by value, and what should be bought next? | What does embodied training eat the most?: Combine value, price, shelf life, effective size and transport into one procedure for what to buy next. | The ledger says what to buy next for one policy on one robot, at this year's prices and exchange rates. Prices move, simulators improve and models get larger, and some columns, force and recovery and intent, are absent from every cheap source. Which of these prices will move first, and which of the zeros will not? |
python3 tools/chain/validate_chain.py --series rob --oracles from the repository root: it re-derives each one with an independent script.
Where this track sits
Before it: World Models builds what a robot would plan with and, in its last fifteen lessons, asks how such a model is trained; it ends on the question this track begins with. Beside it: World Models · 27 Data mixture asks the data question for a world model, Reinforcement Learning · 17 Imitation and IRL, 19 Offline RL and 56 Robotics cover the same problems from the decision-process side, Synthetic Vision covers the simulated data that lesson 11 asks for, Data Engineering and Data-Intensive Systems cover the pipelines of lesson 23, and Financials covers unit economics and discounting. After it: the index of all lessons.