Is data perishable? Shelf life and the supply of corrections
Lesson 18 spent a budget by the slope of each source's curve and assumed that an hour keeps its value once bought. A demonstration does. A correction was made in the states an older policy visited, so it may expire when the policy is retrained, and this lesson measures whether it does. On the Bench it does not: per frame, the first clone's set is worth 1.83 times a fresh set to the fourth clone, and the lowest of fifteen cells is 0.87. A correction loses value only when the world moves outward, by about a fifth when the gust doubles. What runs out is the supply: a fresh set of 800 labelled frames adds 18.5 points of success to the first clone and 0.85 to the fourth, which fails one run in 18. The lesson cannot say what a fleet costs to supervise.
New idea: a record's shelf life is set by what its distribution depends on, and a correction depends on the policy that made it: it keeps its value for as long as the policy being trained stays among the states it was made in, which a better policy does. So the stock of corrections does not expire, and what runs out is the supply of failures to correct.
Forces next: A correction does not lose its value when the policy it was made for is retrained: a better policy stays among the states of a worse one, and on the Bench a set made for the first clone is worth as much per frame to the fourth clone as a fresh set, losing value only when the world moves outward, by about a fifth when the gust doubles. Demonstrations and force recordings, which the expert made and no policy did, keep theirs whoever is trained. What runs out is the supply: a correction teaches what its policy got wrong, the first set adds about eighteen points of success and the fourth about one, and a policy that fails one run in twenty has little left to show. The cheapest continuing supply of corrections is the deployed robots themselves, working, with a person stepping in when they fail. Does a fleet like that improve the policy fast enough to pay for the people who supervise it?
1 · What expiry would cost
The recovery task of lessons 1 to 3 is a cup carried by a two-joint arm along a path between five posts, in a gust of 0.05. A clone averages the stored commands of the frames nearest to the state it is in, within a bandwidth of 0.02 rad, and a run is lost when the cup touches a post. Twenty calm demonstrations train clone 1. Then DAgger (Ross, Gordon and Bagnell, 2011): run the clone, have the expert program label every state it visited, store those frames with the rest (the stock) and train on the union. A round is five runs and 824 labelled frames on average, one frame being 0.05 s; clone n + 1 is trained on the demonstrations and rounds 1 to n. The sets below are 800 frames, 40 s of motion, $0.99 at lesson 17's assumed price of $89 an hour of corrections. Scored on 1000 runs of each clone in each of four lineages of collection runs, the clones complete the course 62.0, 77.5, 86.9, 94.4, 95.8 and 96.6 % of the time.
Lessons 17 and 18 book these frames as hours: four rounds are four rounds. Suppose retraining took a share of a correction's value, so that a set made j clones before the one being trained is worth qj of a fresh one (q is the share kept per retraining). A stock of four rounds, the newest fresh and the others one, two and three retrainings old, would be worth 1 + q + q² + q³ fresh sets: 2.95 at q = 0.8 and 1.88 at q = 0.5, 26 % and 53 % less than four. The rule of lesson 18 would then be a rule for renting a stock, not buying one, because every retraining would send part of it back. Which is true is a measurement, and it needs a ruler.
2 · A ruler for an old set
A set is the first 800 labelled frames of eight fresh runs of one clone under the gust, about one round. Its age k, for the clone n being trained, is the number of clones between them: the set was made by clone n − k, and k = 0 is a fresh set. Add the set to clone n's stock, rebuild (rebuilding is storing) and score the rebuilt clone two ways. The success gain is the points of success it adds on 200 runs with the same starts and gusts as clone n itself, so the set is the only difference; it is what the ledger pays for. The copy error removed is the share by which the root-mean-square gap between the clone's command and the expert's, relative to the expert's size (lesson 1), falls on clone n's own frames from 30 runs. The relevance rk of a set of age k is its value divided by the value of a fresh set. Rent predicts rk = qk, falling with age; capital predicts rk = 1. Four choices fix the ruler, each excluding a comparison that looks fair:
| choice | what it excludes |
|---|---|
| both sets join the same stock of the same clone, 800 frames each | comparing the gains of different clones (a gain falls with success whatever the data) or a set with a stock (more frames, more value) |
| the old set is a new draw of the old clone's runs | re-adding the stored round, which brings no new state; the rule is conservative against old sets, because the stock holds a twin of an old draw and none of a fresh one |
| scored on the clone's own runs and frames | scoring on the expert's frames: a clone is lost on the frames it makes itself (lesson 3) |
| two measures | success alone: from clone 4 on no set adds more than about 2 points, and a ruler that reads zero cannot rank |
3 · The measurement: no decay
Each cell is 40 draws of a set, ten in each of four lineages (a lineage is one draw of the collection runs' gusts and starts), scored by the copy error removed. The second column is the value of a fresh set; the others are rk. Standard errors run from 0.08 to 0.71.
| clone retrained | a fresh set removes | r1 | r2 | r3 | r4 | r5 |
|---|---|---|---|---|---|---|
| 2 | 7.28 % | 0.87 | ||||
| 3 | 3.72 % | 1.09 | 1.28 | |||
| 4 | 1.76 % | 1.28 | 1.43 | 1.83 | ||
| 5 | 1.41 % | 1.21 | 1.49 | 2.03 | 2.23 | |
| 6 | 0.71 % | 1.70 | 1.43 | 1.95 | 2.47 | 3.33 |
Nothing falls with age. Along each row r rises with age, to 3.33 for the oldest set of clone 6, with one dip (clone 6, age 2) inside its standard errors; the lowest of the fifteen cells is 0.87 (± 0.08, clone 2, age 1), and the five cells at age 1 average 1.23. Rent at q = 0.8 predicts 0.80, 0.64, 0.51, 0.41 and 0.33 at ages 1 to 5. The first clone's set at clone 4 measures 1.83 ± 0.21, 6.3 standard errors above rent's 0.51, and still 4.6 above the 0.86 of a rent of only 5 % a retraining. Rent is rejected, and the older the set, the further.
The success measure agrees where it has room. At clones 2 and 3 the sets one and two clones old give r = 0.96 ± 0.10, 0.88 ± 0.12 and 0.87 ± 0.12: consistent with 1, though cells this wide cannot one at a time separate 0.8 from 1 at age 1. At clone 4 the stock is at 94.4 %, a fresh set adds 0.85 ± 0.34 points and sets made 1, 2 and 3 clones back add 1.28, 1.62 and 1.67 (± 0.3 to 0.4), none less than the fresh one. So the ledger may book a stock of corrections at the frames it holds, as far as retraining goes. The question becomes why.
4 · Why it does not fall: a label is a function of the state, and supports nest
An old set can fail a new clone in two ways: its labels are wrong for the new clone, or the new clone never goes where the set is. The first cannot happen with a program for a labeller. The label is the expert's command at a state, a function of the two joint angles (lesson 3), so the same state gets the same label whichever clone visited it and in whichever generation. (A person who relabels would not give the same answer twice: lesson 4.) What is left is where the clones go. From 100 runs of each clone in each lineage:
| clone | 99 % of its frames lie within (cm of the path) | frames beyond clone 1's reach | frames within a bandwidth of one clone-1 set |
|---|---|---|---|
| 1 | 4.90 | 1.00 % | 91.7 % |
| 2 | 4.67 | 0.54 % | 95.4 % |
| 3 | 4.50 | 0.29 % | 95.5 % |
| 4 | 4.40 | 0.25 % | 96.3 % |
| 5 | 4.28 | 0.14 % | 96.1 % |
| 6 | 4.26 | 0.14 % | 96.1 % |
Call the distance that holds 99 % of a clone's frames its reach, and the frames more than 4 cm from the path, where lesson 3 counts a step as lost, the tail. Both shrink: each round pulls a clone back toward the path (lesson 3), so a better policy visits a subset of the states of a worse one, and the set the worse one made still covers it. The bulk is covered whatever the age: 92 to 96 % of a clone's frames lie within a bandwidth of one set, rising and then level with the clone.
Old sets also hold more of the tail. 6.4 % of the frames of a set made by clone 1 are in the tail, against 2.2 % of a set made by clone 4, and among the 160 sets of every age drawn for clone 4 a set's value tracks that share (correlation 0.69). So r can rise with age. A frame in the tail is a state in which a clone went wrong, and its label says what to do there; the value of a correction frame comes from the failures of the policy that made it, and an older policy fails more.
5 · Where value does fall: the world moves outward
Nesting has a direction. If the retrained clone goes where the set was never made, the set cannot serve it, and a world that gets harder does that to a clone that has not changed. Sets made by clone 1 at gust 0.05, 0.075 and 0.10, the retrained clone run at each of the three (100 draws a cell). Cells are the success points a set of 800 frames adds, and in brackets r, the gain relative to a set made at the gust the clone runs in:
| set made at gust ↓, clone runs at → | 0.05 | 0.075 | 0.10 |
|---|---|---|---|
| 0.05 | 18.4 (1) | 16.8 (0.80) | 13.0 (0.81) |
| 0.075 | 21.6 (1.18) | 20.9 (1) | 15.9 (0.98) |
| 0.10 | 22.0 (1.20) | 21.0 (1.01) | 16.2 (1) |
Above the diagonal the set was made in a gentler world than the clone runs in, and it keeps 0.80 of a fresh set's value at ×1.5 and 0.81 at ×2 (± 0.04): about a fifth is lost (19 % when the gust doubles, the same at ×1.5). By the copy error the loss is larger, 0.76 and 0.70 kept. It is not proportional to the move: by success a set made at 0.075 loses almost nothing at 0.10 (0.98). Below the diagonal a harsher world's set is worth more than a fresh one. The reach says why. The farthest 1 % of clone 1's frames lies 4.90 cm from the path at gust 0.05, 5.35 cm at 0.075 and 6.18 cm at 0.10, and the share of its frames beyond the reach of the 0.05 world rises from 1.0 % to 2.2 % and 3.8 %: the clone goes where the set was not made. A set made at a harsher gust holds more of the tail (5.8 % of its frames at 0.05, 9.2 % at 0.075, 10.9 % at 0.10) and gains more.
The other way a policy changes is sideways. Replace the 20 demonstrations by 20 others and train from scratch: corrections made for the clone of the first twenty, added to the second twenty, give 17.5 points against 19.0 for corrections made for the new clone, r = 0.92 ± 0.06 over 48 draws. Both clones stay on the expert's path, so little is lost. In dollars the rent measured is none while the policy improves in a world that stays put. When the gust doubles it is a fifth of $0.99, $0.19 a set, conditional on the world moving.
The widget
What to try. Leave the defaults: clone 4, a set made by clone 1, three clones back, gust 0.05. On this lineage the set removes 3.42 % of the copy error (± 0.63) against 2.29 % for a fresh set, r = 1.49; the four lineages give 1.83 ± 0.21 and, by success, 1.95 ± 0.89. 97.1 % of the clone's frames lie within a bandwidth of the set, and 7.0 % of the set's frames are beyond 4 cm against 1.7 % for a fresh set: the old set has the tail. At clone 6 and age 5 the table says 3.33 ± 0.71. Choose clone 1, age 0 and the world set 0.05, clone 0.10: r is 0.69 live, 0.70 in the table by the copy error and 0.81 by success, and coverage falls to 73.2 %. The right panel is what runs out: the bars fall from 18.5 to 0.85 points at a constant $0.99.
6 · What runs out, and the three shelf lives
If old corrections keep their value, why does buying them stop paying? Because a set is worth the failures in it (§4). A correction teaches what its policy got wrong, and a clone that completes the course nine times in ten makes few mistakes to teach. A run fails here only by ending in a post, and with independent runs the table follows from the share that do. A set costs $0.99 whatever its clone:
| clone | runs ending in a post | failing runs in a round of five | rounds that show none | a fresh set adds (points) | points per dollar |
|---|---|---|---|---|---|
| 1 | 38.1 % | 1.90 | 9 % | 18.5 | 18.7 |
| 2 | 22.6 % | 1.13 | 28 % | 9.8 | 9.9 |
| 3 | 13.1 % | 0.66 | 50 % | 5.4 | 5.5 |
| 4 | 5.6 % | 0.28 | 75 % | 0.85 | 0.86 |
| 5 | 4.2 % | 0.21 | 81 % | 1.0 | 1.0 |
| 6 | 3.4 % | 0.17 | 84 % | 0.72 | 0.73 |
The first set adds 18.5 points, the fourth 0.85: 22 times fewer per dollar at the same price, and at clone 4 three rounds in four show the clone no failure at all. A clone that fails one run in twenty lies between clone 4, one in 18, and clone 5, one in 24, and has little left to show. The stock holds what earlier clones got wrong and keeps it (§3): the first clone's set, added to clone 4, still adds 1.67 points, twice a fresh set's 0.85 though a tenth of the 18.5 it added to clone 1, because most of what clone 1 failed at is taught already. What is scarce is the failures nobody has corrected yet, and the clone now running is the only source of them. The stock does not expire and the supply runs out.
Sort the sources by what the distribution of their records depends on:
| class | records | the distribution depends on | after a retraining | what is spent |
|---|---|---|---|---|
| capital | demonstrations, force recordings | the expert only | nothing changes: last year's set is a draw from the distribution of today's | nothing; what a set adds depends on the stock it joins (below) |
| inventory | footage, simulation | the world: a person's hand, the simulator's arm model | the records do not change; they are worth lesson 16's rate, up to a ceiling: lesson 18's pools stop at 56.0 % (footage) and 82.3 % (the 1 % simulator) at 256 layouts | the headroom: each layout is worth less than the last |
| corrections | rounds of DAgger, takeovers (HG-DAgger, Kelly et al., 2019) | the policy that made them | kept while the retrained policy stays among their states (§3, §4); a fifth lost when the gust doubles (§5) | the supply: failures of the policy now running |
Capital never goes stale and can still be worth less than nothing. In a stock that holds two or more rounds, the clone trained on the rounds alone beats the clone trained on the rounds and the 20 calm demonstrations (1000 runs in each of four lineages): at clone 3 96.9 % against 86.9 %, at clone 4 99.8 % against 94.4, at clone 5 99.7 % against 95.8, at clone 6 99.6 % against 96.6. Shelf life and marginal value are different columns of the ledger, as lesson 18's pools showed for sources. The ledger's corrections curve, 20 demonstrations plus rounds, levels off near 96 % in part because of them, and lesson 20 should read it as a floor for what corrections reach.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
Retraining does not spoil a correction. Per frame, the first clone's set is worth 1.83 ± 0.21 times a fresh set to the fourth clone (1.67 points against 0.85 by success) and no cell of fifteen is below 0.87, because the expert's label is a function of the state and a better policy stays among the states of a worse one. Value goes when the world moves outward, 0.81 kept when the gust doubles, and a record that the expert made and no policy did keeps its value whoever is trained. What runs out is the supply: a fresh set adds 18.5 points to the first clone and 0.85 to the fourth, 22 times fewer per dollar, and at the fourth clone 75 % of rounds show no failure. The cheapest continuing supply of corrections is the deployed robots themselves, working, with a person stepping in when they fail. Does a fleet like that improve the policy fast enough to pay for the people who supervise it?
Interview prompts
- Why can corrections made for an older policy be as good as fresh ones for the retrained policy? (§4 — the expert's label is a function of the state, and a better policy visits a subset of the states of a worse one.)
- How do you measure the shelf life of a set of data? (§2, §3 — the relevance of a set of age k at equal frames: rent predicts qk, capital 1; measured between 0.87 and 3.33.)
- When does a correction lose value, and how much on the Bench? (§5 — when the retrained policy goes where the set was never made: 0.81 kept when the gust doubles; a harsher world's set gains.)
- What is the difference between the stock and the supply of corrections? (§6 — the stock keeps its value; a fresh set adds 18.5 points to clone 1 and 0.85 to clone 4, which shows no failure in 75 % of rounds.)
- A demonstration never goes stale: can it be worth less than nothing? (§6 — dropping the 20 demonstrations raises clone 4 from 94.4 to 99.8 %: value in use depends on the stock.)
- Why not book the stock of corrections as rent and rebuy it? (§6 — keeping only the newest round costs 21 points; a set rebought adds 1.0.)
Companion reads: Lesson 3 · Labels on your own states, Lesson 4 · People are not functions and Reinforcement Learning · 17 Imitation and inverse RL.