all_lessons/Robot Model Training/19 · Perishable datalesson 19 / 24

Is data perishable? Shelf life and the supply of corrections

Lesson 18 spent a budget by the slope of each source's curve and assumed that an hour keeps its value once bought. A demonstration does. A correction was made in the states an older policy visited, so it may expire when the policy is retrained, and this lesson measures whether it does. On the Bench it does not: per frame, the first clone's set is worth 1.83 times a fresh set to the fourth clone, and the lowest of fifteen cells is 0.87. A correction loses value only when the world moves outward, by about a fifth when the gust doubles. What runs out is the supply: a fresh set of 800 labelled frames adds 18.5 points of success to the first clone and 0.85 to the fourth, which fails one run in 18. The lesson cannot say what a fleet costs to supervise.

The thesis, here
A correction is a state and the expert's command for it. The command is a function of the state, so a stored correction is never wrong, only unused: it serves any policy that visits near its state. A policy that improves visits a subset of the states of the one before it, so old corrections keep serving it, and what ends their life is a world that moves outward, past the states they cover. What a better policy runs out of is failures to correct: the stock keeps its value and the supply thins.
Linear position
Forced by: The value of every source falls as a power of the hours bought, so the rule is to buy from each source until its marginal value per dollar equals that of the next, and which source is cheapest changes as you buy. This treats the hours as an asset that keeps its value once bought. Demonstrations do, for the policy trained next year as well as for the one trained today. Corrections were made in the states an older policy visited, and the policy trained next year will not visit quite the same ones. How long does a correction keep its value?
New idea: a record's shelf life is set by what its distribution depends on, and a correction depends on the policy that made it: it keeps its value for as long as the policy being trained stays among the states it was made in, which a better policy does. So the stock of corrections does not expire, and what runs out is the supply of failures to correct.
Forces next: A correction does not lose its value when the policy it was made for is retrained: a better policy stays among the states of a worse one, and on the Bench a set made for the first clone is worth as much per frame to the fourth clone as a fresh set, losing value only when the world moves outward, by about a fifth when the gust doubles. Demonstrations and force recordings, which the expert made and no policy did, keep theirs whoever is trained. What runs out is the supply: a correction teaches what its policy got wrong, the first set adds about eighteen points of success and the fourth about one, and a policy that fails one run in twenty has little left to show. The cheapest continuing supply of corrections is the deployed robots themselves, working, with a person stepping in when they fail. Does a fleet like that improve the policy fast enough to pay for the people who supervise it?
The plan
Six moves. (1) Price the question: what would expiry cost the ledger? (2) Build a ruler for the value of an old set to a retrained clone. (3) Read it by age. (4) Find why it does not fall: the label is a function of the state, and supports nest. (5) Find where it does fall: the world moves outward. (6) Find what runs out, and sort the sources by how long they last.

1 · What expiry would cost

The recovery task of lessons 1 to 3 is a cup carried by a two-joint arm along a path between five posts, in a gust of 0.05. A clone averages the stored commands of the frames nearest to the state it is in, within a bandwidth of 0.02 rad, and a run is lost when the cup touches a post. Twenty calm demonstrations train clone 1. Then DAgger (Ross, Gordon and Bagnell, 2011): run the clone, have the expert program label every state it visited, store those frames with the rest (the stock) and train on the union. A round is five runs and 824 labelled frames on average, one frame being 0.05 s; clone n + 1 is trained on the demonstrations and rounds 1 to n. The sets below are 800 frames, 40 s of motion, $0.99 at lesson 17's assumed price of $89 an hour of corrections. Scored on 1000 runs of each clone in each of four lineages of collection runs, the clones complete the course 62.0, 77.5, 86.9, 94.4, 95.8 and 96.6 % of the time.

Lessons 17 and 18 book these frames as hours: four rounds are four rounds. Suppose retraining took a share of a correction's value, so that a set made j clones before the one being trained is worth qj of a fresh one (q is the share kept per retraining). A stock of four rounds, the newest fresh and the others one, two and three retrainings old, would be worth 1 + q + q² + q³ fresh sets: 2.95 at q = 0.8 and 1.88 at q = 0.5, 26 % and 53 % less than four. The rule of lesson 18 would then be a rule for renting a stock, not buying one, because every retraining would send part of it back. Which is true is a measurement, and it needs a ruler.

2 · A ruler for an old set

A set is the first 800 labelled frames of eight fresh runs of one clone under the gust, about one round. Its age k, for the clone n being trained, is the number of clones between them: the set was made by clone n − k, and k = 0 is a fresh set. Add the set to clone n's stock, rebuild (rebuilding is storing) and score the rebuilt clone two ways. The success gain is the points of success it adds on 200 runs with the same starts and gusts as clone n itself, so the set is the only difference; it is what the ledger pays for. The copy error removed is the share by which the root-mean-square gap between the clone's command and the expert's, relative to the expert's size (lesson 1), falls on clone n's own frames from 30 runs. The relevance rk of a set of age k is its value divided by the value of a fresh set. Rent predicts rk = qk, falling with age; capital predicts rk = 1. Four choices fix the ruler, each excluding a comparison that looks fair:

choicewhat it excludes
both sets join the same stock of the same clone, 800 frames eachcomparing the gains of different clones (a gain falls with success whatever the data) or a set with a stock (more frames, more value)
the old set is a new draw of the old clone's runsre-adding the stored round, which brings no new state; the rule is conservative against old sets, because the stock holds a twin of an old draw and none of a fresh one
scored on the clone's own runs and framesscoring on the expert's frames: a clone is lost on the frames it makes itself (lesson 3)
two measuressuccess alone: from clone 4 on no set adds more than about 2 points, and a ruler that reads zero cannot rank

3 · The measurement: no decay

Each cell is 40 draws of a set, ten in each of four lineages (a lineage is one draw of the collection runs' gusts and starts), scored by the copy error removed. The second column is the value of a fresh set; the others are rk. Standard errors run from 0.08 to 0.71.

clone retraineda fresh set removesr1r2r3r4r5
27.28 %0.87
33.72 %1.091.28
41.76 %1.281.431.83
51.41 %1.211.492.032.23
60.71 %1.701.431.952.473.33

Nothing falls with age. Along each row r rises with age, to 3.33 for the oldest set of clone 6, with one dip (clone 6, age 2) inside its standard errors; the lowest of the fifteen cells is 0.87 (± 0.08, clone 2, age 1), and the five cells at age 1 average 1.23. Rent at q = 0.8 predicts 0.80, 0.64, 0.51, 0.41 and 0.33 at ages 1 to 5. The first clone's set at clone 4 measures 1.83 ± 0.21, 6.3 standard errors above rent's 0.51, and still 4.6 above the 0.86 of a rent of only 5 % a retraining. Rent is rejected, and the older the set, the further.

The success measure agrees where it has room. At clones 2 and 3 the sets one and two clones old give r = 0.96 ± 0.10, 0.88 ± 0.12 and 0.87 ± 0.12: consistent with 1, though cells this wide cannot one at a time separate 0.8 from 1 at age 1. At clone 4 the stock is at 94.4 %, a fresh set adds 0.85 ± 0.34 points and sets made 1, 2 and 3 clones back add 1.28, 1.62 and 1.67 (± 0.3 to 0.4), none less than the fresh one. So the ledger may book a stock of corrections at the frames it holds, as far as retraining goes. The question becomes why.

4 · Why it does not fall: a label is a function of the state, and supports nest

An old set can fail a new clone in two ways: its labels are wrong for the new clone, or the new clone never goes where the set is. The first cannot happen with a program for a labeller. The label is the expert's command at a state, a function of the two joint angles (lesson 3), so the same state gets the same label whichever clone visited it and in whichever generation. (A person who relabels would not give the same answer twice: lesson 4.) What is left is where the clones go. From 100 runs of each clone in each lineage:

clone99 % of its frames lie within (cm of the path)frames beyond clone 1's reachframes within a bandwidth of one clone-1 set
14.901.00 %91.7 %
24.670.54 %95.4 %
34.500.29 %95.5 %
44.400.25 %96.3 %
54.280.14 %96.1 %
64.260.14 %96.1 %

Call the distance that holds 99 % of a clone's frames its reach, and the frames more than 4 cm from the path, where lesson 3 counts a step as lost, the tail. Both shrink: each round pulls a clone back toward the path (lesson 3), so a better policy visits a subset of the states of a worse one, and the set the worse one made still covers it. The bulk is covered whatever the age: 92 to 96 % of a clone's frames lie within a bandwidth of one set, rising and then level with the clone.

Old sets also hold more of the tail. 6.4 % of the frames of a set made by clone 1 are in the tail, against 2.2 % of a set made by clone 4, and among the 160 sets of every age drawn for clone 4 a set's value tracks that share (correlation 0.69). So r can rise with age. A frame in the tail is a state in which a clone went wrong, and its label says what to do there; the value of a correction frame comes from the failures of the policy that made it, and an older policy fails more.

5 · Where value does fall: the world moves outward

Nesting has a direction. If the retrained clone goes where the set was never made, the set cannot serve it, and a world that gets harder does that to a clone that has not changed. Sets made by clone 1 at gust 0.05, 0.075 and 0.10, the retrained clone run at each of the three (100 draws a cell). Cells are the success points a set of 800 frames adds, and in brackets r, the gain relative to a set made at the gust the clone runs in:

set made at gust ↓, clone runs at →0.050.0750.10
0.0518.4 (1)16.8 (0.80)13.0 (0.81)
0.07521.6 (1.18)20.9 (1)15.9 (0.98)
0.1022.0 (1.20)21.0 (1.01)16.2 (1)

Above the diagonal the set was made in a gentler world than the clone runs in, and it keeps 0.80 of a fresh set's value at ×1.5 and 0.81 at ×2 (± 0.04): about a fifth is lost (19 % when the gust doubles, the same at ×1.5). By the copy error the loss is larger, 0.76 and 0.70 kept. It is not proportional to the move: by success a set made at 0.075 loses almost nothing at 0.10 (0.98). Below the diagonal a harsher world's set is worth more than a fresh one. The reach says why. The farthest 1 % of clone 1's frames lies 4.90 cm from the path at gust 0.05, 5.35 cm at 0.075 and 6.18 cm at 0.10, and the share of its frames beyond the reach of the 0.05 world rises from 1.0 % to 2.2 % and 3.8 %: the clone goes where the set was not made. A set made at a harsher gust holds more of the tail (5.8 % of its frames at 0.05, 9.2 % at 0.075, 10.9 % at 0.10) and gains more.

The other way a policy changes is sideways. Replace the 20 demonstrations by 20 others and train from scratch: corrections made for the clone of the first twenty, added to the second twenty, give 17.5 points against 19.0 for corrections made for the new clone, r = 0.92 ± 0.06 over 48 draws. Both clones stay on the expert's path, so little is lost. In dollars the rent measured is none while the policy improves in a world that stays put. When the gust doubles it is a fifth of $0.99, $0.19 a set, conditional on the world moving.

The widget

Add an old set to a retrained clone
Top: the course, the cup moving left to right. Grey dots are where the retrained clone puts the cup in 30 runs, red ones lie a bandwidth or more from every frame of the set, teal dots are the set (800 labelled frames of eight runs of the clone that made it). Bottom left: relevance r by age of the set, from four lineages (teal, ± 1 s.e.), by success (purple) and live for one lineage (amber), against rent. Bottom right: a fresh set's gain (bars, points) and the runs ending in a post (red, %). The slider is the set's age.
set made by
-
copy error removed by the set (live)
-
by a fresh set (live)
-
relevance r (live)
-
r, four lineages, copy error
-
r, success
-
success gain, set (fresh)
-
clone's frames within a bandwidth
-
set frames beyond 4 cm
-
clone's runs ending in a post
-
fresh set, points per dollar
-
Show the core JS
SL.draw = function (nw, seed, gust) { return SL.frames(SL.runs(nw, SL.NRUN, seed, gust), SL.NSET); };
SL.copyError = function (nw, probe) {
  var se = 0, ss = 0, i;
  for (i = 0; i < probe.length; i++) {
    var p = DL.predict(nw, probe[i][0]), a = probe[i][1];
    se += (p[0] - a[0]) * (p[0] - a[0]) + (p[1] - a[1]) * (p[1] - a[1]); ss += a[0] * a[0] + a[1] * a[1];
  }
  return Math.sqrt(se / ss);
};
SL.step = function (c) {
  var lin = c.lin, j = c.n - c.k, pr = SL.probe(lin, c.n, c.gj), d = c.done;
  var X = SL.draw(SL.clone(lin, j), SL.setSeed(lin.seed, c.n - 1, c.k, d, c.gi), SL.GUSTS[c.gi]);
...
  var e1 = SL.copyError(SL.nwOf(SL.stock(lin, c.n).concat(X)), pr.frames);
  c.red.push((pr.e0 - e1) / pr.e0); c.done++;

What to try. Leave the defaults: clone 4, a set made by clone 1, three clones back, gust 0.05. On this lineage the set removes 3.42 % of the copy error (± 0.63) against 2.29 % for a fresh set, r = 1.49; the four lineages give 1.83 ± 0.21 and, by success, 1.95 ± 0.89. 97.1 % of the clone's frames lie within a bandwidth of the set, and 7.0 % of the set's frames are beyond 4 cm against 1.7 % for a fresh set: the old set has the tail. At clone 6 and age 5 the table says 3.33 ± 0.71. Choose clone 1, age 0 and the world set 0.05, clone 0.10: r is 0.69 live, 0.70 in the table by the copy error and 0.81 by success, and coverage falls to 73.2 %. The right panel is what runs out: the bars fall from 18.5 to 0.85 points at a constant $0.99.

6 · What runs out, and the three shelf lives

If old corrections keep their value, why does buying them stop paying? Because a set is worth the failures in it (§4). A correction teaches what its policy got wrong, and a clone that completes the course nine times in ten makes few mistakes to teach. A run fails here only by ending in a post, and with independent runs the table follows from the share that do. A set costs $0.99 whatever its clone:

cloneruns ending in a postfailing runs in a round of fiverounds that show nonea fresh set adds (points)points per dollar
138.1 %1.909 %18.518.7
222.6 %1.1328 %9.89.9
313.1 %0.6650 %5.45.5
45.6 %0.2875 %0.850.86
54.2 %0.2181 %1.01.0
63.4 %0.1784 %0.720.73

The first set adds 18.5 points, the fourth 0.85: 22 times fewer per dollar at the same price, and at clone 4 three rounds in four show the clone no failure at all. A clone that fails one run in twenty lies between clone 4, one in 18, and clone 5, one in 24, and has little left to show. The stock holds what earlier clones got wrong and keeps it (§3): the first clone's set, added to clone 4, still adds 1.67 points, twice a fresh set's 0.85 though a tenth of the 18.5 it added to clone 1, because most of what clone 1 failed at is taught already. What is scarce is the failures nobody has corrected yet, and the clone now running is the only source of them. The stock does not expire and the supply runs out.

Sort the sources by what the distribution of their records depends on:

classrecordsthe distribution depends onafter a retrainingwhat is spent
capitaldemonstrations, force recordingsthe expert onlynothing changes: last year's set is a draw from the distribution of today'snothing; what a set adds depends on the stock it joins (below)
inventoryfootage, simulationthe world: a person's hand, the simulator's arm modelthe records do not change; they are worth lesson 16's rate, up to a ceiling: lesson 18's pools stop at 56.0 % (footage) and 82.3 % (the 1 % simulator) at 256 layoutsthe headroom: each layout is worth less than the last
correctionsrounds of DAgger, takeovers (HG-DAgger, Kelly et al., 2019)the policy that made themkept while the retrained policy stays among their states (§3, §4); a fifth lost when the gust doubles (§5)the supply: failures of the policy now running

Capital never goes stale and can still be worth less than nothing. In a stock that holds two or more rounds, the clone trained on the rounds alone beats the clone trained on the rounds and the 20 calm demonstrations (1000 runs in each of four lineages): at clone 3 96.9 % against 86.9 %, at clone 4 99.8 % against 94.4, at clone 5 99.7 % against 95.8, at clone 6 99.6 % against 96.6. Shelf life and marginal value are different columns of the ledger, as lesson 18's pools showed for sources. The ledger's corrections curve, 20 demonstrations plus rounds, levels off near 96 % in part because of them, and lesson 20 should read it as a floor for what corrections reach.

Road not taken · book the stock as rent and rebuy it
It is how a firm treats a perishable, and the entry gives the premise: a stored set was made in states an older policy visited. At q = 0.8 four rounds are worth 2.95 sets; replace the oldest first and rebuy at $0.99 a set. The Bench says what that buys. Keep the demonstrations and only the newest round and clone 5 falls from 95.8 % to 74.8 %, 21 points; keep only the oldest round and it is at 77.5 %, 18 points lost. Rebuying does not repair it: a fresh set added to clone 5 is worth 1.0 point. Rent is the right book for the one case measured, a world that moved outward, a fifth of $0.99 a set.
What this lesson did not do
The clone is a table of stored frames, so retraining is storing and nothing is forgotten; a network trained by gradient steps on a changing stock can lose old frames (lesson 14). The labeller is a program: a person's labels drift (lesson 4), a way for old sets to go stale that nothing here measures. The world moved in one way, a stronger gust; a different arm or posts is another outward move, for which §4's reach is the test, not a measured ratio. Footage and simulation were sorted by lesson 18's ceilings and not run under retraining. At age 1 the success cells cannot separate 0.8 from 1. The Bench's task is small (lesson 16): ratios and shapes carry, the points and the dollars do not. Who supplies failures when the policy is good is lesson 20.

Common mistakes / failure modes

"the policy changed, so its old corrections are stale"
By the copy error no cell of the fifteen is below 0.87, and the first clone's set is worth 1.83 fresh sets to clone 4 (§3).
"a better policy needs different corrections"
Its states are inside the old policy's: 99 % of clone 6's frames lie within 4.26 cm of the path, clone 1's within 4.90 (§4).
"keep only the newest round"
Clone 5 falls from 95.8 % to 74.8 % (§6).
"corrections never expire, so keep buying them"
The stock keeps its value, the supply does not: 18.5 points for the first set, 0.85 for the fourth (§6).
"a demonstration can only help"
At clone 4 the rounds alone reach 99.8 %, the rounds with the demonstrations 94.4 % (§6).

Checkpoint exercise

Try it
Clone 4 ends in a post on 5.6 % of its runs and the runs are independent; a round is five runs and a set costs $0.99. (a) How many runs until the first failure, on average? (b) What is the chance that a round shows none? (c) What does one failing run found cost at clone 4, and at clone 1 (38.1 %)? Answer: (a) 1 / 0.05625 = 17.8 runs, 3.6 rounds. (b) (1 − 0.05625)5 = 74.9 %. (c) A round holds 5 × 0.05625 = 0.281 failing runs, so a failing run costs $0.989 / 0.281 = $3.52 at clone 4 and $0.989 / 1.90 = $0.52 at clone 1, 6.8 times dearer. The stock has not lost value; the next failure has become dear to find.

Where this points next

Retraining does not spoil a correction. Per frame, the first clone's set is worth 1.83 ± 0.21 times a fresh set to the fourth clone (1.67 points against 0.85 by success) and no cell of fifteen is below 0.87, because the expert's label is a function of the state and a better policy stays among the states of a worse one. Value goes when the world moves outward, 0.81 kept when the gust doubles, and a record that the expert made and no policy did keeps its value whoever is trained. What runs out is the supply: a fresh set adds 18.5 points to the first clone and 0.85 to the fourth, 22 times fewer per dollar, and at the fourth clone 75 % of rounds show no failure. The cheapest continuing supply of corrections is the deployed robots themselves, working, with a person stepping in when they fail. Does a fleet like that improve the policy fast enough to pay for the people who supervise it?

Takeaway
The value of an old correction set to a retrained clone, against a fresh set of the same size, is its relevance r, and on the Bench it does not fall with the set's age: by the copy error it is between 0.87 and 3.33 in fifteen cells. The expert's label is a function of the state and a better policy visits a subset of the states of a worse one, so an old set keeps covering the new clone and holds more of the tail. Value is lost when the retrained policy goes where the set was never made: 0.81 kept when the gust doubles, 0.92 when the clone is replaced sideways, a conditional rent of $0.19 on a $0.99 set. What runs out is the supply: a fresh set adds 18.5 points to the first clone and 0.85 to the fourth, and three rounds in four show the fourth clone no failure. Sort sources by what their records depend on: the expert (capital, which never goes stale and can still dilute a stock), the world (inventory, which saturates) and the policy (corrections, whose supply thins).

Interview prompts

Companion reads: Lesson 3 · Labels on your own states, Lesson 4 · People are not functions and Reinforcement Learning · 17 Imitation and inverse RL.