all_lessons/Robot Model Training/22 · Effective hourslesson 22 / 24

Effective hours: duplicates, quality and collapse

Lesson 21 spent a budget by reading curves of success against hours bought, every hour a layout the policy had not seen, and a real corpus counts recorded episodes. On the Bench 80 layouts recorded 32 times each are 2,560 episodes and 6.61 hours, and the policy trained on them reaches what 80 layouts reach once, 84.4 %. This lesson defines the size of a dataset by what the policy reaches, in the unit of lesson 16, and tests three discounts against it: counting layouts and not episodes; a batch recorded with an offset that passes every check made inside it and is worth less than nothing; and a generator trained on its own samples, which forgets its rarest layouts first. It does not say what it costs to move the hours that survive.

The thesis, here
The size of a dataset is what the policy can use: the number of your own layouts, each demonstrated once, that give it the same success. A count of recorded episodes overstates it by the repeat factor, a count of different layouts predicts it for a policy that stores what it is shown, a batch with an offset can make it smaller than the same dataset without that batch, and a generator trained on its own samples lowers it by removing the rare layouts first. Kish's formula for unequal weights, applied to the episodes per layout, orders what a learner that averages reaches and does not size what a learner that stores reaches.
Linear position
Forced by: Spending by marginal value per dollar gives a sequence and not a mixture: cheap breadth first, the scarce column when the cheap source saturates, and the instrument for that column before its price rank says so. The plan counts hours as if an hour were an hour of information. A hundred thousand recorded episodes can hold eight hundred distinct situations, a policy trained on its own samples loses its tail, and a bad batch looks like a good one. What is the effective size of a dataset?
New idea: the effective size of a dataset is the number of your own layouts, demonstrated once each, that give the policy the same success, and each loss between that number and the recorded count can be computed from the corpus: repeats, a batch's offset and what a loop forgets. The ledger's curves count hours of different layouts, which is to say effective hours.
Forces next: Counting effective hours instead of recorded ones shrinks a corpus, sometimes by orders of magnitude, and the hours that survive still have to reach the accelerators: a model that trains on video spends its time decoding it. If the processors wait for the data, the data budget is not the only bill. What does it cost to move the effective hours, and where does the machine stall?
The plan
Six moves. (1) Record the same layouts again and see what the policy gets. (2) Define a size by what the policy reaches, and test the counts against it on five corpora. (3) Ask what a duplicate does to a learner that averages. (4) Record a batch with an offset, and check it against something the tracker did not record. (5) Train a generator on its own samples. (6) Price the hours that survive.

1 · The same layouts, recorded again

The layouts curves of lessons 16 to 21 count different layouts: the ledger's own curve counts own layouts, each demonstrated once, and a source adds attempted layouts, each a new one. A corpus need not be built that way. Take the first 80 layouts of your own file and record each m times; the expert runs the layout again from a start about a centimetre off, so every recording is a new file of the same situation. The policy is the lookup of lesson 7: it copies the stored layout nearest to a new one, here with every repeat of that layout in its table, and success is the share of 1000 new layouts it completes. A recorded hour is 387.1 demonstrations of 9.3 s.

repeats per layout mepisodes recordedmotion hourssuccess
1800.2184.4 %
21600.4184.4 %
43200.8384.7 %
86401.6584.5 %
161,2803.3184.3 %
322,5606.6184.5 %

Thirty-two times the hours buy the same policy: 84.3 to 84.7 % for every m, which is well inside the sampling error of the 1000 test layouts. A filter that drops identical files has nothing to drop, because the start offset makes all 2,560 files different. At real scale duplicates are measured as well: near-duplicate filtering removes 3.04 % of the examples of C4 and 13.63 % of those of RealNews, and one 61-word sequence occurs 61,036 times in the C4 training set (Lee et al., 2022). The baton's corpus has the same ratio on a larger scale: 100,000 episodes of 9.3 s are 258 hours and 800 situations are 2.07 hours, a factor of 125.

At lesson 17's assumed price of $123.6 for a motion hour of your own robot, the 6.61 hours cost $817 and bought what 0.207 hours, $25.5, would have bought: the hours that did any work cost $3,956 each. The plan of lesson 21 reads its curves at the hours it has bought, and for this corpus the number to read at is 0.207. What is the size of a dataset?

2 · A size defined by what the policy reaches

Four constraints. The size has the unit of lesson 16, an own layout demonstrated once; it is measured on the policy and not asserted; it is tied to the task, the policy and the test layouts; and a planner needs it before the money is spent, so it must be predictable from the corpus. The first three give a definition. Let S be the success of the policy trained on the dataset on the 1000 test layouts, and Neq(S) the number of own layouts that reach S on the ledger's own curve, read off by the inversion of lesson 16 (its interval is the Wilson interval of S pushed through it):

neff = Neq(S)

For the corpus of §1 it is 80 layouts (73 to 86). The fourth constraint asks for a count that predicts it, and five corpora give the candidates something to disagree about. In a field corpus the episodes are recorded where the world presents them: layout c has probability pc proportional to rank−a, the ranks assigned at random over the 256 layouts of your file (a = 0: every layout alike; a = 1: a long tail; a = 2: a few layouts take most episodes), and R episodes are drawn. Write R for the episodes recorded, kc for those of layout c, D for the number of layouts with kc > 0, and nK = (Σ kc)² / Σ kc² for Kish's formula applied to the counts. One draw of each corpus:

corpusepisodes Rdifferent layouts DKish nKmeasured size neffR / neff
80 layouts, 32 repeats each2,5608080.080 (73–86)32.0
field, a = 0806152.565 (60–72)1.2
field, a = 132010519.092 (86–101)3.5
field, a = 22,560622.458 (53–62)44.5
field, a = 210,2401072.595 (89–109)107.7

Recorded episodes overstate the size by a factor from 1.2, where every layout is alike, to 107.7, where a few dominate; files are no better, since every recording is a different file. Kish's formula on the counts falls below the measured size in every row but the first, by a factor of 5 at a = 1 and 24 to 39 at a = 2. The number of different layouts comes within a fifth: over eight draws of each corpus the measured size is 0.80 to 0.96 of D at a = 1 and R = 320 and 0.81 to 1.14 at a = 2 and R = 2,560 (0.81 to 1.38 at R = 10,240, where success passes 95 % and the own curve flattens). The intervals of the table carry the sampling error of the 1000 test layouts, and the spread over draws the luck of which layouts were drawn.

The reason is in the policy. For each test layout it copies one demonstration of the nearest stored layout, so the number of repeats behind a layout never enters (§1): success depends on which layouts are present and where they lie, and a layout recorded once and one recorded a hundred times are the same entry in the table. The count that matches is the number of entries. Exact repeats are found by a hash of the layout; layouts rebuilt by hand to within a few millimetres are clusters of layouts closer than the distance within which a stored layout serves a new one, which this lesson does not run.

Kish (1965) defined (Σ w)² / Σ w² for the precision of a weighted mean: n for equal weights, never more, and a measure of the loss from unequal weights, not of the number of different units. Applied to episodes per layout it is this lesson's construct, with one exact reading: Σ (kc / R)² is the probability that two episodes drawn at random are of the same layout, so nK is its inverse (the checkpoint works it by hand). It is 80 for the first corpus, which cannot tell it from D.

3 · A duplicate is also a weight

The lookup cannot be moved by counts, by construction; a learner trained on a mean loss can. Over the recorded episodes i, the loss Σi Ki (y − yi)², with Ki a kernel of the distance between the new layout and the layout of episode i, is minimised by the weighted average y = Σ Ki yi / Σ Ki, and a layout recorded kc times enters it kc times: a duplicate is up-weighted in the loss and not merely wasted. Run that learner on the Bench: for each test layout take the 8 nearest stored layouts, weight each by kc exp(−d² / 2h²) with h = 1.5 cm (d is how far the farthest post moves between the two layouts, as in lesson 7), average their hand paths and follow the average. The same 80 layouts and 2,560 episodes, with the counts drawn at exponent a (ranges over eight draws):

exponent alayouts presentKish nKchance two episodes are of one layoutlookupaverage of neighbours
08078.31.3 %84.4 %93.4 %
18014.86.8 %84.4 %86.6 % (83.2–89.6)
2502.540.6 %70.9 % (69.9–79.4)59.3 % (54.1–74.7)

With equal counts the average interpolates and beats the lookup, 93.4 % to 84.4 %. At a = 1 every layout is still present and the lookup is where it was, while the average loses 6.8 points: the popular layouts pull the average at every neighbouring test layout toward their paths. At a = 2, 30 layouts are never recorded, the lookup loses what they covered, and the average loses 34.1 points. Kish's formula orders the three corpora as the averaging learner does, 78.3 > 14.8 > 2.5 against 93.4 > 86.6 > 59.3 %, and sizes neither learner: it names no number of layouts that either is worth. For a mean-loss learner a duplicate costs twice, as hours that taught nothing and as weight taken from the rare layouts; for the lookup only the first.

Road not taken · the size as Kish's formula over episodes per layout
It is standard, takes one line, needs no training, equals the number of layouts when the counts are equal and falls as they become unequal, as duplicates make them. On the Bench it names the wrong quantity: 19.0 where the policy that stores layouts measured 92, 2.4 where it measured 58; it ranks the learner that averages and does not size it; and it is nearly blind to the tail, since in §5 a corpus loses 60 % of its layouts while nK moves from 22 to 18. What survives is the collision probability, the number for how often a run that draws episodes at random sees the same layout twice.

4 · A batch that passes its own checks

Quality enters as a weight, and the weight can be negative. Let a rig have recorded a share of the layouts with a hand tracker offset by 1.5 cm sideways: every recorded hand target of its layouts is shifted, the layouts it states are right, and every recording succeeded in the operator's frame. In the corpus of §1 take 25 % of the layouts, 20 of 80. Nothing inside the batch shows the offset. None of its recordings touches a post of its layout (the closest comes within 5.1 cm of one, against a contact distance of 5 cm, and 6.0 cm unshifted; at 2 cm 34 of the 80 would touch), its hours look like any batch's, and the spread of its clearance to the posts (the sideways distance from the recorded path to a post where the hand passes it), 0.8 mm, is the rest's 0.7 mm: an offset that every recording shares cannot show in how much recordings differ.

The policy shows it. Keeping the batch gives 76.4 % and a size of 59 layouts (55 to 64), where the clean corpus gave 84.4 % and 80. Dropping the 20 layouts gives 80.8 % and 69, so each layout of the batch is worth (59.5 − 69.3) / 20 = −0.49 own layouts (−0.53 to −0.01 over the eight choices below). A lookup trusts its labels: 15.7 % of the test layouts have one of the rig's layouts as their nearest stored layout, and the arm then follows a path displaced by 1.5 cm; it completes 40.8 % of them, against 68.8 % when the rig's layouts are dropped and the next nearest layout serves (the dilution of lesson 18, by quality instead of by source). Over eight choices of which layouts the rig recorded, keeping gives 73.1 to 77.8 % and dropping 76.5 to 79.1 %, so dropping wins every time; with 10 % of the layouts on the rig the figures are 79.8 and 82.7 %.

What does see it is held-out consistency: a comparison with something the tracker did not record. The posts a layout states are known without the hand, and the expert passes them at a clearance that does not depend on the layout (at least 7.0 cm sideways). Compare the clearance of the rig's layouts with the rest's: the difference of means is 1.49 cm, z = 76 standard errors (Welch), so the batch is flagged and its offset estimated, 1.49 cm against the 1.5 cm applied. Dropping the flagged batch gives 80.8 %; subtracting the estimated offset gives 84.4 %, the clean figure. At these batch sizes the check flags offsets down to 0.59 mm (z = 3); the Bench's expert differs by under a millimetre between recordings, and a noisier demonstrator would raise the threshold in proportion to that spread, over the square root of the batch size.

5 · A generator trained on its own samples

A model that generates episodes can feed the next corpus. Take the simplest generator: it memorises what it was trained on and emits R episodes with the shares of the previous corpus, pc = kc / R; the policy trains on the new corpus, and the next generation starts from it. A layout with k of the R episodes is missing from the next corpus with probability (1 − k / R)R, about e−k: a layout recorded once survives a generation with probability 63.3 % (a simulation gives 63.6), and the layouts left after one generation number Σc [1 − (1 − kc / R)R] in expectation, 79.2 for the corpus at a = 1 and R = 320 that has 105 layouts (79.6 over 300 generations drawn). The diversity 1 − Σ pc² is multiplied by 1 − 1 / R per generation in expectation, 0.9876 over eight generations at R = 640: the measure that the head dominates hardly moves while the layouts with small counts, the tail, go extinct. Median of eight draws, a = 1 and R = 640:

generationslayouts presentsuccessoriginal episodes still representeda tenth of the original kept: layoutssuccess
015196.1 %100 %15196.1 %
112592.8 %95 %12592.7 %
210789.7 %90 %11290.9 %
48384.9 %83 %9587.9 %
86175.0 %74 %8686.1 %

The tail goes first. After eight generations 60 % of the layouts are gone and success on the uniform test has fallen 21.1 points (71.7 to 78.1 % over the draws), while 74 % of the original episodes still have their layout present and nK has moved from 22 to 18: a check that averages over recorded episodes sees less than half of the loss, because the layouts that go are the rare ones. Keeping a random tenth of the original corpus in every generation leaves 86 layouts, 57 % of the original, and 86.1 %: the loss is slowed and not stopped, since the original holds only the layouts it recorded. Shumailov et al. (2024) show for language models, variational autoencoders and Gaussian mixtures that training on a model's own output makes the tails of the original distribution disappear, and in their OPT-125m runs a random tenth of the original data kept in each generation gave only minor degradation; they tested no robot data.

The widget

Record a corpus, corrupt part of it, recycle it, and read its size
Left: the box of layouts, shift across and tilt up. A dot is one of the 1000 test layouts, teal when the policy trained on this corpus completes it and red when not; a ring is a stored layout, larger the more episodes were recorded of it (log scale), amber when the rig recorded it and dashed when corrected; a grey cross is a layout that is gone, dropped as flagged or forgotten by the generator. Right: hours on a log scale (recorded, different layouts, Kish's formula, and the measured size with its interval). The first slider sets the episodes recorded. The widget keeps one demonstration of each stored layout, which the table of §1 shows costs nothing.
episodes recorded
-
recorded hours
-
different layouts
-
Kish's formula
-
measured size, own layouts [95 %]
-
effective hours
-
success on 1000 new layouts
-
recorded per effective layout
-
cost per effective hour
-
layouts forgotten or dropped
-
original episodes still represented
-
rig check, z
-
estimated offset
-
Show the core JS
ES.score = function (off) {
...
  for (t = 0; t < NT; t++) {
    bd = Infinity; bc = -1;
    for (c = 0; c < NS; c++) { if (off[c] !== off[c]) continue; d = DM[t * NS + c]; if (d < bd) { bd = d; bc = c; } }
    pick[t] = bc; if (bc >= 0) { okv[t] = outcome(t, bc, off[bc]); ok += okv[t]; }
  }
...
ES.generation = function (cnt, cnt0, R, fresh, seed) {
  var nf = Math.round(fresh * R), a = ES.draw(ES.share(cnt, R), R - nf, seed), b, i;
  if (nf > 0) { b = ES.draw(ES.share(cnt0, R), nf, seed + 500); for (i = 0; i < NS; i++) a[i] += b[i]; }
  return a;
};
ES.neq = function (s) { return LG.ownEquivalent(LG.ownCurve('hand'), s); };

What to try. Leave the defaults, 80 scripted layouts and 2,560 episodes: 6.61 hours recorded, 80 layouts, Kish 80.0, a measured size of 80 layouts (73 to 86), success 84.4 %, so 32 episodes per effective layout and $3,956 per effective hour. Move the slider to 80 and to 10,240 episodes: success stays 84.4 and 84.4 % while recorded hours go to 26.5 and the ratio to 128. Take the field corpus with a = 2 at 10,240 episodes: 107 layouts, Kish 2.5, a size of 95 (89 to 109), 26.5 hours recorded for 0.25 effective. Put 25 % of the scripted layouts on the rig: 76.4 % kept (size 59), 80.8 % dropped, 84.4 % corrected, at z = 76. On the field corpus with a = 1 and 640 episodes (150 layouts, 97.1 %) raise the generations to 8: 60 layouts, 75.5 %, 76 % of the original episodes still represented; keep a tenth of the original and it is 86 layouts and 86.1 %. On the scripted corpus the same eight generations keep 79 of the 80 layouts and 84.3 %: 32 episodes per layout make extinction slow, not impossible. These are single seeded draws; the table of §5 gives medians.

6 · The ledger line, and what finding it costs

The ledger has a price per motion hour (lesson 17), a rate per hour of a source (lesson 16) and a curve of value against hours (lesson 18). This lesson puts a factor in front of all three, the effective fraction e = neff / R. The scripted corpus has e = 1 / 32 (0.207 of 6.61 hours); the field corpus of §2 with a = 2 and 10,240 episodes has 1 / 107.7, and over eight draws 1 / 65 to 1 / 112. Every price of the ledger is paid per recorded hour and returns per effective one, so an effective hour costs the price over e: $3,956 against $123.6 for the scripted corpus. A plan that reads its curves at effective hours buys correctly; one that reads them at recorded hours buys 32 to 108 times too much.

The effective hours are fewer than the recorded ones, and finding them is a pass over the recorded ones. A duplicate is found by comparing an episode with others, a batch's offset by comparing the batch with posts recorded elsewhere, a generator's loss by comparing with the original corpus, and each of the three reads every recorded episode. The 80 layouts that survive are cheap to keep; the 2,560 episodes around them were stored, decoded and compared to find them, and a model trained on the whole corpus decodes every frame of every episode in every epoch.

What this lesson did not do
The policy is a lookup, and the learner that averages was run once, with one bandwidth: the size of a trained network, with its capacity and optimiser, is not measured, here or in a later lesson. Repeats were exact; layouts rebuilt by hand to a few millimetres need a distance below which two layouts count as one, and were not run. The generator memorises; a smoother one invents layouts of its own, and the Bench did not run one. The offset is one kind of bad batch, a constant shift checked against posts the Bench records without the tracker; noise, wrong labels and a wrong body (the older arm and the footage of lesson 16 are weights of that kind) need other references. Which episodes to drop when a budget forces a choice, by difficulty as in the pruning of Sorscher et al. (2022) for image classifiers, was not run. The Bench task is small, 80 to 256 layouts in under an hour of motion: the ratios carry and the totals do not (lesson 18). What it costs to move the hours that survive is lesson 23, and lesson 24 reads the ledger with this column in it.

Common mistakes / failure modes

"more recorded hours buy a better policy"
Thirty-two times the hours left success between 84.3 and 84.7 % (§1), and the size per recorded hour was 1 / 107.7 in the field corpus of §2.
"Kish's formula is the effective size"
It gives 2.4 where the policy measured 58 layouts; it ranks an averaging learner's success and sizes neither (§2, §3).
"duplicates are only wasted hours"
For a learner that averages they are weight taken from the rest: 93.4 % with equal counts and 59.3 % with a few dominating (§3).
"a batch that checks out inside itself is good"
Its spread is 0.8 mm against the rest's 0.7, yet keeping it costs 8.0 points and each of its layouts is worth −0.49 (§4).
"my validation on recorded episodes will warn me of collapse"
After eight generations 74 % of the original episodes are still represented while 60 % of the layouts are gone (§5).

Checkpoint exercise

Try it
A corpus has 600 episodes: 300 of one layout, and one episode each of 300 other layouts. (a) How many episodes and how many different layouts? (b) What is Kish's nK, and what is the probability that two episodes drawn at random are of the same layout? (c) A training run draws 20 episodes at random with replacement: how many different layouts does it see on average? Answer: (a) 600 and 301. (b) nK = 600² / (300² + 300) = 360,000 / 90,300 = 3.99, the inverse of the collision probability 25.1 %: half the episodes are the heavy layout, (1/2)² = 25 %, and the 300 singles add 300 × (1/600)² = 0.08 %. (c) The heavy layout is seen with probability 1 − 0.520, which is 0.999999, and each single one with 1 − (599/600)20 = 0.0328, so 1 + 300 × 0.0328 = 10.8 layouts in 20 draws: half of the draws land on the heavy layout.

Where this points next

Counting effective hours shrinks the corpus. The 2,560 episodes of §1 hold 80 layouts, a factor of 32 and, at the baton's scale, 125; a field corpus of 10,240 episodes whose popularity falls as the square of the rank is worth 95 layouts, 65 to 112 episodes recorded per effective layout over eight draws. The layouts that survive are cheap to keep, but finding them reads every recorded episode: a duplicate is found by comparing episodes, a bad batch by comparing it with another, a generator's lost tail by comparing with the original, and a policy that trains on video decodes the video it reads, in all 2,560 episodes, every epoch. The hours that survive still have to reach the accelerators, and if the processors wait for the data, the data budget is not the only bill. What does it cost to move the effective hours, and where does the machine stall?

Takeaway
The size of a dataset is the number of your own layouts, demonstrated once each, that give the policy the same success. Recorded episodes overstate it by the repeat factor, 32 for 80 layouts recorded 32 times and up to 107.7 where a few layouts dominate, and files do not help, since every recording is a different file. For a policy that stores layouts the count of different layouts predicts the size within a fifth where the own curve is steep; Kish's formula on the counts does not, though it orders the success of a learner that averages, for which a duplicate is also a weight. A batch recorded with an offset passes every check made inside it, can be worth less than nothing (−0.49 layouts each at 1.5 cm), and is found by comparison with something the tracker did not record. A generator trained on its own samples forgets the rare layouts first, 60 % of them in eight generations, while 74 % of the original episodes are still represented. The ledger's hours are effective hours, and finding them reads the recorded ones.

Interview prompts

Companion reads: Lesson 7 · How much data: the coverage law, Lesson 8 · Other bodies and World Models · 27 The mixture is the model.