The policy writes its own test set
The clone of lesson 1 copies the expert to 3.4 % on frames it has never seen, makes four times that error on the frames it produces itself, and fails more often the longer the course. This lesson puts a ruler on the difference. Every frame is measured by its distance from the nearest stored demonstration. The expert's held-out frames sit on top of the data and the clone's own frames do not, and the clone's error is a curve in that distance, not a property of the function. Nothing in the data teaches the clone to come back once it has left, so each slip stays, and the steps lost to slips grow as the square of the length of the task. More demonstrations of the same kind leave all of it where it was. What would change it is the next lesson's question.
New idea: a policy's error depends on how far a frame is from its data, the policy produces its own frames, and nothing in the data teaches it to return, so slips are never undone and their cost grows as the square of the horizon. The missing ingredient is in the data, not in the learner: the states the clone visits have no labels.
Forces next: A policy's mistakes land it in states its training data never showed, where its next action is a guess, so one slip is charged for every step that follows: the cost grows as the square of the horizon unless the loop pulls the policy back or the training data include the states the policy reaches. More demonstrations recorded the same way leave the curve where it is. The only one who knows what to do in a state the policy has just wandered into is the expert. Can the expert label the states the policy visits, and what does that cost?
1 · A ruler for frames
Lesson 1 scored the clone in two places and the scores disagreed. To see why, give every frame one more number: how far it is from what the clone was built from. The clone is a table of stored frames, so the natural ruler is the distance from a frame to the nearest stored one, in the learner's own unit, the bandwidth, 0.02 rad, about 1.5 cm at the cup at the start of the course. Call it d. The kernel reaches three bandwidths; beyond d = 3 the clone copies one stored frame, and the last bin below takes every frame from 2 up.
Three groups of frames, five bins of d. The held-out expert frames are what lesson 1 scored: ten fresh runs of the expert. The expert's frames under gust are 100 runs of the expert itself under the gusts the clone will meet. The clone's own frames are its 200 runs, five posts, 20 calm demonstrations, gust 0.05, as in lesson 1.
| distance from the data, d (bandwidths) | held-out expert frames | expert's frames under gust | clone's own frames |
|---|---|---|---|
| under 0.5 | 99.6 % | 91.1 % | 46.7 % |
| 0.5 to 1 | 0.4 % | 8.5 % | 29.8 % |
| 1 to 1.5 | 0.1 % | 0.4 % | 15.5 % |
| 1.5 to 2 | 0.0 % | 0.0 % | 6.1 % |
| 2 or more | 0.0 % | 0.0 % | 1.9 % |
Where the expert goes is where the data are: 99.6 % of the held-out frames lie within half a bandwidth of a stored frame. The gusts alone do not change that. The expert under gust is within one bandwidth in 99.6 % of its frames, because it is the controller that pulls a run back. The clone's own frames are spread: 46.7 % within half a bandwidth and 23.4 % at least one bandwidth away, against 0.1 % of the held-out frames and 0.4 % of the expert's. The held-out frames are a test set that lies on the training set. The clone's own frames are a test set the clone wrote itself, by acting.
2 · The error is a function of the distance
Score the clone's command against the expert's at the same state, as in lesson 1, bin by bin. The expert's command at a state the clone has reached is available, because the expert is a function of the joint angles; this is only possible on the Bench.
| distance from the data (bandwidths) | mean distance in the bin | copy error of the clone's own frames |
|---|---|---|
| under 0.5 | 0.23 | 6.0 % |
| 0.5 to 1 | 0.73 | 12.5 % |
| 1 to 1.5 | 1.21 | 19.6 % |
| 1.5 to 2 | 1.71 | 26.2 % |
| 2 or more | 2.26 | 34.2 % |
The error climbs by 14 points for every bandwidth of distance, along a line (r² = 0.999). The reason is what a kernel answer is: the average of commands that the expert gave at stored places near the query. A command given at another place is a copy of something meant for somewhere else, and the further the query is from every stored place the less it is meant for here. The nearest-frame fallback is the extreme case.
Now average the curve over where each group's frames are. The held-out expert frames sit at its left end, at a mean distance of 0.02 bandwidths, where their error is 3.1 % (the line gives 2.9 % there; over all of them it is 3.4 %). The clone's own frames lie along all of it. Errors are root-mean-square, so the bins combine as squares: the square root of the share-weighted mean of the squared bin errors is 13.6 %, against the 13.9 % measured (the exact weight of a bin is its share of the expert's squared commands, within a few percent of its share of frames). The ratio of 4.1 in lesson 1 is not a worse function on the clone's frames. It is the same function, evaluated further from where it was fitted.
3 · Nothing pulls it back
A clone slightly off its data would be harmless if it drifted back. Does it? Measure how d changes from one step to the next: for every frame in a bin, the mean of d(t+1) − d(t). The fraction of its distance that a run recovers per step is the pull-back p = −(mean change of d) / (mean d). The first bin cannot be read this way (d cannot be negative, so it drifts up whatever happens), so take the bin from 0.5 to 1 bandwidth.
| in the bin 0.5 to 1 bandwidth | steps measured | pull-back per step p | loop gain J = 1 − p |
|---|---|---|---|
| the expert under gust | 1568 | 7.5 % | 0.925 |
| the clone | 8966 | −0.1 % | 1.001 |
The expert recovers 7.5 % of its distance per step. The clone recovers nothing: it moves 0.1 % of its distance outward per step. The expert is a feedback controller: it aims at the point 6 cm ahead on its path, so a displacement produces a command that cancels part of it. The clone has a copy of what the expert did on its own path and almost nothing else. Its demonstrations were recorded calm, from starts within 1 cm of the path, and contain no displaced frames, so near the edge of the data the kernel returns the commands of on-path frames, which point along the path. Nothing it stored says "if you are 2 cm to the left, go right". The clone has the expert's commands and not the expert's feedback.
What that does is a one-line recursion. Let et be the displacement from the data after t steps and ηt what one step adds, the gust and the clone's own error:
et+1 = J · et + ηt so et = Σk Jk ηt−k, spread = η · √((1 − J2t) / (1 − J²))
A disturbance of age k is multiplied by Jk. With J < 1 the spread settles at η / √(1 − J²); with J = 1 it grows as η√t without limit. The gust alone moves each joint by 0.05 rad/s × 0.05 s = 0.0025 rad per step, 0.125 bandwidths. With the expert's J the spread settles at 0.33 bandwidths, 0.49 cm. With the clone's it reaches 1.25 bandwidths, 1.9 cm, after 100 steps, and keeps going. The clone is an integrator: a velocity command is added to the joint angles, and whatever is not corrected is kept. Every gust stays, and a clone that has gone d away is worse at its job there (§2), so its own error adds to the gust.
4 · A slip that nothing undoes
What does a slip cost if nothing undoes it? Take a run that must stay within 4 cm of the expert's path and must not touch a post. The expert under the same gusts does: it loses 0.0 of its 235 steps on five posts. Count a step as lost if the cup has touched a post, is farther than 4 cm from the path, or never reached the mat, and let C(T) be the expected number of the first T steps that are lost. Suppose each step carries a chance h of a slip that is never undone. A slip at step s costs the T − s steps after it, so
P(lost by step t) = 1 − (1 − h)t ≈ h t, C(T) = Σt ≤ T (1 − (1 − h)t) ≈ h T² / 2 (h T small)
Slips are equally likely at every step, so a slip costs on average T/2 steps and there are hT of them. If the model is right, h = 2C/T² is the same at every course length, and C(T) is a straight line of slope 2 on log–log axes. A slope of 1 would mean a slip costs the same however long the course; slope 2 says it costs in proportion to the steps left. Courses of one to six posts, 20 calm demonstrations, gust 0.05:
| posts | steps allowed T | reaches the mat | steps lost C(T) | share of steps lost | h = 2C/T² |
|---|---|---|---|---|---|
| 1 | 65 | 87.5 % | 5.2 | 8.0 % | 2.48 × 10⁻³ |
| 2 | 107 | 83.0 % | 14.7 | 13.7 % | 2.56 × 10⁻³ |
| 3 | 150 | 75.5 % | 29.5 | 19.6 % | 2.62 × 10⁻³ |
| 4 | 193 | 63.5 % | 45.0 | 23.3 % | 2.42 × 10⁻³ |
| 5 | 235 | 63.5 % | 60.5 | 25.7 % | 2.19 × 10⁻³ |
| 6 | 278 | 42.0 % | 107.5 | 38.7 % | 2.78 × 10⁻³ |
Both predictions hold. h stays between 2.19 and 2.78 × 10⁻³, within 13 % of its mean of 2.51 × 10⁻³, and the fitted slope of ln C against ln T across the six courses is 2.00 (r² = 0.994). About one step in 399 carries a slip the clone does not recover from; a five-post course loses 25.7 % of its steps and a six-post course 38.7 %. A caution: the hazard is not exactly constant, because a displacement random-walks for a while before it crosses a margin, and for large hT the exact sum bends below hT²/2. Over these lengths the square is what is measured.
The widget
What to try. Leave the defaults: five posts, 20 calm demonstrations, gust 0.05. The clone reaches the mat in 63.5 % of runs, its copy error is 3.4 % on expert frames and 13.9 % on its own, and 23.4 % of its own frames are at least one bandwidth from the data. In the 0.5 to 1 bandwidth bin the expert pulls back 7.5 % of its distance per step and the clone −0.1 %. It loses 60.5 of its 235 steps, 2.19 × 10⁻³ per step. Slide the course to three posts: 29.5 of 150 steps lost, 2.62 × 10⁻³; to six: 107.5 of 278, 2.78 × 10⁻³. The hazard barely moves while the cost nearly quadruples, and the exponent readout, once the six courses are in, says 2.00. Return to five posts and quadruple the demonstrations to 80: nothing moves (success 63.5 %, 65.2 steps lost, the clone's pull-back still 0.1 % outward). With 5 demonstrations success falls to 51.5 %: below some number the clone has not seen the course, above it more do nothing. Now record the 20 demonstrations under gusts of 0.05: success 81.5 %, the clone's pull-back 2.7 %, steps lost 31.8; under gusts of 0.10: success 97.5 %, pull-back 14.5 %, 6.7 steps lost. The first copy error readout rises meanwhile, 3.4 % to 4.9 % to 7.0 %: by lesson 1's scorecard the clone got worse as it got better. Finally put the demonstrations back to calm and move the run-time gust: at 0 the clone reaches the mat 96.0 % of the time and loses 9.5 steps; at 0.08 it reaches it 31.0 % of the time and loses 114.1.
5 · What more data can change
The pull-back is what the clone is missing, and it is present only where the data contain displaced frames labelled with the expert's response. The widget's own experiments sort the kinds of data by that test:
| 20 calm demonstrations, unless stated | reaches the mat | copy error, expert frames | copy error, own frames | own frames beyond 1 bandwidth | clone's pull-back | steps lost |
|---|---|---|---|---|---|---|
| 5 calm | 51.5 % | 3.4 % | 16.9 % | 28.8 % | −0.2 % | 86.2 |
| 20 calm | 63.5 % | 3.4 % | 13.9 % | 23.4 % | −0.1 % | 60.5 |
| 80 calm | 63.5 % | 3.3 % | 14.6 % | 23.8 % | −0.1 % | 65.2 |
| recorded under gusts of 0.05 | 81.5 % | 4.9 % | 11.9 % | 5.4 % | 2.7 % | 31.8 |
| recorded under gusts of 0.10 | 97.5 % | 7.0 % | 8.1 % | 0.2 % | 14.5 % | 6.7 |
More of the same. Four times the calm demonstrations (80, against 20) leave every column where it was: they repeat the same thin band of frames, so they add copies of the same commands and no displaced frames. The loss is not a shortage of data of this kind. Data of another kind. Demonstrations recorded under gusts include displaced frames, each labelled with the expert's response, and the clone learns a pull-back from them: 2.7 % per step at gust 0.05 and 14.5 % at 0.10, the second from only 954 steps in a bin that holds 2.6 % of the clone's frames. The frames beyond one bandwidth fall from 23.4 % to 0.2 %, and the success rises to 97.5 %. A different loop. If the command were an absolute position tracked by a stiff controller, J would be below 1 whatever the data: a displacement is corrected by the controller, not by anything the clone learned.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
A clone's own frames are a different set from the expert's, and the clone's error is a curve in the distance from its data: 13.9 % on its own frames against 3.4 % on the expert's. Nothing pulls a displaced clone back (−0.1 % per step, against 7.5 % for the expert), so every slip stays, and the five-post course loses 60.5 of its 235 steps, with the loss growing as the square of the horizon. Eighty calm demonstrations change none of it (63.5 % success). Demonstrations recorded under gusts change it by supplying displaced frames, at the price of noisier labels and a noise level somebody has to guess, and the frames they supply are the expert's, not the ones the clone reaches by its own errors. The only one who knows what to do in a state the policy has just wandered into is the expert. Can the expert label the states the policy visits, and what does that cost?
Interview prompts
- A cloned policy has the same function on its training frames and on its own, yet three to four times the error on the second. How can that be? (§1, §2 — the error is a curve in the distance from the data, and the policy's own frames lie further out on it.)
- Why does the cost of compounding errors grow as T² and not T? (§4 — a slip that is never undone costs the T − s steps after it; slips are equally likely at every step, so the cost per slip averages T/2 and there are hT of them.)
- How would you measure, from rollouts, whether a policy pulls itself back toward its data? (§3 — the mean change of the distance to the nearest training frame per step, per bin of distance; the loop gain J = 1 − p.)
- Why do 80 calm demonstrations do no better than 20? (§3, §5 — they repeat the same band of frames and contain no displaced states, so the pull-back stays zero.)
- Why can demonstrations recorded under noise give a better policy and a worse validation error at once? (§5 — they contain displaced frames labelled with the expert's response, which teach a pull-back; the commands to copy now vary more from frame to frame, so the copy error on expert frames rises.)
- A policy has 1 % training error and finishes 67 % of its runs. Contradiction? (§2 — no: training error is measured on the training frames, success on the frames the policy produces, which lie further from them.)
- In a loop et+1 = J et + η, what does J = 1 mean for the spread of e? (§3 — it grows as η√t without limit; with J < 1 it settles at η / √(1 − J²).)
Companion reads: World Models · 21 Teacher forcing to rollout (the same compounding in a world model's own rollouts), World Models · 06 Errors compound (the horizon at which a model's predictions stop being trusted) and Reinforcement Learning · 17 Imitation and IRL (the quadratic bound and its linear repair from the MDP side).