all_lessons/Robot Model Training/02 · Closed looplesson 2 / 24

The policy writes its own test set

The clone of lesson 1 copies the expert to 3.4 % on frames it has never seen, makes four times that error on the frames it produces itself, and fails more often the longer the course. This lesson puts a ruler on the difference. Every frame is measured by its distance from the nearest stored demonstration. The expert's held-out frames sit on top of the data and the clone's own frames do not, and the clone's error is a curve in that distance, not a property of the function. Nothing in the data teaches the clone to come back once it has left, so each slip stays, and the steps lost to slips grow as the square of the length of the task. More demonstrations of the same kind leave all of it where it was. What would change it is the next lesson's question.

The thesis, here
A clone's error is not one number. It is a curve in the distance from the data: 3.1 % at the data, 34.2 % at 2.3 bandwidths away. The expert's held-out frames sit at the left end of the curve. The clone's own frames, moved by its errors and the gusts, sit along all of it. The expert pulls itself back when displaced and the clone does not, because nothing it was shown was displaced; so every slip stays, a slip that stays costs every step after it, and the lost steps grow as the square of the horizon.
Linear position
Forced by: A policy cloned from twenty demonstrations copies the expert to within a few percent on frames it has never seen, makes about four times that error on the frames it produces itself, and fails more than a third of the time on a course of five posts and more than half the time on six. The two kinds of frame come from different places: the first from where the expert went, the second from where the policy goes. Why is a policy that is accurate where the expert went lost where it goes, and how fast does the loss grow with the length of the task?
New idea: a policy's error depends on how far a frame is from its data, the policy produces its own frames, and nothing in the data teaches it to return, so slips are never undone and their cost grows as the square of the horizon. The missing ingredient is in the data, not in the learner: the states the clone visits have no labels.
Forces next: A policy's mistakes land it in states its training data never showed, where its next action is a guess, so one slip is charged for every step that follows: the cost grows as the square of the horizon unless the loop pulls the policy back or the training data include the states the policy reaches. More demonstrations recorded the same way leave the curve where it is. The only one who knows what to do in a state the policy has just wandered into is the expert. Can the expert label the states the policy visits, and what does that cost?
The plan
Five moves. (1) Put a ruler on frames: the distance from the nearest stored demonstration. (2) Show that the error is a function of that ruler and that the two kinds of frame sit at different places on it. (3) Ask whether a displaced policy comes back, for the expert and for the clone. (4) Derive what a slip that nothing undoes costs, and measure how it grows with the length of the task. (5) Ask what more data, and data of another kind, can change.

1 · A ruler for frames

Lesson 1 scored the clone in two places and the scores disagreed. To see why, give every frame one more number: how far it is from what the clone was built from. The clone is a table of stored frames, so the natural ruler is the distance from a frame to the nearest stored one, in the learner's own unit, the bandwidth, 0.02 rad, about 1.5 cm at the cup at the start of the course. Call it d. The kernel reaches three bandwidths; beyond d = 3 the clone copies one stored frame, and the last bin below takes every frame from 2 up.

Three groups of frames, five bins of d. The held-out expert frames are what lesson 1 scored: ten fresh runs of the expert. The expert's frames under gust are 100 runs of the expert itself under the gusts the clone will meet. The clone's own frames are its 200 runs, five posts, 20 calm demonstrations, gust 0.05, as in lesson 1.

distance from the data, d (bandwidths)held-out expert framesexpert's frames under gustclone's own frames
under 0.599.6 %91.1 %46.7 %
0.5 to 10.4 %8.5 %29.8 %
1 to 1.50.1 %0.4 %15.5 %
1.5 to 20.0 %0.0 %6.1 %
2 or more0.0 %0.0 %1.9 %

Where the expert goes is where the data are: 99.6 % of the held-out frames lie within half a bandwidth of a stored frame. The gusts alone do not change that. The expert under gust is within one bandwidth in 99.6 % of its frames, because it is the controller that pulls a run back. The clone's own frames are spread: 46.7 % within half a bandwidth and 23.4 % at least one bandwidth away, against 0.1 % of the held-out frames and 0.4 % of the expert's. The held-out frames are a test set that lies on the training set. The clone's own frames are a test set the clone wrote itself, by acting.

2 · The error is a function of the distance

Score the clone's command against the expert's at the same state, as in lesson 1, bin by bin. The expert's command at a state the clone has reached is available, because the expert is a function of the joint angles; this is only possible on the Bench.

distance from the data (bandwidths)mean distance in the bincopy error of the clone's own frames
under 0.50.236.0 %
0.5 to 10.7312.5 %
1 to 1.51.2119.6 %
1.5 to 21.7126.2 %
2 or more2.2634.2 %

The error climbs by 14 points for every bandwidth of distance, along a line (r² = 0.999). The reason is what a kernel answer is: the average of commands that the expert gave at stored places near the query. A command given at another place is a copy of something meant for somewhere else, and the further the query is from every stored place the less it is meant for here. The nearest-frame fallback is the extreme case.

Now average the curve over where each group's frames are. The held-out expert frames sit at its left end, at a mean distance of 0.02 bandwidths, where their error is 3.1 % (the line gives 2.9 % there; over all of them it is 3.4 %). The clone's own frames lie along all of it. Errors are root-mean-square, so the bins combine as squares: the square root of the share-weighted mean of the squared bin errors is 13.6 %, against the 13.9 % measured (the exact weight of a bin is its share of the expert's squared commands, within a few percent of its share of frames). The ratio of 4.1 in lesson 1 is not a worse function on the clone's frames. It is the same function, evaluated further from where it was fitted.

Road not taken · a smoother learner
The clone is a table, and a table is the plainest case: it knows nothing beyond its entries. A smooth fit, random Fourier features with ridge regression, answers smoothly everywhere. On the same 20 demonstrations, 300 features and a length scale of 0.1 rad, it has a copy error of 0.9 % on its training frames and reaches the mat in 67.0 % of runs under the same gusts. The numbers move a great deal with the settings (over five settings, 100 to 1000 features and length scales of 0.05 to 0.2 rad, the training error is at most 2.3 % and the success runs from 25.5 to 79.0 %), and in none does it come near the expert's 100 %: the disagreement between the two scores stays.

3 · Nothing pulls it back

A clone slightly off its data would be harmless if it drifted back. Does it? Measure how d changes from one step to the next: for every frame in a bin, the mean of d(t+1) − d(t). The fraction of its distance that a run recovers per step is the pull-back p = −(mean change of d) / (mean d). The first bin cannot be read this way (d cannot be negative, so it drifts up whatever happens), so take the bin from 0.5 to 1 bandwidth.

in the bin 0.5 to 1 bandwidthsteps measuredpull-back per step ploop gain J = 1 − p
the expert under gust15687.5 %0.925
the clone8966−0.1 %1.001

The expert recovers 7.5 % of its distance per step. The clone recovers nothing: it moves 0.1 % of its distance outward per step. The expert is a feedback controller: it aims at the point 6 cm ahead on its path, so a displacement produces a command that cancels part of it. The clone has a copy of what the expert did on its own path and almost nothing else. Its demonstrations were recorded calm, from starts within 1 cm of the path, and contain no displaced frames, so near the edge of the data the kernel returns the commands of on-path frames, which point along the path. Nothing it stored says "if you are 2 cm to the left, go right". The clone has the expert's commands and not the expert's feedback.

What that does is a one-line recursion. Let et be the displacement from the data after t steps and ηt what one step adds, the gust and the clone's own error:

et+1 = J · et + ηt so et = Σk Jk ηt−k, spread = η · √((1 − J2t) / (1 − J²))

A disturbance of age k is multiplied by Jk. With J < 1 the spread settles at η / √(1 − J²); with J = 1 it grows as η√t without limit. The gust alone moves each joint by 0.05 rad/s × 0.05 s = 0.0025 rad per step, 0.125 bandwidths. With the expert's J the spread settles at 0.33 bandwidths, 0.49 cm. With the clone's it reaches 1.25 bandwidths, 1.9 cm, after 100 steps, and keeps going. The clone is an integrator: a velocity command is added to the joint angles, and whatever is not corrected is kept. Every gust stays, and a clone that has gone d away is worse at its job there (§2), so its own error adds to the gust.

4 · A slip that nothing undoes

What does a slip cost if nothing undoes it? Take a run that must stay within 4 cm of the expert's path and must not touch a post. The expert under the same gusts does: it loses 0.0 of its 235 steps on five posts. Count a step as lost if the cup has touched a post, is farther than 4 cm from the path, or never reached the mat, and let C(T) be the expected number of the first T steps that are lost. Suppose each step carries a chance h of a slip that is never undone. A slip at step s costs the T − s steps after it, so

P(lost by step t) = 1 − (1 − h)t ≈ h t, C(T) = Σt ≤ T (1 − (1 − h)t) ≈ h T² / 2 (h T small)

Slips are equally likely at every step, so a slip costs on average T/2 steps and there are hT of them. If the model is right, h = 2C/T² is the same at every course length, and C(T) is a straight line of slope 2 on log–log axes. A slope of 1 would mean a slip costs the same however long the course; slope 2 says it costs in proportion to the steps left. Courses of one to six posts, 20 calm demonstrations, gust 0.05:

postssteps allowed Treaches the matsteps lost C(T)share of steps losth = 2C/T²
16587.5 %5.28.0 %2.48 × 10⁻³
210783.0 %14.713.7 %2.56 × 10⁻³
315075.5 %29.519.6 %2.62 × 10⁻³
419363.5 %45.023.3 %2.42 × 10⁻³
523563.5 %60.525.7 %2.19 × 10⁻³
627842.0 %107.538.7 %2.78 × 10⁻³

Both predictions hold. h stays between 2.19 and 2.78 × 10⁻³, within 13 % of its mean of 2.51 × 10⁻³, and the fitted slope of ln C against ln T across the six courses is 2.00 (r² = 0.994). About one step in 399 carries a slip the clone does not recover from; a five-post course loses 25.7 % of its steps and a six-post course 38.7 %. A caution: the hazard is not exactly constant, because a displacement random-walks for a while before it crosses a margin, and for large hT the exact sum bends below hT²/2. Over these lengths the square is what is measured.

The widget

Watch a clone leave its data
Top: the table seen from above. Grey, the demonstrations; each of 12 of the clone's 200 runs is coloured by its distance from the data at every step (teal under 0.5 bandwidth, amber 0.5 to 1.5, red beyond). Bottom left: where the frames of the expert's held-out runs (teal) and the clone's own runs (amber) fall, with the clone's copy error in each bin. Bottom right: steps lost against steps allowed over the six course lengths on log–log axes; the amber point is the current course. The first control is the course length. The readouts include the pull-back of the expert and the clone in the 0.5 to 1 bandwidth bin, with the number of steps behind each.
clone reaches the mat
—
copy error, expert frames
—
copy error, clone's own frames
—
own frames at least 1 bandwidth from the data
—
pull-back per step, expert
—
pull-back per step, clone
—
steps lost, C(T)
—
hazard h = 2C/T², × 10⁻³ per step
—
exponent of C against T, courses 1 to 6
—
steps allowed, T
—
Show the core JS
LL.dist = function (nw, q) {
  var best = Infinity;
  nw.each(q, function (id) { var e = nw.dist2(id, q); if (e < best) best = e; });
  return best > 9 ? 3 : Math.sqrt(best);
};
LL.bin = function (d) { var b = 0; while (b < LL.NB - 1 && d >= LL.EDGES[b + 1]) b++; return b; };
...
    for (t = 0; t + 1 < d.length; t++) { var bb = bins[LL.bin(d[t])]; bb.dd += d[t + 1] - d[t]; bb.ds += d[t]; bb.m++; }
...
    out.bins.push({ n: b.n, m: b.m, share: tot ? b.n / tot : 0, err: withErr && b.ss > 0 ? Math.sqrt(b.se / b.ss) : NaN, drift: dr, meanD: dm, pull: b.m ? -dr / dm : NaN });
...
  var C = BN.costCurve(ev.rolls, ev.T, LL.TUBE);
  return { C: C, T: ev.T, total: C[ev.T], hazard: 2 * C[ev.T] / (ev.T * ev.T), succ: ev.succ };

What to try. Leave the defaults: five posts, 20 calm demonstrations, gust 0.05. The clone reaches the mat in 63.5 % of runs, its copy error is 3.4 % on expert frames and 13.9 % on its own, and 23.4 % of its own frames are at least one bandwidth from the data. In the 0.5 to 1 bandwidth bin the expert pulls back 7.5 % of its distance per step and the clone −0.1 %. It loses 60.5 of its 235 steps, 2.19 × 10⁻³ per step. Slide the course to three posts: 29.5 of 150 steps lost, 2.62 × 10⁻³; to six: 107.5 of 278, 2.78 × 10⁻³. The hazard barely moves while the cost nearly quadruples, and the exponent readout, once the six courses are in, says 2.00. Return to five posts and quadruple the demonstrations to 80: nothing moves (success 63.5 %, 65.2 steps lost, the clone's pull-back still 0.1 % outward). With 5 demonstrations success falls to 51.5 %: below some number the clone has not seen the course, above it more do nothing. Now record the 20 demonstrations under gusts of 0.05: success 81.5 %, the clone's pull-back 2.7 %, steps lost 31.8; under gusts of 0.10: success 97.5 %, pull-back 14.5 %, 6.7 steps lost. The first copy error readout rises meanwhile, 3.4 % to 4.9 % to 7.0 %: by lesson 1's scorecard the clone got worse as it got better. Finally put the demonstrations back to calm and move the run-time gust: at 0 the clone reaches the mat 96.0 % of the time and loses 9.5 steps; at 0.08 it reaches it 31.0 % of the time and loses 114.1.

5 · What more data can change

The pull-back is what the clone is missing, and it is present only where the data contain displaced frames labelled with the expert's response. The widget's own experiments sort the kinds of data by that test:

20 calm demonstrations, unless statedreaches the matcopy error, expert framescopy error, own framesown frames beyond 1 bandwidthclone's pull-backsteps lost
5 calm51.5 %3.4 %16.9 %28.8 %−0.2 %86.2
20 calm63.5 %3.4 %13.9 %23.4 %−0.1 %60.5
80 calm63.5 %3.3 %14.6 %23.8 %−0.1 %65.2
recorded under gusts of 0.0581.5 %4.9 %11.9 %5.4 %2.7 %31.8
recorded under gusts of 0.1097.5 %7.0 %8.1 %0.2 %14.5 %6.7

More of the same. Four times the calm demonstrations (80, against 20) leave every column where it was: they repeat the same thin band of frames, so they add copies of the same commands and no displaced frames. The loss is not a shortage of data of this kind. Data of another kind. Demonstrations recorded under gusts include displaced frames, each labelled with the expert's response, and the clone learns a pull-back from them: 2.7 % per step at gust 0.05 and 14.5 % at 0.10, the second from only 954 steps in a bin that holds 2.6 % of the clone's frames. The frames beyond one bandwidth fall from 23.4 % to 0.2 %, and the success rises to 97.5 %. A different loop. If the command were an absolute position tracked by a stiff controller, J would be below 1 whatever the data: a displacement is corrected by the controller, not by anything the clone learned.

Road not taken · noisier demonstrations
Recording the demonstrations under gusts looks like the answer, and on this table it nearly is: 97.5 % at gusts of 0.10. It has three costs. The displacements it covers are the expert's under gust, and the clone's displacements come from the clone's own errors, which are not the same ones. The right amount of noise has to be guessed, and it was twice the gust the clone then met. And the old scorecard gets worse: the copy error on expert frames, recorded the same way, rises from 3.4 % to 7.0 %, the opposite of what the clone does in the loop. It returns in lesson 3 as noise injection and is compared there with labelling the states the clone reaches.
What this lesson did not do
It measured the loop on one learner (a table of stored frames, with a check on a smooth fit) and a two-number observation; a policy that reads pictures has far more room between its data and its inputs, so the picture here is the optimistic one. It counted a slip as a collision or a departure from a 4 cm tube; real tasks have slips that cost more or less, and recoveries that cost something. The constant-hazard model is a first-order account; the Bench's hazard rises a little with time. It did not say how to put the states the clone reaches into the data or what that costs (lesson 3), what a policy should output when the right action is not a single one (lesson 4), or how to change the loop gain with the form of the action (lessons 5 and 6).

Common mistakes / failure modes

"the clone is accurate, so it stays close to the expert"
23.4 % of its own frames lie at least one bandwidth from the data, against 0.1 % of the held-out expert frames (§1).
"its errors are small and random, so they average out"
With no pull-back (J = 1.001) disturbances are summed, not averaged: the spread grows as the square root of the steps, 1.25 bandwidths after 100 (§3).
"the cost of errors grows with the horizon"
It grows with the square of it: the slope of ln C against ln T is 2.00 across the six courses (§4).
"more demonstrations will close the gap"
20 to 80 calm demonstrations: success 63.5 % to 63.5 %, steps lost 60.5 to 65.2 (§5).
"a smoother learner avoids it"
A fit with 0.9 % training error completes 67.0 % of runs (§2).
"the gust is what breaks the clone"
The expert under the same gusts loses 0.0 steps; what the clone lacks is the pull-back, −0.1 % against 7.5 % (§3).
"noise in the demonstrations is just noise"
Recorded under gusts they teach a pull-back and lift success to 81.5 % and 97.5 %, while the old copy error rises to 4.9 % and 7.0 % (§5).

Checkpoint exercise

Try it
A policy has a chance 0.001 per step of a slip it never recovers from. (a) To first order, how many of the first 100 steps does it lose on average? (b) And of the first 200? (c) The plant's gust is 0.04 rad/s per joint (a bandwidth is 0.02 rad, a step 0.05 s). How far from its data, in bandwidths, does a policy with no pull-back wander after 100 steps, and where does a policy with J = 0.925 settle? Answer: (a) hT²/2 = 0.001 × 100² / 2 = 5 steps (the exact sum gives 4.9). (b) 0.001 × 200² / 2 = 20 steps: doubling the horizon quadruples the cost and doubles the share lost, from 5 % to 10 %. (c) One step adds 0.04 × 0.05 / 0.02 = 0.1 bandwidth. With no pull-back the displacement is a random walk, 0.1 × √100 = 1.0 bandwidth and still growing. With J = 0.925 it settles at 0.1 / √(1 − 0.925²) = 0.26 bandwidth, 3.8 times smaller, and stops growing.

Where this points next

A clone's own frames are a different set from the expert's, and the clone's error is a curve in the distance from its data: 13.9 % on its own frames against 3.4 % on the expert's. Nothing pulls a displaced clone back (−0.1 % per step, against 7.5 % for the expert), so every slip stays, and the five-post course loses 60.5 of its 235 steps, with the loss growing as the square of the horizon. Eighty calm demonstrations change none of it (63.5 % success). Demonstrations recorded under gusts change it by supplying displaced frames, at the price of noisier labels and a noise level somebody has to guess, and the frames they supply are the expert's, not the ones the clone reaches by its own errors. The only one who knows what to do in a state the policy has just wandered into is the expert. Can the expert label the states the policy visits, and what does that cost?

Takeaway
A policy writes its own test set: the frames it sees next are the ones its actions produce. Measured by the distance from the nearest stored frame, the expert's held-out frames sit on the data (99.6 % within half a bandwidth) and the clone's own do not (23.4 % at least one bandwidth away), and the clone's copy error rises by 14 points per bandwidth of distance, which turns 3.4 % on the data into 13.9 % on its own frames. The expert recovers 7.5 % of a displacement per step and the clone nothing, because its demonstrations contain no displaced frames, so the clone is an integrator: gusts and errors stay. A slip that stays costs every step after it, so the lost steps grow as the square of the horizon (slope 2.00 across one to six posts; one step in 399 carries such a slip). More calm demonstrations change nothing; displaced frames with the expert's response change a lot, at a price.

Interview prompts

Companion reads: World Models · 21 Teacher forcing to rollout (the same compounding in a world model's own rollouts), World Models · 06 Errors compound (the horizon at which a model's predictions stop being trusted) and Reinforcement Learning · 17 Imitation and IRL (the quadratic bound and its linear repair from the MDP side).