all_lessons/Robot Model Training/03 · Own-state labelslesson 3 / 24

Labels on your own states: DAgger and corrections

Lesson 2 left a clone that copies the expert to 3.4 % on frames it has not seen, is lost on the frames it makes itself, and loses steps as the square of the horizon however many demonstrations it has. This lesson puts the clone's own frames into the data: it runs, the expert says what it should have done at every state it visited, those frames join everything stored, and the clone is rebuilt. That is DAgger. Four rounds of five runs take it from 63.5 % to 95.6 % on five posts (mean of eight seeds). The labels are the price: a person gives them late when replaying a run, scarcely when taking over, and two people asked from the same pose can send the cup either side of a post.

The thesis, here
A clone fits the frames it is given, the expert's (copy error 3.4 %), while the robot is scored on its own frames, where the same function errs by 13.9 %. DAgger gives it its own frames, labelled by the expert, and keeps everything: the share of its own frames a bandwidth or more from its data falls from 23.4 % to 0.1 %, it recovers from a displacement as the expert does, and it finishes about nineteen runs in twenty. What it cannot repair is a labeller who is not a function of the state.
Linear position
Forced by: A policy's mistakes land it in states its training data never showed, where its next action is a guess, so one slip is charged for every step that follows: the cost grows as the square of the horizon unless the loop pulls the policy back or the training data include the states the policy reaches. More demonstrations recorded the same way leave the curve where it is. The only one who knows what to do in a state the policy has just wandered into is the expert. Can the expert label the states the policy visits, and what does that cost?
New idea: the training set should be the states the policy itself visits, labelled by the expert, with nothing already stored thrown away. It buys the pull-back that lesson 2 found missing; it costs labelled frames and a labeller who can answer for states it did not drive, which a program can and a person only in part.
Forces next: Training on the policy's own states, with the expert's actions as labels, takes a policy that finishes five posts about six times in ten to one that finishes them about nineteen times in twenty after four rounds of five rollouts, and a person supplies those labels best by taking over when the policy goes wrong. But people are not functions: asked from the same pose, or asked by two operators, one goes left of the post and another right, and a regression asked to fit both returns their average, a route that runs into the post. What should a policy output when the right action is not one point?
The plan
Five moves. (1) What the training set should be, and what that asks of the labeller. (2) Three ways to buy displaced frames at the same number of labelled frames. (3) The loop on the Bench: what it changes, what to keep, whether the rounds matter. (4) What a label costs when a person gives it. (5) Two people.

1 · What the training set should be

Lesson 2 found why the two scores of lesson 1 disagree: the clone's error is a curve in the distance from its data, and its own frames sit far out along it. A learner can only fit the frames it is given, so what matters is which frames. The theory agrees. Ross and Bagnell (2010) bound the extra cost of a clone trained on the expert's frames by T²ε, with T the horizon in steps and ε the clone's error on the expert's states; it is an upper bound, tight only for constructed examples, and on the Bench ε is the 3.4 % of lesson 1. Ross, Gordon and Bagnell (2011) move the measurement: with εN the loss, after N rounds, on the states the learner's own policies visit, the extra cost is bounded by about TεN, linear in T. The two ε are different numbers: for the first clone, 3.4 % on the expert's states and 13.9 % on its own. The move changes not how well the function fits but where it is asked to fit.

Four constraints follow. (a) Only the expert knows what to do in a state the clone has wandered into, so the labels come from the expert. (b) The clone's cost is paid on its own states, so the states should come from the clone. (c) A label is a query: count labelled frames and, if a person answers, person-seconds at 0.05 s a frame. (d) The clone's states depend on the clone, which depends on the data, so they cannot be listed in advance; the clone has to be run to find them. That makes a loop. Round 0 stores the demonstrations, the expert driving. Each later round runs the current clone for five rollouts under the gust, the expert labels every state they visited, the labelled frames are added to everything stored, and the clone is rebuilt (for a table of stored frames, rebuilt means stored). This is DAgger (Ross, Gordon and Bagnell, 2011) in the variant they report often does best in practice: the expert drives the first round only.

2 · Three ways to buy displaced frames

Lesson 2 found what the clone lacks: displaced frames, each labelled with the expert's response. A fair comparison fixes what is spent, the number of labelled frames. Four rounds of five rollouts leave 7,104 frames stored in the default run, 3,684 of demonstrations and 3,420 labelled in the loop. The alternatives add whole demonstrations until at least that many are stored.

where the frames come fromwhat the expert doesclone reaches the matdemonstrations in which the expert touches a post
39 calm demonstrationsdrives, as in lesson 159.0 %0 %
demonstrations under a gust of 0.05 rad/s, the gust the clone meetsdrives while the plant is shaken; each label is its own command78.5 %0 %
under a gust of 0.10, twice thatthe same, shaken harder95.5 %4 %
under a gust of 0.15, three timesthe same, harder again99.0 %36 %
the clone's own states, four rounds of five rolloutslabels what the clone did, never drives95.0 %none

More of the same does nothing, as lesson 2 found. Injecting noise is a method in its own right (Laskey et al., 2017, DART): noise goes into the supervisor's control while it demonstrates, which forces it to show how to recover from errors. On the Bench the equivalent is a gust on the plant while the expert drives, so the demonstrations contain displaced states with the recovery labelled. What sets its strength is the level, not the frames: 78.5 % at the gust the clone meets (81.5 % with the 20 demonstrations of lesson 2), 95.5 % at twice it and 99.0 % at three times, which is above the loop at every number of frames the widget below offers. Even 3,851 frames, fewer than the loop holds after one round, reach 98.5 % at twice the gust.

Road not taken · noise injection
On this Bench, with the level chosen well, it beats the loop, and the lesson has to say so. What it asks is not free. The level has to be two or three times the gust the clone will meet, which is the number you would like to know; DART fits it to the error of a learner trained on earlier demonstrations. The supervisor is the one pushed around: the expert touches a post in 4 % of its own demonstrations at 0.10, 36 % at 0.15 and 75 % at 0.20, and in DART's human trials the stronger of two levels did worse (72 % against 79 %), which its authors suggest may have been too much for the supervisor. And the Bench is generous: two joints and a smooth expert, where shaking in every direction covers the displacements the clone's slips make; on DART's higher-dimensional simulated tasks plain isotropic noise did not do well. The loop never pushes the expert and has no level to guess: the displacements are the ones the clone makes. Its price is the robot's own mistakes, 5.2 of every 20 collection runs ending in a post (§3), which DART's authors count, with the tedium of the corrections, among the costs of the loop.

3 · The loop, measured

Each round runs the clone five times and labels every state those runs visited, about 807 frames (a run ends early when it touches a post). The widget opens on the typical run: of eight collection seeds (eight draws of the collection runs' gusts and starts), the one whose success curve lies closest to their average. A step is lost if the cup has touched a post, is farther than 4 cm from the path or never reached the mat, and steps lost is the expected number of the first T = 235 steps that are lost (lesson 2's C(T)). The ruler is lesson 2's: how far each of the clone's own frames is from the nearest stored frame, in kernel bandwidths.

roundlabelled framesreaches the matsteps losterror on own framesown frames at least 1 bandwidth from datapull-back per step
0063.5 %60.513.9 %23.4 %−0.1 %
183478.5 %40.811.2 %1.6 %7.8 %
21,71387.0 %24.69.6 %0.5 %7.6 %
43,42095.0 %10.98.3 %0.06 %10.4 %

Every measure of lesson 2 moves as the argument says it should. The clone's own frames end up on its data. The pull-back lesson 2 found missing is there after one round, 7.8 % of a displacement recovered per step in this run (4.2 % over the eight seeds) against the expert's 7.5 %, because the data now hold displaced states labelled with the expert's response. The error on the clone's own frames, the number training now minimises, falls from 13.9 to 8.3 %, though not to the expert-frame 3.4 %. Over eight collection seeds the success at round 4 is 95.6 % (the seeds range from 93.5 to 97.5 %), about nineteen runs in twenty, and it creeps up to 98.4 % by round 8.

Keep everything. A round could replace the stored frames instead of adding to them. After four rounds, averaged over eight seeds, keeping everything reaches 95.6 %, keeping the demonstrations and only the newest round 76.7 %, and keeping only the newest round 3.6 %. The ruler says why: the share of the clone's own frames a bandwidth or more from its data is 0.1 % keeping everything, 6.4 % once the earlier rounds are dropped, and 54 % once the expert's own path is dropped too, when the data hold only the states the last clone visited.

Do the rounds matter? Take all the rollouts from the first clone instead and rebuild once. Twenty rollouts reach 96.2 % against 95.6 % for four rounds of five, forty rollouts 98.4 % against 98.4 % for eight rounds of five. The first clone is the worst driver on the Bench, so its runs wander widest, which would explain why its batch already covers where the later, better clones go. Labelled frames count; rounds do not. What the rounds buy is a robot that crashes less while the data are collected: 8.4 of the twenty runs of the first clone end in a post against 5.2 for retrained clones, 16.9 of forty against 5.6. The theory asks for rounds because each new clone can go where the last one did not; the check is the ruler, stopping when almost none of the clone's own frames is a bandwidth from its data (0.1 % at round 4) and success has stopped rising.

Does it bend the square? Not on this Bench. From one to six posts the loop lowers the steps lost by a factor of 4.0 at one post and 4.3 at six, and the exponent reads 1.94 ± 0.27 (one standard error) against lesson 2's 2.00, which these data cannot tell apart: a touch is permanent, so a small hazard per step is still summed over every step after it. The linear bound concerns a loss measured on the learner's own states; the Bench does not show the curve turning linear, and this lesson does not claim it.

4 · What a label costs when a person gives it

Four rounds label 3,228 frames on average. A person who answers one control step at a time spends 0.05 s on each: 161 s. That is the bill if the person is as good as the expert, and the expert is not a person: it is a program, a function of the joint angles, so it answers for any state at once and exactly. A person falls short of the program in three ways: imprecise, late, and not single-valued; the first two are here, the third in §5. To price them the Bench needs a model, and the simplest has one number, a reaction time τ of 0.3 s, 6 frames. It is an assumption, not a measurement of people, and the table varies it.

Imprecise. Add noise of 40 % of the command's own size to every label, 12 times the clone's copy error, and the clone still reaches the mat in 94.6 % of runs (95.6 with exact labels). Noise 1.6 times the command's own size costs 8.2 points. A regression answers with the mean of the labels near a query (lesson 1), and noise that cancels in a mean costs almost nothing.

Late. A person replaying a run steers along it and sees no effect of what they do, so their command at frame t answers the state they saw τ frames earlier. Kelly et al. (2019) name the mechanism: asked for labels without being in control, a person's labels are likely to degrade through perceived actuator lag. On the expert's own runs a 0.3 s lag makes the label wrong by 42.5 % of the command, about as much as the 40 % noise. But the clone answers with the average of nearby labels, and on those frames that average keeps 97 % of the lag's error against 10 % of the noise's: neighbouring frames carry nearly the same lag error, so averaging removes almost none of it.

Taking over. The remedy is to put the person in control. In HG-DAgger (Kelly et al., 2019) the clone drives, the person takes control when they judge it necessary and keeps it until they hand it back, and only the frames the person drove are labelled. On the Bench the person takes over when the cup is more than 2.5 cm from the path, after τ frames, drives until it is within 1 cm, and hands back. The 2.5 cm is not free: the clone's failures cross it a median of 12 frames before the touch, 90 % of them more than the 6 frames a person needs to react, but cross 4 cm only a median of 3 frames before, 15 % of them more than 6. In control the person sees what each command does, so the labels are exact (our assumption again).

after four rounds, unless statedreaches the mat, mean (lowest to highest of 8 seeds)labelled framescollection runs ending in a post, of 20
the program labels every frame95.6 % (93.5 to 97.5)3,2285.2
labels with 40 % noise94.6 % (91.5 to 97.5)3,2864.9
a person replaying, 0.1 s late96.3 % (92.0 to 100)3,4283.8
a person replaying, 0.3 s late79.1 % (58.5 to 90.0)3,1669.0
a person replaying, 0.4 s late6.6 % (2.0 to 14.0)2,58913.8
a person taking over, 0.3 s72.3 % (66.5 to 76.5)1,2090.8
taking over, 0.3 s, ten rounds79.2 % (72.5 to 84.5)3,0721.2 of 50

A person supplies labels best by taking over in one sense: it is the only way whose labels do not depend on how late the person is (70.9, 72.3 and 67.7 % at 0, 0.3 and 0.5 s, against 96.3, 79.1 and 6.6 % for a replay at 0.1, 0.3 and 0.4 s), and it keeps the robot out of the posts while the data are collected (0.8 of 20 collection runs end in one, against 3.8 to 13.8 for a replay). It is steadiest, not most accurate: a replay with a short lag does better. Kelly et al. (2019) report faster, more stable learning than DAgger in a driving simulation. On the Bench it is safer but not faster: it labels only the frames the person drove, so the person's tolerance bounds the displacements the data ever show, and ten rounds reach 79.2 % where the program's 3,228 frames reached 95.6 %.

The widget

Label what the clone did
Dots are the stored frames (grey: demonstrations, purple: earlier rounds, amber: the newest round); lines are 12 of the 200 runs of the clone after the chosen round (cyan reached the mat, red touched a post). Bottom left: success against frames stored, and, if chosen, another way of buying the same frames; the amber point is the slider's round. Bottom right: the start frame, with the heading A gives (teal), B (purple, dashed) and the clone (amber); A and B are the two operators of §5, A passing the first post above and B below. One collection seed runs here, the typical one of eight; the tables of §4 and §5 average the eight.
clone reaches the mat
—
95 % interval
—
frames labelled so far
—
person-seconds at 20 labels a second
—
collection runs ending in a post
—
steps lost, C(T)
—
error on the clone's own frames
—
own frames at least 1 bandwidth from data
—
pull-back per step
—
runs ending at the first post
—
clone's heading at the start frame
—
same frames bought the other way
—
Show the core JS
DL.collect = function (s) {
...
        if (mode === 0 && dev > DL.TOL) { mode = 1; cnt = 0; }
        if (mode === 1) { cnt++; if (cnt > tau) mode = 2; }
        if (mode === 2 && dev < DL.BACK) mode = 0;
        if (mode === 2) { var a = BN.slalom.expertAct(wA, q); fresh.push([q, a]); return a; }
        return DL.predict(nw, q);
...
      var op = s.who === 'two' && s.pick() < 0.5 ? wB : wA;
      for (t = 0; t < ro.S.length; t++) fresh.push([ro.S[t], BN.slalom.expertAct(op, ro.S[Math.max(0, t - s.tau)])]);
...
  fresh.forEach(function (f) { s.X.push(f[0]); s.Y.push(f[1]); });

What to try. Leave the defaults: round 0, the program labels, the latest clone drives. That is lesson 2's clone: 63.5 % reach the mat (interval 56.6 to 69.9 %), 60.5 of 235 steps are lost, the error on its own frames is 13.9 %, and 23.4 % of them are a bandwidth or more from the data. Slide to round 4: 95.0 %, 10.9 steps lost, error 8.3 %, 0.06 % far from the data, pull-back 10.4 % a step against the expert's 7.5 %, from 3,420 labelled frames, with 4 of the 20 collection runs ending in a post. Under compare with, choose demonstrations under gust 0.10: 95.5 % from the same frames. Under runs are made by, choose the first clone: 95.5 % at round 4, with 7 of 20 runs ending in a post. Under who labels, choose a person replaying the run: 84.5 % at round 4 with the reaction time left at 0.3 s, 2.0 % at 0.4 s; choose a person taking over: 76.0 % from 1,223 labelled frames. Last, choose two people: 24.0 % at round 0, the clone's command at the start frame heading 4.8° against 37.1° for A; 0.0 % at round 4, with 99 % of the runs ending at the first post.

5 · Two people

A person is also not a function of the state. Asked twice from the same pose, or by two operators, one goes above the first post and the other below, and both are right: the course has room for either. On the Bench operator A passes the first post above it (the expert of every lesson so far; above means larger y) and operator B below it; the posts and the 5 cm margin are the same. Give the clone ten demonstrations of each, from the same ten starting poses, as many frames as before (3,682).

At the start frame the cup velocity A commands points 37.1° above the x axis and B's points −37.1°, as far below it. The clone's answer is the kernel-weighted mean of the commands stored near the start frame, and the mean of two mirror-image commands points almost along the axis, 4.8°. Followed in a straight line to the column of the first post, 20 cm ahead, A's heading arrives 15.1 cm above the post's centre, outside its 5 cm halo; the clone's arrives 1.7 cm from it, inside. A regression minimises the squared difference to both, and for squared error the best single answer is the mean (lesson 1): a route that runs into the post.

the clone's data (mean of eight seeds)reaches the matruns ending at the first post
one operator, 20 demonstrations63.5 %13.0 %
two operators, 10 + 10 demonstrations24.0 %18.0 %
after one round, each run labelled by A or B6.8 %62.8 %
after two rounds0.3 %91.1 %
after four rounds0.0 %92.6 %

At round 0 the first post ends 18.0 % of the runs and the later posts another 58.0 %: the two routes are mirror images about the line of posts, so their average is the straight line through every post. The loop does not repair it; it makes it worse. Each collection run is labelled by one of the two operators, correctly for that operator. The clone, now heading for the post, visits states head-on to it, and each of those gets a label from above or from below. The mean of the labels is still the straight line into the post, and every round adds more frames that say so. DAgger assumed that the labeller is a function of the state, and the commands of two people are not.

What this lesson did not do
It used one learner, a table of stored frames, on a two-number state; a policy that reads pictures may differ. It did not show the loop's cost curve turning linear, and it did not need the rounds (§3). The person is a model with one number, a reaction time of 0.3 s, and one assumption, that a person in control gives exact labels: there are no human data here. It did not say what a policy should output when the right action is not one point (lesson 4), how to commit to one route for a while (lesson 5), where corrections return as a stage of a training recipe (lesson 14), or how many trials a verdict on any of this needs (lesson 15).

Common mistakes / failure modes

"the loop is just more data"
The same 7,104 frames as 39 calm demonstrations reach the mat in 59.0 % of runs; as the clone's own labelled states, 95.0 % (§2).
"noise injection is a poor man's DAgger"
At three times the gust it reaches 99.0 %, above the loop, with a demonstrator that touches a post in 36 % of its runs (§2).
"more rounds, each replacing the last"
Twenty rollouts from the first clone in one batch reach 96.2 %, against 95.6 % in four rounds, so rounds spare the robot, not the labels; keeping only the newest round reaches 3.6 % (§3).
"a person labels like the program, only slower"
Noisy labels are nearly free (94.6 %); late labels are not: 79.1 % at 0.3 s and 6.6 % at 0.4 s (§4).
"two good teachers beat one"
Ten demonstrations from each of two operators: 24.0 %, against 63.5 % from one; after four rounds of their labels, 0.0 % (§5).

Checkpoint exercise

Try it
Operators A and B command the same speed at one pose, A at +37° and B at −37° to the x axis; the first post stands 20 cm ahead, with a 5 cm halo. (a) What heading does the mean of their commands have? (b) How far above the post's centre does a heading of +37° cross the post's column, and does it clear the halo? (c) Four rounds label 3,228 frames at 0.05 s a frame. How long is that? Answer: (a) 0°: along the axis, straight at the post, which is what a regression returns. (b) 20 cm × tan 37° = 15.1 cm above the centre, 10.1 cm beyond the halo; the mean heading crosses at 0 cm, inside it. (c) 3,228 × 0.05 s = 161 s, about 2.7 minutes of a person's time at the control rate.

Where this points next

Labelling the states the clone itself visits works on the Bench: four rounds of five runs take it from 63.5 % to 95.6 %. The labels are the price, about 3,228 frames. A person supplies them late when replaying (79.1 % after four rounds at a lag of 0.3 s, 6.6 % at 0.4 s) and safely, but with few frames to show for it, when taking over (72.3 %). And a person is not a function of the state. Two operators who pass the first post on opposite sides leave a regression to answer with the average of their commands, which heads 4.8° at the post: with the same number of frames, ten demonstrations from each, the clone reaches the mat in 24.0 % of runs, and in 0.0 % after four rounds of labels. What should a policy output when the right action is not one point?

Takeaway
A clone is trained on the expert's states and scored on its own, so put the states it visits into the training set, labelled by the expert, and keep everything: DAgger. On the Bench four rounds of five runs take success from 63.5 % to 95.6 %, the steps lost from 60.5 to 9.6, and the pull-back per step from −0.1 % to 11.9 % (means over eight seeds). What counts is the number of own-state frames labelled; the rounds only spare the robot. With the level well chosen, noise injection beats the loop here, but its demonstrator touches a post in 36 % of its own runs at three times the gust. A person labels late on a replay and exactly but scarcely when taking over, and two people who pass the first post on opposite sides leave the regression to average them: 24.0 % from ten demonstrations each, then 0.0 % after the loop.

Interview prompts

Companion reads: Reinforcement Learning · 17 Imitation and IRL (DAgger and its linear bound from the MDP side), World Models · 21 Teacher forcing to rollout (training on inputs the model will not see, in a world model) and World Models · 06 Errors compound (the horizon at which a model's own predictions stop being trusted).