Labels on your own states: DAgger and corrections
Lesson 2 left a clone that copies the expert to 3.4 % on frames it has not seen, is lost on the frames it makes itself, and loses steps as the square of the horizon however many demonstrations it has. This lesson puts the clone's own frames into the data: it runs, the expert says what it should have done at every state it visited, those frames join everything stored, and the clone is rebuilt. That is DAgger. Four rounds of five runs take it from 63.5 % to 95.6 % on five posts (mean of eight seeds). The labels are the price: a person gives them late when replaying a run, scarcely when taking over, and two people asked from the same pose can send the cup either side of a post.
New idea: the training set should be the states the policy itself visits, labelled by the expert, with nothing already stored thrown away. It buys the pull-back that lesson 2 found missing; it costs labelled frames and a labeller who can answer for states it did not drive, which a program can and a person only in part.
Forces next: Training on the policy's own states, with the expert's actions as labels, takes a policy that finishes five posts about six times in ten to one that finishes them about nineteen times in twenty after four rounds of five rollouts, and a person supplies those labels best by taking over when the policy goes wrong. But people are not functions: asked from the same pose, or asked by two operators, one goes left of the post and another right, and a regression asked to fit both returns their average, a route that runs into the post. What should a policy output when the right action is not one point?
1 · What the training set should be
Lesson 2 found why the two scores of lesson 1 disagree: the clone's error is a curve in the distance from its data, and its own frames sit far out along it. A learner can only fit the frames it is given, so what matters is which frames. The theory agrees. Ross and Bagnell (2010) bound the extra cost of a clone trained on the expert's frames by T²ε, with T the horizon in steps and ε the clone's error on the expert's states; it is an upper bound, tight only for constructed examples, and on the Bench ε is the 3.4 % of lesson 1. Ross, Gordon and Bagnell (2011) move the measurement: with εN the loss, after N rounds, on the states the learner's own policies visit, the extra cost is bounded by about TεN, linear in T. The two ε are different numbers: for the first clone, 3.4 % on the expert's states and 13.9 % on its own. The move changes not how well the function fits but where it is asked to fit.
Four constraints follow. (a) Only the expert knows what to do in a state the clone has wandered into, so the labels come from the expert. (b) The clone's cost is paid on its own states, so the states should come from the clone. (c) A label is a query: count labelled frames and, if a person answers, person-seconds at 0.05 s a frame. (d) The clone's states depend on the clone, which depends on the data, so they cannot be listed in advance; the clone has to be run to find them. That makes a loop. Round 0 stores the demonstrations, the expert driving. Each later round runs the current clone for five rollouts under the gust, the expert labels every state they visited, the labelled frames are added to everything stored, and the clone is rebuilt (for a table of stored frames, rebuilt means stored). This is DAgger (Ross, Gordon and Bagnell, 2011) in the variant they report often does best in practice: the expert drives the first round only.
2 · Three ways to buy displaced frames
Lesson 2 found what the clone lacks: displaced frames, each labelled with the expert's response. A fair comparison fixes what is spent, the number of labelled frames. Four rounds of five rollouts leave 7,104 frames stored in the default run, 3,684 of demonstrations and 3,420 labelled in the loop. The alternatives add whole demonstrations until at least that many are stored.
| where the frames come from | what the expert does | clone reaches the mat | demonstrations in which the expert touches a post |
|---|---|---|---|
| 39 calm demonstrations | drives, as in lesson 1 | 59.0 % | 0 % |
| demonstrations under a gust of 0.05 rad/s, the gust the clone meets | drives while the plant is shaken; each label is its own command | 78.5 % | 0 % |
| under a gust of 0.10, twice that | the same, shaken harder | 95.5 % | 4 % |
| under a gust of 0.15, three times | the same, harder again | 99.0 % | 36 % |
| the clone's own states, four rounds of five rollouts | labels what the clone did, never drives | 95.0 % | none |
More of the same does nothing, as lesson 2 found. Injecting noise is a method in its own right (Laskey et al., 2017, DART): noise goes into the supervisor's control while it demonstrates, which forces it to show how to recover from errors. On the Bench the equivalent is a gust on the plant while the expert drives, so the demonstrations contain displaced states with the recovery labelled. What sets its strength is the level, not the frames: 78.5 % at the gust the clone meets (81.5 % with the 20 demonstrations of lesson 2), 95.5 % at twice it and 99.0 % at three times, which is above the loop at every number of frames the widget below offers. Even 3,851 frames, fewer than the loop holds after one round, reach 98.5 % at twice the gust.
3 · The loop, measured
Each round runs the clone five times and labels every state those runs visited, about 807 frames (a run ends early when it touches a post). The widget opens on the typical run: of eight collection seeds (eight draws of the collection runs' gusts and starts), the one whose success curve lies closest to their average. A step is lost if the cup has touched a post, is farther than 4 cm from the path or never reached the mat, and steps lost is the expected number of the first T = 235 steps that are lost (lesson 2's C(T)). The ruler is lesson 2's: how far each of the clone's own frames is from the nearest stored frame, in kernel bandwidths.
| round | labelled frames | reaches the mat | steps lost | error on own frames | own frames at least 1 bandwidth from data | pull-back per step |
|---|---|---|---|---|---|---|
| 0 | 0 | 63.5 % | 60.5 | 13.9 % | 23.4 % | −0.1 % |
| 1 | 834 | 78.5 % | 40.8 | 11.2 % | 1.6 % | 7.8 % |
| 2 | 1,713 | 87.0 % | 24.6 | 9.6 % | 0.5 % | 7.6 % |
| 4 | 3,420 | 95.0 % | 10.9 | 8.3 % | 0.06 % | 10.4 % |
Every measure of lesson 2 moves as the argument says it should. The clone's own frames end up on its data. The pull-back lesson 2 found missing is there after one round, 7.8 % of a displacement recovered per step in this run (4.2 % over the eight seeds) against the expert's 7.5 %, because the data now hold displaced states labelled with the expert's response. The error on the clone's own frames, the number training now minimises, falls from 13.9 to 8.3 %, though not to the expert-frame 3.4 %. Over eight collection seeds the success at round 4 is 95.6 % (the seeds range from 93.5 to 97.5 %), about nineteen runs in twenty, and it creeps up to 98.4 % by round 8.
Keep everything. A round could replace the stored frames instead of adding to them. After four rounds, averaged over eight seeds, keeping everything reaches 95.6 %, keeping the demonstrations and only the newest round 76.7 %, and keeping only the newest round 3.6 %. The ruler says why: the share of the clone's own frames a bandwidth or more from its data is 0.1 % keeping everything, 6.4 % once the earlier rounds are dropped, and 54 % once the expert's own path is dropped too, when the data hold only the states the last clone visited.
Do the rounds matter? Take all the rollouts from the first clone instead and rebuild once. Twenty rollouts reach 96.2 % against 95.6 % for four rounds of five, forty rollouts 98.4 % against 98.4 % for eight rounds of five. The first clone is the worst driver on the Bench, so its runs wander widest, which would explain why its batch already covers where the later, better clones go. Labelled frames count; rounds do not. What the rounds buy is a robot that crashes less while the data are collected: 8.4 of the twenty runs of the first clone end in a post against 5.2 for retrained clones, 16.9 of forty against 5.6. The theory asks for rounds because each new clone can go where the last one did not; the check is the ruler, stopping when almost none of the clone's own frames is a bandwidth from its data (0.1 % at round 4) and success has stopped rising.
Does it bend the square? Not on this Bench. From one to six posts the loop lowers the steps lost by a factor of 4.0 at one post and 4.3 at six, and the exponent reads 1.94 ± 0.27 (one standard error) against lesson 2's 2.00, which these data cannot tell apart: a touch is permanent, so a small hazard per step is still summed over every step after it. The linear bound concerns a loss measured on the learner's own states; the Bench does not show the curve turning linear, and this lesson does not claim it.
4 · What a label costs when a person gives it
Four rounds label 3,228 frames on average. A person who answers one control step at a time spends 0.05 s on each: 161 s. That is the bill if the person is as good as the expert, and the expert is not a person: it is a program, a function of the joint angles, so it answers for any state at once and exactly. A person falls short of the program in three ways: imprecise, late, and not single-valued; the first two are here, the third in §5. To price them the Bench needs a model, and the simplest has one number, a reaction time τ of 0.3 s, 6 frames. It is an assumption, not a measurement of people, and the table varies it.
Imprecise. Add noise of 40 % of the command's own size to every label, 12 times the clone's copy error, and the clone still reaches the mat in 94.6 % of runs (95.6 with exact labels). Noise 1.6 times the command's own size costs 8.2 points. A regression answers with the mean of the labels near a query (lesson 1), and noise that cancels in a mean costs almost nothing.
Late. A person replaying a run steers along it and sees no effect of what they do, so their command at frame t answers the state they saw τ frames earlier. Kelly et al. (2019) name the mechanism: asked for labels without being in control, a person's labels are likely to degrade through perceived actuator lag. On the expert's own runs a 0.3 s lag makes the label wrong by 42.5 % of the command, about as much as the 40 % noise. But the clone answers with the average of nearby labels, and on those frames that average keeps 97 % of the lag's error against 10 % of the noise's: neighbouring frames carry nearly the same lag error, so averaging removes almost none of it.
Taking over. The remedy is to put the person in control. In HG-DAgger (Kelly et al., 2019) the clone drives, the person takes control when they judge it necessary and keeps it until they hand it back, and only the frames the person drove are labelled. On the Bench the person takes over when the cup is more than 2.5 cm from the path, after τ frames, drives until it is within 1 cm, and hands back. The 2.5 cm is not free: the clone's failures cross it a median of 12 frames before the touch, 90 % of them more than the 6 frames a person needs to react, but cross 4 cm only a median of 3 frames before, 15 % of them more than 6. In control the person sees what each command does, so the labels are exact (our assumption again).
| after four rounds, unless stated | reaches the mat, mean (lowest to highest of 8 seeds) | labelled frames | collection runs ending in a post, of 20 |
|---|---|---|---|
| the program labels every frame | 95.6 % (93.5 to 97.5) | 3,228 | 5.2 |
| labels with 40 % noise | 94.6 % (91.5 to 97.5) | 3,286 | 4.9 |
| a person replaying, 0.1 s late | 96.3 % (92.0 to 100) | 3,428 | 3.8 |
| a person replaying, 0.3 s late | 79.1 % (58.5 to 90.0) | 3,166 | 9.0 |
| a person replaying, 0.4 s late | 6.6 % (2.0 to 14.0) | 2,589 | 13.8 |
| a person taking over, 0.3 s | 72.3 % (66.5 to 76.5) | 1,209 | 0.8 |
| taking over, 0.3 s, ten rounds | 79.2 % (72.5 to 84.5) | 3,072 | 1.2 of 50 |
A person supplies labels best by taking over in one sense: it is the only way whose labels do not depend on how late the person is (70.9, 72.3 and 67.7 % at 0, 0.3 and 0.5 s, against 96.3, 79.1 and 6.6 % for a replay at 0.1, 0.3 and 0.4 s), and it keeps the robot out of the posts while the data are collected (0.8 of 20 collection runs end in one, against 3.8 to 13.8 for a replay). It is steadiest, not most accurate: a replay with a short lag does better. Kelly et al. (2019) report faster, more stable learning than DAgger in a driving simulation. On the Bench it is safer but not faster: it labels only the frames the person drove, so the person's tolerance bounds the displacements the data ever show, and ten rounds reach 79.2 % where the program's 3,228 frames reached 95.6 %.
The widget
What to try. Leave the defaults: round 0, the program labels, the latest clone drives. That is lesson 2's clone: 63.5 % reach the mat (interval 56.6 to 69.9 %), 60.5 of 235 steps are lost, the error on its own frames is 13.9 %, and 23.4 % of them are a bandwidth or more from the data. Slide to round 4: 95.0 %, 10.9 steps lost, error 8.3 %, 0.06 % far from the data, pull-back 10.4 % a step against the expert's 7.5 %, from 3,420 labelled frames, with 4 of the 20 collection runs ending in a post. Under compare with, choose demonstrations under gust 0.10: 95.5 % from the same frames. Under runs are made by, choose the first clone: 95.5 % at round 4, with 7 of 20 runs ending in a post. Under who labels, choose a person replaying the run: 84.5 % at round 4 with the reaction time left at 0.3 s, 2.0 % at 0.4 s; choose a person taking over: 76.0 % from 1,223 labelled frames. Last, choose two people: 24.0 % at round 0, the clone's command at the start frame heading 4.8° against 37.1° for A; 0.0 % at round 4, with 99 % of the runs ending at the first post.
5 · Two people
A person is also not a function of the state. Asked twice from the same pose, or by two operators, one goes above the first post and the other below, and both are right: the course has room for either. On the Bench operator A passes the first post above it (the expert of every lesson so far; above means larger y) and operator B below it; the posts and the 5 cm margin are the same. Give the clone ten demonstrations of each, from the same ten starting poses, as many frames as before (3,682).
At the start frame the cup velocity A commands points 37.1° above the x axis and B's points −37.1°, as far below it. The clone's answer is the kernel-weighted mean of the commands stored near the start frame, and the mean of two mirror-image commands points almost along the axis, 4.8°. Followed in a straight line to the column of the first post, 20 cm ahead, A's heading arrives 15.1 cm above the post's centre, outside its 5 cm halo; the clone's arrives 1.7 cm from it, inside. A regression minimises the squared difference to both, and for squared error the best single answer is the mean (lesson 1): a route that runs into the post.
| the clone's data (mean of eight seeds) | reaches the mat | runs ending at the first post |
|---|---|---|
| one operator, 20 demonstrations | 63.5 % | 13.0 % |
| two operators, 10 + 10 demonstrations | 24.0 % | 18.0 % |
| after one round, each run labelled by A or B | 6.8 % | 62.8 % |
| after two rounds | 0.3 % | 91.1 % |
| after four rounds | 0.0 % | 92.6 % |
At round 0 the first post ends 18.0 % of the runs and the later posts another 58.0 %: the two routes are mirror images about the line of posts, so their average is the straight line through every post. The loop does not repair it; it makes it worse. Each collection run is labelled by one of the two operators, correctly for that operator. The clone, now heading for the post, visits states head-on to it, and each of those gets a label from above or from below. The mean of the labels is still the straight line into the post, and every round adds more frames that say so. DAgger assumed that the labeller is a function of the state, and the commands of two people are not.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
Labelling the states the clone itself visits works on the Bench: four rounds of five runs take it from 63.5 % to 95.6 %. The labels are the price, about 3,228 frames. A person supplies them late when replaying (79.1 % after four rounds at a lag of 0.3 s, 6.6 % at 0.4 s) and safely, but with few frames to show for it, when taking over (72.3 %). And a person is not a function of the state. Two operators who pass the first post on opposite sides leave a regression to answer with the average of their commands, which heads 4.8° at the post: with the same number of frames, ten demonstrations from each, the clone reaches the mat in 24.0 % of runs, and in 0.0 % after four rounds of labels. What should a policy output when the right action is not one point?
Interview prompts
- Why is DAgger's bound linear in the horizon where behaviour cloning's is quadratic? (§1 — cloning bounds the cost by T²ε with ε measured on the expert's states, DAgger by TεN with εN measured on the learner's own: the quantity changes, not only the exponent.)
- Four rounds of five runs or one batch of twenty? (§3 — on the Bench the batch does as well; rounds spare the robot, and a learner that moves needs them.)
- What is the case for and against noise injection instead of the loop? (§2 — at three times the gust it beats the loop and needs no rounds; but the level has to be guessed and the demonstrator pushed around.)
- Why does label noise barely hurt a kernel regression while label lag ruins it? (§4 — noise cancels in the mean of nearby labels; a lag shifts every label the same way.)
- Why have a person take over instead of labelling a replay, and what does it cost? (§4 — in control the labels are exact whatever the lag, but only the frames the person drove are labelled.)
- Why does DAgger with two operators get worse each round? (§5 — each round labels states head-on to the post with commands from above and below; their mean is the straight line into it.)
Companion reads: Reinforcement Learning · 17 Imitation and IRL (DAgger and its linear bound from the MDP side), World Models · 21 Teacher forcing to rollout (training on inputs the model will not see, in a world model) and World Models · 06 Errors compound (the horizon at which a model's own predictions stop being trusted).