Copy a person who knows: behaviour cloning
A trained world model says what would happen if the robot did something, and it has never chosen anything. A robot has to choose about twenty times a second, and the cheapest teacher for a chooser is a person who already is one. This lesson turns a person's demonstrations into a controller with the plainest tool there is, regression. It sets up the Bench, a two-link arm carrying a cup past a row of vases; says what a demonstration and a policy are; weighs a stored policy against planning at every step; copies twenty demonstrations with a learner that looks up what the expert did in the most similar moment; and scores the result three ways. The scores disagree. The clone copies the expert to within a few percent on frames it has never seen, makes four times that error on the frames it produces itself, and touches a post in more than a third of its runs on a five-post course. Why they disagree is the next lesson.
New idea: a policy can be cloned from demonstrations by regression, and it has to be scored in two places that must never be confused: on the frames the expert produced, and on the frames the policy produces when it runs. The first is what training minimises; the second decides whether the task is done.
Forces next: A policy cloned from twenty demonstrations copies the expert to within a few percent on frames it has never seen, makes about four times that error on the frames it produces itself, and fails more than a third of the time on a course of five posts and more than half the time on six. The two kinds of frame come from different places: the first from where the expert went, the second from where the policy goes. Why is a policy that is accurate where the expert went lost where it goes, and how fast does the loss grow with the length of the task?
1 · The Bench, the clock and the expert
Picture a cup of coffee in a gripper, a row of vases on a table, and a two-link arm that has to carry the cup past the vases and set it on a mat at the far end without touching one. That is the Bench. Everything in this track happens on it, because it is small enough to simulate in your browser and exact enough that every claim is a number you can check.
The arm has two links of 0.5 m with its base at the origin, x to the right and y up. Its two joint angles q = (q1, q2), in radians, fix where the cup is, up to 1 m from the base. The robot is told what to do every 0.05 s, twenty times a second, by a pair of joint velocities u = (u1, u2) in rad/s; one tick of that clock is a step. The arm obeys, with a gust added: every step each joint's velocity is disturbed by a Gaussian of standard deviation 0.05 rad/s, so that a cup left alone with no command wanders about 10.0 mm in one second.
q ← q + 0.05 · (u + gust), gust ~ N(0, 0.05²) per joint per step
The course is n vases (posts) in a row, 0.2 m apart, each 3.5 cm in radius. The cup, 1.5 cm in radius, touches a post when their centres come within 5 cm. The cup starts 20 cm before the first post and has to get 10 cm past the last. The path that passes each post on alternating sides, 10 cm off its centre, leaves a margin of 5 cm. With five posts the policy is given T = 235 steps (11.75 s), about a quarter more than the expert needs (186 steps, 9.3 s).
Someone has to show how. In a laboratory that is a person with a joystick; here it is a controller you can read: it aims at the point 6 cm further along its path, moves there at 15 cm/s, and a damped Jacobian turns the cup's velocity into joint velocities. Under the same gusts it completes every course. What the robot senses is only its joint angles, two numbers, the proprioception of a real arm; there is no camera, and the posts never move, so the layout is the same in every demonstration. The symbols for the rest of the track:
| symbol | meaning | unit |
|---|---|---|
| ot = qt | what the robot senses at step t (here its joint angles) | rad |
| at = ut | what it does: the joint-velocity command | rad/s |
| π | a policy: a function from what is sensed to what is done, a = π(o) | |
| D | demonstrations: the list of pairs (oi, ai) from runs of the expert | |
| n, T | posts in the course; steps the policy is allowed | posts; steps |
2 · Two ways to get a chooser
A world model tells you what a command would do. It does not pick the command. There are two ways to turn it, or anything like it, into a robot that moves.
Plan at every step. Draw K candidate command sequences of length H, run each through the model, score the outcomes, execute the first command of the best, and do it all again 50 ms later.
Store the choice. Learn once a function π from what is sensed to what is done, and evaluate it in one pass at each step.
| plan at every step | stored policy | |
|---|---|---|
| needs | a model of consequences, a score that says what "best" means, a search | a teacher |
| work per decision | K·H model steps: with K = 100 and H = 20 that is 2000 | one evaluation of π |
| per second at 20 Hz | 40000 model steps | 20 evaluations |
| what limits it | the model's trustworthy horizon (World Models, lesson 6) and the search budget | the quality and coverage of what it was taught |
| a new goal | handled at once: change the score | only if the goal was an input it was taught with |
The stored policy is cheaper by the factor 2000 in work per decision (if one evaluation of π costs about what one model step does), and it gives up the planner's flexibility. What it costs is a teacher, and there are three. A demonstration: someone who can do the task does it, and we record pairs of what they sensed and what they did. A score: a number for how well a whole attempt went, from which a policy can be improved by trial. The planner itself: run the planner offline and record its choices, which needs the model and the search we were trying to avoid. Only the first needs nothing the robot does not already have, except a person. We start there.
3 · Cloning is regression
A demonstration is one run of the expert, the sequence of pairs (o0, a0), (o1, a1), …; D is all the pairs from m runs. Behaviour cloning asks for the function π that makes π(oi) ≈ ai over D. The usual measure of "≈" is squared error, and it is not arbitrary. Suppose the expert's command at o is some function π*(o) plus Gaussian scatter of variance σ². The likelihood of D under a candidate π is then a product of Gaussians, and its negative logarithm is
−ln p(D | π) = (1 / 2σ²) · Σi ‖ai − π(oi)‖² + constant
so maximum likelihood and least squares are the same problem: minimise L(π) = (1/N) Σi ‖π(oi) − ai‖² over the N pairs of D. (For any one input the best single output is the conditional mean of the commands the expert gave there.)
The learner can be as plain as you like. The plainest takes no training at all: look up what the expert did in the most similar moments. For a query o, give every stored frame the weight wi = exp(−½ ‖(o − oi)/h‖²), set to zero beyond three bandwidths, and answer the weighted mean of the stored commands:
π(o) = Σi wi ai / Σi wi (the choice of c that minimises Σi wi ‖ai − c‖²)
With h = 0.02 rad, about 1.5 cm at the cup, and an expert that moves 0.0118 rad per step, the kernel reaches 5.1 frames either way along a demonstration. If nothing stored lies within three bandwidths, the clone copies the single nearest stored frame. Training is storing, acting is looking up, and nothing outside what was stored can be answered except by that nearest copy. That is why it is a useful learner here: it is instant, it is deterministic, you can read it, and it generalises only to what resembles its data, which is the honest model of a policy that reads pictures. A smooth fit would behave differently in detail.
Two scores follow. The copy error ε is the root-mean-square difference between the clone's command and the expert's command at the same state, divided by the root-mean-square size of the expert's commands. It is scored on two sets of frames: ten held-out expert runs, for the frames the expert produces, and the clone's own runs, relabelled by the expert, for the frames the clone produces. The success is the share of 200 runs from jittered starts, under the gusts, that reach the mat without touching a post within T steps, with a Wilson 95 % interval. The expert, run the same way (100 runs), is the reference.
The widget
What to try. Leave the defaults: five posts, 20 demonstrations recorded in calm conditions, gust 0.05. The clone's copy error on fresh expert frames is 3.4 %, and it reaches the mat in 63.5 % of its runs (interval 56.6 to 69.9 %); the expert, under the same gusts, reaches it in 100 %. Slide the course down to one post: the copy error is 4.4 % and success 87.5 %. Slide it up to six: the copy error is 2.9 % and success 42.0 %. The copy error never leaves 2.9 to 4.4 %, while success drops 45.5 points. Turn the gust to 0: the same clone, the same demonstrations, reaches the mat in 96 % of its runs, so it copies well enough when nothing disturbs it; turn it to 0.08 and success is 31 %. Return to 0.05 and double the demonstrations to 40: the copy error is 3.3 % and success 55.0 %, no better than with 20; with 10 demonstrations they are 3.3 % and 63.0 %. Now record the demonstrations under the same gusts: success rises to 81.5 % and the copy error to 4.9 %, and the gap to the expert is still there. Finally watch the second error readout: on the frames the clone itself produces its error is 13.9 %, 4.1 times the error on the expert's frames, and even with no gust it is 8.6 % against 3.4 %.
4 · Two places to be scored, and they disagree
The same experiment on every course length, with 20 calm demonstrations and gust 0.05:
| posts | steps allowed | copy error, expert frames | reaches the mat (95 % interval) | touches a post |
|---|---|---|---|---|
| 1 | 65 | 4.4 % | 87.5 % (82.2–91.4) | 12.5 % |
| 2 | 107 | 3.3 % | 83.0 % (77.2–87.6) | 17.0 % |
| 3 | 150 | 3.3 % | 75.5 % (69.1–80.9) | 24.5 % |
| 4 | 193 | 3.2 % | 63.5 % (56.6–69.9) | 36.5 % |
| 5 | 235 | 3.4 % | 63.5 % (56.6–69.9) | 36.5 % |
| 6 | 278 | 2.9 % | 42.0 % (35.4–48.9) | 58.0 % |
Read the table in three passes. The copy error is flat. It sits between 2.9 and 4.4 % for every course, so the clone copies about as well on a six-post course as on a one-post course. Success is not flat. It falls from 87.5 % to 42.0 %, and the intervals for one post and six do not touch. The failures are collisions, not stalls. The share of runs that neither finish nor collide is 0 % in every row: the clone does not freeze, it walks into a post. On the five-post course 73 of the 200 runs end in a collision, and 26 of them are at the first post and 26 at the fourth, so 71.2 % of the collisions are at two of the five posts.
So the number training reports does not predict the number the robot lives by. They are measured on different frames. The copy error asks: given a frame the expert produced, how close is the clone's command to the expert's? Success asks what happens when the clone's own commands, and the gusts, decide which frames come next. Score the clone on its frames, relabelled by the expert, and the error is 13.9 % against 3.4 % on the expert's: about 4.1 times as large, in the same function, on the same course. The clone is accurate where the expert went and less accurate where it goes, and where it goes is not where it was trained.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
A person's demonstrations can be turned into a controller by regression, and on the Bench a clone from twenty of them copies the expert to 3.4 % on frames it has not seen. It touches a post in 36.5 % of its runs on five posts and in 58.0 % on six, and on the frames it produces for itself its error is 4.1 times larger. The copy error is measured where the expert went; the clone's success is decided where the clone goes, and nothing in training looks there. Why is a policy that is accurate where the expert went lost where it goes, and how fast does the loss grow with the length of the task?
Interview prompts
- Why is behaviour cloning just regression, and why squared error? (§3 — if the expert's command is a function plus Gaussian scatter, maximising the likelihood of the demonstrations is minimising the sum of squared differences.)
- Give two reasons to store a policy instead of planning at every step, and one reason not to. (§2 — one evaluation against K·H model steps per decision; no dependence on the model's horizon; but it needs a teacher and does not follow a new goal it was not taught.)
- A cloned policy has 3 % validation error and fails half its runs. Is that a contradiction? (§4 — no: validation error is measured on the expert's frames and success is decided on the policy's own; here the policy's error on its own frames is about four times larger.)
- What does a nearest-demo regressor do for an input far from every stored frame, and why is that the honest model of a policy that reads pictures? (§3 — it copies the single nearest frame; it generalises only to what resembles its data.)
- Doubling the demonstrations neither lowered the copy error nor raised the success. Why? (§3, widget — calm demonstrations cover the same thin band of frames; the error is in the frames outside it.)
- How would you put an honest error bar on a success rate, and how wide is it for 200 runs? (§3, §4 — a Wilson interval; about plus or minus 6.6 points at 63.5 % for 200 runs.)
Companion reads: World Models · 09 Plan with it (the planner this lesson declines to run every step), World Models · 06 Errors compound (why a planner can only look a few steps ahead), Reinforcement Learning · 17 Imitation and IRL (the same problem from the MDP side) and World Models · 31 Capstone (the track this one follows).