all_lessons/Robot Model Training/01 · Cloninglesson 1 / 24

Copy a person who knows: behaviour cloning

A trained world model says what would happen if the robot did something, and it has never chosen anything. A robot has to choose about twenty times a second, and the cheapest teacher for a chooser is a person who already is one. This lesson turns a person's demonstrations into a controller with the plainest tool there is, regression. It sets up the Bench, a two-link arm carrying a cup past a row of vases; says what a demonstration and a policy are; weighs a stored policy against planning at every step; copies twenty demonstrations with a learner that looks up what the expert did in the most similar moment; and scores the result three ways. The scores disagree. The clone copies the expert to within a few percent on frames it has never seen, makes four times that error on the frames it produces itself, and touches a post in more than a third of its runs on a five-post course. Why they disagree is the next lesson.

The thesis, here
A demonstration is a list of pairs: what the robot sensed, what the person did. Fitting that list is ordinary regression, and it is the cheapest teacher there is. But the number training minimises, the copy error on frames the expert visited, is not the number the robot lives by. The robot lives by whether the task gets done in a loop where it visits its own frames, and the same function can be accurate on the first kind of frame and not on the second.
Linear position
Forced by: A trained world model answers what would happen if the robot did something, and it has never chosen anything: it has no goal, and the data it is tested on do not change when it is used. A robot has to choose, about twenty times a second, from what it has just sensed, and the cheapest teacher for a chooser is a person who already is one. What does a robot learn by copying a person doing the task, and how would we know it had learned it?
New idea: a policy can be cloned from demonstrations by regression, and it has to be scored in two places that must never be confused: on the frames the expert produced, and on the frames the policy produces when it runs. The first is what training minimises; the second decides whether the task is done.
Forces next: A policy cloned from twenty demonstrations copies the expert to within a few percent on frames it has never seen, makes about four times that error on the frames it produces itself, and fails more than a third of the time on a course of five posts and more than half the time on six. The two kinds of frame come from different places: the first from where the expert went, the second from where the policy goes. Why is a policy that is accurate where the expert went lost where it goes, and how fast does the loss grow with the length of the task?
The plan
Four moves. (1) Put the problem on the Bench: the arm, the course, the clock and the expert. (2) Compare two ways to get a chooser, planning at every step and a stored policy, and the teachers that can fill one. (3) Derive behaviour cloning as regression and pick the simplest learner. (4) Score the clone in two places, read the disagreement, and say what it leaves open.

1 · The Bench, the clock and the expert

Picture a cup of coffee in a gripper, a row of vases on a table, and a two-link arm that has to carry the cup past the vases and set it on a mat at the far end without touching one. That is the Bench. Everything in this track happens on it, because it is small enough to simulate in your browser and exact enough that every claim is a number you can check.

The arm has two links of 0.5 m with its base at the origin, x to the right and y up. Its two joint angles q = (q1, q2), in radians, fix where the cup is, up to 1 m from the base. The robot is told what to do every 0.05 s, twenty times a second, by a pair of joint velocities u = (u1, u2) in rad/s; one tick of that clock is a step. The arm obeys, with a gust added: every step each joint's velocity is disturbed by a Gaussian of standard deviation 0.05 rad/s, so that a cup left alone with no command wanders about 10.0 mm in one second.

q ← q + 0.05 · (u + gust), gust ~ N(0, 0.05²) per joint per step

The course is n vases (posts) in a row, 0.2 m apart, each 3.5 cm in radius. The cup, 1.5 cm in radius, touches a post when their centres come within 5 cm. The cup starts 20 cm before the first post and has to get 10 cm past the last. The path that passes each post on alternating sides, 10 cm off its centre, leaves a margin of 5 cm. With five posts the policy is given T = 235 steps (11.75 s), about a quarter more than the expert needs (186 steps, 9.3 s).

Someone has to show how. In a laboratory that is a person with a joystick; here it is a controller you can read: it aims at the point 6 cm further along its path, moves there at 15 cm/s, and a damped Jacobian turns the cup's velocity into joint velocities. Under the same gusts it completes every course. What the robot senses is only its joint angles, two numbers, the proprioception of a real arm; there is no camera, and the posts never move, so the layout is the same in every demonstration. The symbols for the rest of the track:

symbolmeaningunit
ot = qtwhat the robot senses at step t (here its joint angles)rad
at = utwhat it does: the joint-velocity commandrad/s
πa policy: a function from what is sensed to what is done, a = π(o)
Ddemonstrations: the list of pairs (oi, ai) from runs of the expert
n, Tposts in the course; steps the policy is allowedposts; steps

2 · Two ways to get a chooser

A world model tells you what a command would do. It does not pick the command. There are two ways to turn it, or anything like it, into a robot that moves.

Plan at every step. Draw K candidate command sequences of length H, run each through the model, score the outcomes, execute the first command of the best, and do it all again 50 ms later.

Store the choice. Learn once a function π from what is sensed to what is done, and evaluate it in one pass at each step.

plan at every stepstored policy
needsa model of consequences, a score that says what "best" means, a searcha teacher
work per decisionK·H model steps: with K = 100 and H = 20 that is 2000one evaluation of π
per second at 20 Hz40000 model steps20 evaluations
what limits itthe model's trustworthy horizon (World Models, lesson 6) and the search budgetthe quality and coverage of what it was taught
a new goalhandled at once: change the scoreonly if the goal was an input it was taught with

The stored policy is cheaper by the factor 2000 in work per decision (if one evaluation of π costs about what one model step does), and it gives up the planner's flexibility. What it costs is a teacher, and there are three. A demonstration: someone who can do the task does it, and we record pairs of what they sensed and what they did. A score: a number for how well a whole attempt went, from which a policy can be improved by trial. The planner itself: run the planner offline and record its choices, which needs the model and the search we were trying to avoid. Only the first needs nothing the robot does not already have, except a person. We start there.

Road not taken · plan at every step
Planning is principled, it follows a new goal at once, and a good model makes it good. It costs 2000 times the work per decision, repeated 20 times a second for the life of the robot, and it can only look as far ahead as the model can be trusted, tens of steps in the Courtyard of the World Models track. It returns in lesson 12, where a score is the teacher, and lesson 13, where a slow planner and a fast policy share a robot.

3 · Cloning is regression

A demonstration is one run of the expert, the sequence of pairs (o0, a0), (o1, a1), …; D is all the pairs from m runs. Behaviour cloning asks for the function π that makes π(oi) ≈ ai over D. The usual measure of "≈" is squared error, and it is not arbitrary. Suppose the expert's command at o is some function π*(o) plus Gaussian scatter of variance σ². The likelihood of D under a candidate π is then a product of Gaussians, and its negative logarithm is

−ln p(D | π) = (1 / 2σ²) · Σi ‖ai − π(oi)‖² + constant

so maximum likelihood and least squares are the same problem: minimise L(π) = (1/N) Σi ‖π(oi) − ai‖² over the N pairs of D. (For any one input the best single output is the conditional mean of the commands the expert gave there.)

The learner can be as plain as you like. The plainest takes no training at all: look up what the expert did in the most similar moments. For a query o, give every stored frame the weight wi = exp(−½ ‖(o − oi)/h‖²), set to zero beyond three bandwidths, and answer the weighted mean of the stored commands:

π(o) = Σi wi ai / Σi wi (the choice of c that minimises Σi wi ‖ai − c‖²)

With h = 0.02 rad, about 1.5 cm at the cup, and an expert that moves 0.0118 rad per step, the kernel reaches 5.1 frames either way along a demonstration. If nothing stored lies within three bandwidths, the clone copies the single nearest stored frame. Training is storing, acting is looking up, and nothing outside what was stored can be answered except by that nearest copy. That is why it is a useful learner here: it is instant, it is deterministic, you can read it, and it generalises only to what resembles its data, which is the honest model of a policy that reads pictures. A smooth fit would behave differently in detail.

Two scores follow. The copy error ε is the root-mean-square difference between the clone's command and the expert's command at the same state, divided by the root-mean-square size of the expert's commands. It is scored on two sets of frames: ten held-out expert runs, for the frames the expert produces, and the clone's own runs, relabelled by the expert, for the frames the clone produces. The success is the share of 200 runs from jittered starts, under the gusts, that reach the mat without touching a post within T steps, with a Wilson 95 % interval. The expert, run the same way (100 runs), is the reference.

The widget

Clone the expert, then run the clone
The table seen from above: posts with their 5 cm halo, the expert's path (dashed), the demonstrations the clone was built from (grey) and 12 of its 200 runs (cyan reached the mat, red touched a post). Below: success against the number of posts (the default setting fills in, one point per course; the amber point is the current setting), and three errors side by side. Course length is the first control. The demonstrations can be recorded in calm conditions or under the same gusts the clone will meet.
copy error, expert frames
—
copy error, clone's own frames
—
clone reaches the mat
—
95 % interval
—
clone touches a post
—
expert reaches the mat
—
steps allowed
—
frames stored
—
Show the core JS
NW.predict = function (x, o) {
  var dy = this.Y.length ? this.Y[0].length : 0, out = new Array(dy), tot = 0, self = this, j;
  for (j = 0; j < dy; j++) out[j] = 0;
  this.each(x, function (id) {
    var e = self.dist2(id, x); if (e > 9) return;
    var w = Math.exp(-0.5 * e), y = self.Y[id]; tot += w; for (var m = 0; m < dy; m++) out[m] += w * y[m];
  });
  if (tot < 1e-12) {
    var nb = this.nearest(x); return nb < 0 ? out : this.Y[nb].slice();
  }
  for (j = 0; j < dy; j++) out[j] /= tot;
  return out;
};
...
function frameError(w, nw, runs, own) {
  var se = 0, ss = 0;
  runs.forEach(function (ro) {
    for (var t = 0; t < ro.A.length; t++) {
      var a = BN.slalom.expertAct(w, ro.S[t]), p = own ? ro.A[t] : nw.predict(ro.S[t]);
      se += Math.pow(p[0] - a[0], 2) + Math.pow(p[1] - a[1], 2); ss += a[0] * a[0] + a[1] * a[1];
    }
  });
  return Math.sqrt(se / ss);
}

What to try. Leave the defaults: five posts, 20 demonstrations recorded in calm conditions, gust 0.05. The clone's copy error on fresh expert frames is 3.4 %, and it reaches the mat in 63.5 % of its runs (interval 56.6 to 69.9 %); the expert, under the same gusts, reaches it in 100 %. Slide the course down to one post: the copy error is 4.4 % and success 87.5 %. Slide it up to six: the copy error is 2.9 % and success 42.0 %. The copy error never leaves 2.9 to 4.4 %, while success drops 45.5 points. Turn the gust to 0: the same clone, the same demonstrations, reaches the mat in 96 % of its runs, so it copies well enough when nothing disturbs it; turn it to 0.08 and success is 31 %. Return to 0.05 and double the demonstrations to 40: the copy error is 3.3 % and success 55.0 %, no better than with 20; with 10 demonstrations they are 3.3 % and 63.0 %. Now record the demonstrations under the same gusts: success rises to 81.5 % and the copy error to 4.9 %, and the gap to the expert is still there. Finally watch the second error readout: on the frames the clone itself produces its error is 13.9 %, 4.1 times the error on the expert's frames, and even with no gust it is 8.6 % against 3.4 %.

4 · Two places to be scored, and they disagree

The same experiment on every course length, with 20 calm demonstrations and gust 0.05:

postssteps allowedcopy error, expert framesreaches the mat (95 % interval)touches a post
1654.4 %87.5 % (82.2–91.4)12.5 %
21073.3 %83.0 % (77.2–87.6)17.0 %
31503.3 %75.5 % (69.1–80.9)24.5 %
41933.2 %63.5 % (56.6–69.9)36.5 %
52353.4 %63.5 % (56.6–69.9)36.5 %
62782.9 %42.0 % (35.4–48.9)58.0 %

Read the table in three passes. The copy error is flat. It sits between 2.9 and 4.4 % for every course, so the clone copies about as well on a six-post course as on a one-post course. Success is not flat. It falls from 87.5 % to 42.0 %, and the intervals for one post and six do not touch. The failures are collisions, not stalls. The share of runs that neither finish nor collide is 0 % in every row: the clone does not freeze, it walks into a post. On the five-post course 73 of the 200 runs end in a collision, and 26 of them are at the first post and 26 at the fourth, so 71.2 % of the collisions are at two of the five posts.

So the number training reports does not predict the number the robot lives by. They are measured on different frames. The copy error asks: given a frame the expert produced, how close is the clone's command to the expert's? Success asks what happens when the clone's own commands, and the gusts, decide which frames come next. Score the clone on its frames, relabelled by the expert, and the error is 13.9 % against 3.4 % on the expert's: about 4.1 times as large, in the same function, on the same course. The clone is accurate where the expert went and less accurate where it goes, and where it goes is not where it was trained.

Road not taken · more of the same demonstrations
The first reflex when a learner underperforms is more data. It is the wrong reflex here, and the widget says so: 40 demonstrations leave the copy error at 3.3 % and the success at 55.0 %, against 3.4 % and 63.5 % with 20. Calm demonstrations repeat themselves: they cover the same thin band of frames again, so the clone's error on its own frames is untouched (14.4 % against 13.9 % with 20). Recording them under the gusts moves success to 81.5 %, because it widens the band, which is a different kind of data, and lesson 2 says what kind and how much.
What this lesson did not do
It did not say why success falls as the course grows, or what the error on the clone's own frames does as the run goes on; that is lesson 2. It used a nearest-demo learner and a two-number observation; a neural network reading pictures behaves differently in detail, and lesson 2 repeats the measurement with a smooth fit and says why the picture here is the optimistic one. It did not ask where the demonstrations come from or who the expert is (lessons 3 and 4), what happens when the posts move (lesson 7) or when the task depends on a force no camera sees (lesson 10). The expert here is a controller; a person will turn out to be a different kind of teacher.

Common mistakes / failure modes

"a low validation error means a good policy"
The clone's copy error on fresh expert frames is 3.4 %, and 36.5 % of its runs on five posts end in a collision (§4).
"if the clone is wrong, give it more demonstrations"
20 to 40 demonstrations: copy error 3.4 % to 3.3 %, success 63.5 % to 55.0 %. More of the same band is more of the same (§3, widget).
"it works for the first second, so it works for the task"
One post: 87.5 % success. Six posts: 42.0 %, with a copy error no larger: 4.4 % at one post, 2.9 % at six (§4).
"the clone fails because the network is bad"
There is nothing to train: it is a lookup of what the expert did. With no gust it reaches the mat 96 % of the time on five posts; the failure comes from the disturbance and the loop (§3, §4).
"planning is principled, a policy is a shortcut"
Planning costs 2000 times the work per decision and inherits the model's horizon; the policy trades flexibility for a teacher (§2).
"the demonstrations are the frames the robot will see"
The clone's own frames carry 4.1 times the error of the expert's frames, even in this two-number world (§4).

Checkpoint exercise

Try it
A planner evaluates 200 candidate sequences of 25 steps each with a learned model that costs 0.1 ms per step, one after another on one core. (a) How long does one decision take? (b) The robot decides 20 times a second: by what factor is that too slow? (c) A stored policy costs 0.1 ms per evaluation: how many decisions a second could one core make with it? Answer: (a) 200 × 25 × 0.1 ms = 500 ms. (b) A decision is due every 50 ms, so it is 10 times too slow. (c) 1 s / 0.1 ms = 10000 decisions a second, 500 times the 20 the robot needs. Batching the 200 candidates hides the latency and not the work: the planner still does 5000 model steps per decision against one evaluation.

Where this points next

A person's demonstrations can be turned into a controller by regression, and on the Bench a clone from twenty of them copies the expert to 3.4 % on frames it has not seen. It touches a post in 36.5 % of its runs on five posts and in 58.0 % on six, and on the frames it produces for itself its error is 4.1 times larger. The copy error is measured where the expert went; the clone's success is decided where the clone goes, and nothing in training looks there. Why is a policy that is accurate where the expert went lost where it goes, and how fast does the loss grow with the length of the task?

Takeaway
A policy is a function from what the robot senses to what it does, and the cheapest way to get one is regression on a person's demonstrations: it needs no model, no score and no planner, and it costs one evaluation per decision against 2000 times that for planning. On the Bench, a nearest-demo clone of 20 demonstrations has a copy error of 3.4 % on fresh expert frames, yet reaches the mat in only 63.5 % of its runs on five posts (42.0 % on six, 87.5 % on one). The copy error is flat in the length of the course while success is not, because the two are measured on different frames: on the clone's own frames its error is 13.9 %. More demonstrations of the same kind do nothing (55.0 % with 40), and with no disturbance the clone does fine (96 %). Score a policy by what happens when it runs, with an interval, and ask where its frames come from.

Interview prompts

Companion reads: World Models · 09 Plan with it (the planner this lesson declines to run every step), World Models · 06 Errors compound (why a planner can only look a few steps ahead), Reinforcement Learning · 17 Imitation and IRL (the same problem from the MDP side) and World Models · 31 Capstone (the track this one follows).