all_lessons/Robot Model Training/04 · Distributionslesson 4 / 24

People are not functions

Lesson 3 left a clone that reaches the mat on about nineteen runs in twenty when one operator's frames build it (96.5 % on this lesson's draw). Add a second operator, who passes the first post on the other side, pool both sets of frames and fit the same regression, and it reaches the mat on 53.5 %: at the start pose one operator's action takes the cup up and the other's down, and squared error returns their average, an action nobody took, aimed at the post. This lesson changes what a policy outputs: not one action but a distribution over actions, with a hump for each route the data show, drawn from at every step. A draw lands on a route and never between two. But it is made again at every step, so near a fork the arm is sent up, then down, and 26.5 % of runs still end in a post.

The thesis, here
A policy trained by squared error answers every observation with one action, and where the data hold two good actions the answer that costs least is their average, which neither operator took. The repair is in what the policy outputs. Let it output a distribution with as many humps as the data show at that observation, and act by drawing from it: a draw lands on one of the routes, and the average of many draws is still the old answer, so a task with one route is untouched. What a draw does not have is memory: the next step draws again.
Linear position
Forced by: Training on the policy's own states, with the expert's actions as labels, takes a policy that finishes five posts about six times in ten to one that finishes them about nineteen times in twenty after four rounds of five rollouts, and a person supplies those labels best by taking over when the policy goes wrong. But people are not functions: asked from the same pose, or asked by two operators, one goes left of the post and another right, and a regression asked to fit both returns their average, a route that runs into the post. What should a policy output when the right action is not one point?
New idea: output a distribution over actions, with a hump for each route the data show, and act by drawing from it. The two routes stay two routes, and every step now decides again.
Forces next: A policy that outputs a distribution over actions, and acts by sampling it, recovers both routes instead of their collision. But it draws a new sample at every step, and near the fork each draw can pick a different route: the arm dithers between them and a quarter of the rollouts still end in a post. Committing to one route for a while would stop the dithering, at the price of not looking while committed. How long should a policy commit, and what should it commit to?
The plan
Five moves. (1) Pool two operators and watch a regression return their average. (2) Ask what squared error asks for, and what other single answers return. (3) Output the distribution the data show and act by drawing from it. (4) Do it without storing frames: a mixture and a codebook. (5) Count what drawing at every step costs.

1 · A second operator

The expert of the last three lessons passes the first post above it (larger y), the second below it, and so on down the row. Call it operator A. A second operator, B, does the opposite: below the first post, above the second. Both clear every post by the same 5 cm, and their routes are mirror images that cross midway between each pair of posts. There the same pose is asked for opposite actions, A's taking the cup down and B's up. So is the start pose, where both begin. Five posts, five such forks: the start pose and the four crossings.

Each operator supplies lesson 3's data: 20 demonstrations, then four rounds of five runs of their own current clone with every frame labelled by that operator, 7025 frames for A and 7124 for B. Pool the two sets and fit lesson 1's regression, the kernel-weighted mean of the actions stored within three bandwidths (h = 0.02 rad), and score it on 200 runs under the gust.

datareaches the mat95 % intervalends at a post
A alone, 20 demonstrations and four rounds of labels96.5 %93.0 to 98.3 %3.5 %
A and B pooled53.5 %46.6 to 60.3 %46.5 %

Pooling takes back 43.0 points, more than the four rounds had added to A's 20 calm demonstrations (63.5 %, lesson 3's six in ten). The runs that end at a post end at all five: 11, 32, 8, 8 and 34 of the 200 at posts 1 to 5.

Look at the first fork. 350 stored frames lie within reach of the start pose. Turn each action into the velocity it gives the cup (forward, up; cm/s) and average by kernel weight, operator by operator: A's frames give (12.1, +8.8), B's (11.7, −9.3). The regression averages over both, (11.9, +1.9), 9.2° above the axis because B holds 38.1 % of the weight in reach. The post is 20 cm ahead and the cup starts level with its centre, so that heading passes 20 × tan 9.2° = 3.2 cm from the centre, inside the 5 cm halo. Nobody took this action: the nearest stored action is 3.9 cm/s away, while one operator's frames scatter by 0.9 cm/s.

Is the pooled regression under-fit? Score it by lesson 1's copy error, on fresh calm demonstrations. One operator's regression errs by 3.5 % on its own operator's frames; the pooled one by 28.3 % on A's and 26.5 % on B's, and by 29.9 % on the very frames it was fitted to. A better fit cannot help: one action cannot equal two labels.

The routes are far apart for most of the course and share a neighbourhood only at the forks. Call a pose shared if each operator holds at least a tenth of the kernel weight there. Shared poses make up 16.4 % of the route's length, and the expert spends 32 of its 186 steps in them. Away from the forks the regression is right; at them it is wrong, and five forks decide the run.

2 · What squared error asks for

Lesson 1 left a promise: at any one input the best single output of a squared-error fit is the mean of the actions the data give there. At a fork that mean is the trouble. Let the frames in reach have actions ai and kernel weights wi, and let c be the answer:

L(c) = Σ wi ‖ai − c‖² / Σ wi, minimised at c* = Σ wi ai / Σ wi, and L(c) = L(c*) + ‖c − c*‖²

Set the derivative to zero and the weighted mean appears; every other answer pays its squared distance from it. Take two routes, A's action aA holding a share p of the weight and B's aB the rest, a distance d apart. The mean costs L(c*) = p(1 − p)d², and an answer on A's route costs (1 − p)d², which is 1/p times as much. The loss prefers the compromise whatever p is. The data agree (a route here is the weighted mean of one operator's frames in reach):

at the start poseat the first crossing
weight of B's frames within reach38.1 %59.9 %
squared error at A's route ÷ at the mean1.612.49
squared error at B's route ÷ at the mean2.601.67
absolute error at A's route ÷ at the mean0.891.26
absolute error at B's route ÷ at the mean1.330.86

At the start pose A holds 1 − 0.381 = 0.619 of the weight and 1/0.619 = 1.62, against 1.61 measured; at the crossing 1/(1 − 0.599) = 2.49, as measured. Squared error is lowest at the mean at both forks. Why is the loss shaped so? Lesson 1 reached it as maximum likelihood under one assumption: the label is the function's value plus Gaussian scatter of one width, so the actions at an observation form one hump. At a fork they form two, and the single hump that fits them best sits between them, on neither. The fault is in the family of outputs, not in the optimiser: a loss that returns one action cannot return two.

Two other single answers deserve a computation. Absolute error. The answer minimising Σ wi |ai − c| is the weighted median of each component, the point with half the weight on each side. One might expect it to stay between the routes. It does not (the last two rows): with unequal weights it goes to the heavier one, A's at the start pose and B's at the first crossing, and it reaches the mat in 93.5 % of runs. The most likely action. Fit two humps to the frames in reach (§4 shows how) and answer with the centre of the heavier. It reaches the mat in 98.5 % of runs, the highest number in this lesson, so there is something to rule out, and §3 does it.

3 · A distribution, drawn from

What the data show at an observation is a set of actions with weights: the frames within reach, weighted by the kernel. Divided by their sum, the weights are a distribution, p(a | o) = Σ wi δ(a − ai) / Σ wi, with mass wi / Σ wi on each stored action. Lesson 1's regression reported its mean. A policy can report the distribution and act by drawing from it: pick a stored frame in reach with probability proportional to its weight and take its action.

A head must (i) put weight on two separate actions at one observation; (ii) draw well inside the 50 ms budget of a step (lesson 1); (iii) train on the frames we have, which carry no tag saying who drove them (a tag would make each person a function again, but someone must choose it at run time); (iv) agree with lesson 1 where the data show one route. The kernel draw visits the same frames as the mean, so it costs the same, and its expectation is the mean: 200 000 draws at the start pose average to within 0.0003 rad/s of it.

headwhat it returns at the start posereaches the maton B's side of post 1reversals per course (§5)
mean (squared error)between the routes, every time53.5 %–0.00
median (absolute error)A's route, the heavier here93.5 %17.0 %3.09
heaviest humpA's route98.5 %21.0 %3.17
draw one stored frameB's route in 35.3 % of 400 draws, between in 0.0 %73.5 %51.5 %12.68

The draws land on B's route about as often as B's weight in reach (35.3 against 38.1 %), and both routes are in use: B's side of the first post is taken in 51.5 % of the draw's runs, against 17.0 % for the median and 21.0 % for the heaviest hump. The draw reaches the mat in 73.5 % of runs (67.0 to 79.1 %) against 53.5 % for the mean; across eight datasets drawn the same way, 300 runs each, 69.7 % against 55.6 %. A draw is not better in every way: on one operator's data it reaches the mat in 87.5 % of runs where the mean reaches 96.5 % (83.5 against 94.3 % over the eight datasets), because a draw returns one neighbour's action and so adds the spread of the labels in reach to every step. Its virtue is only that it does not average.

Road not taken · the most likely action
The heaviest hump reaches the mat in 98.5 % of pooled runs (97.1 % over eight datasets), more than any head that draws. On this Bench it is the best answer, and its faults are ones the Bench cannot show. It copies whoever holds more weight in reach, A at the start pose and B at the first crossing, so which operator it follows is set by how many frames each left near that pose, not by anything about the people, and it flips abruptly where their weights cross (3.17 reversals per course). It discards the other: B's side of the first post is taken in 21.0 % of runs at a 50 % share of B and in 1.0 % at 30 %. And where the choice has a reason (a post that is sometimes absent, a route blocked today, an operator the policy was asked to imitate) it cannot take the alternative, because it never could. Here either side clears every post, so nothing punishes it. A policy that draws keeps both routes available and needs a rule for when to draw, the next lesson; the heaviest hump returns there as one extreme of that rule.

4 · Heads that do not store frames

A trained network cannot keep every frame; it must emit the distribution's parameters. Two ways are standard, and both can be built from the frames in reach to see what they can express.

A mixture of Gaussians (Bishop's mixture density network, 1994): the head emits K weights πk, centres μk and spreads σk, p(a | o) = Σ πk N(a; μk, σk² I), is trained by maximum likelihood, and acts by drawing a component by its weight and then a point from it. Here the frames in reach are grouped into K clusters by four rounds of weighted k-means; a cluster's weight is its share of the kernel weight, its centre the weighted mean, its spread the rms distance over √2. K = 1 is lesson 1's answer with a spread added and does no better than the mean: 48.5 % reach the mat, and 33.8 % of its draws land between the routes. K = 2 puts one hump on each route: 74.0 %, no draw between; four and eight components change nothing.

A codebook (RT-1, Brohan et al., 2022, discretises every action dimension into 256 bins; Behavior Transformers, Shafiullah et al., 2022, run k-means over all the actions in the data, predict a bin and add an offset): the head emits a categorical distribution over K prototype actions. Here the prototypes come from k-means over every stored action and the categorical is the kernel-weighted histogram of the frames in reach over them, which is what a classifier trained by cross-entropy converges to. There is no offset, so a prototype is a rounded action.

Kmixture: reaches the matmixture: draws between the routescodebook: reaches the matcodebook: runs ending at post 1
148.5 %33.8 %––
274.0 %0.0 %3.0 %177
468.5 %0.0 %26.5 %51
871.0 %0.0 %70.5 %6
16––76.5 %4
32––75.5 %4
64––76.5 %1

A codebook is the kernel draw with its action rounded to the nearest prototype: it cannot beat the draw, and equals it once the rounding is finer than the 5 cm margin. Two prototypes cannot thread a post: 177 of 200 runs end at the first. Two other families learn the distribution by sampling alone. Implicit BC (Florence et al., 2021) learns an energy over observation and action and answers with the action that minimises it, found by search; Diffusion Policy (Chi et al., 2023) draws by iterative denoising. Neither is built here; both keep every hump and pay for a search or a chain of denoising steps at every step.

The widget

Two operators, five heads, one fork
Top: 14 of the 200 runs of the chosen head (teal passes post 1 above it, A's side; amber below, B's; red = ended at a post); dashed, A's route (grey) and B's (amber). Bottom left: the chosen fork in cup velocities (forward, up; cm/s): stored frames in reach (teal A, amber B), 400 draws (rings), the kernel mean (red); "between the routes" is the middle half of the vertical gap from B's mean action to A's. Bottom right: the vertical velocity commanded in the first three runs. The first control sets B's share of the weight.
reaches the mat
—
95 % interval
—
ends at a post
—
on B's side of post 1
—
reversals per course
—
draws between the routes
—
draws on B's route
—
B's weight in reach of the fork
—
route length in a shared neighbourhood
—
frames stored
—
Show the core JS
if (d2 <= 9) { this.ids[m++] = id; SCRATCH[id] = this.c[id] * Math.exp(-0.5 * d2); }
...
var u = rng() * S.tot, k = 0;
while (k < S.m - 1 && u > S.w[k]) { u -= S.w[k]; k++; }
return [S.a0[S.ids[k]], S.a1[S.ids[k]]];
...
var mx = HL.fit(S, K), u = rng(), c = 0;
while (c < mx.k - 1 && u > mx.pi[c]) { u -= mx.pi[c]; c++; }
return [mx.m0[c] + mx.sd[c] * BN.randn(rng), mx.m1[c] + mx.sd[c] * BN.randn(rng)];
...
for (k = 0; k < S.m; k++) HIST[cb.bin[S.ids[k]]] += S.w[k];
u = rng() * S.tot; k = 0;
while (k < cb.K - 1 && u > HIST[k]) { u -= HIST[k]; k++; }
return cb.cen[k].slice();

What to try. Leave the defaults: the mean head, B holding half the weight, kernel 0.020, the start pose. It reaches the mat in 53.5 % of runs and 46.5 % end at a post; the scatter shows two humps of stored frames and the red mean between them, where 100.0 % of its answers lie, with B holding 38.1 % of the weight in reach. Slide the share to 0: one hump is left and the mean reaches the mat in 96.5 % of runs. Choose draw one stored frame at 50 %: 73.5 % reach the mat, 0.0 % of the draws land between the routes and 35.3 % on B's, there are 12.68 reversals per course and run 1 flips between the humps; at a share of 10 % the draw reaches 71.0 %. Choose heaviest hump: 98.5 %, 3.17 reversals, B's side taken in 21.0 % of runs and in 1.0 % at a share of 30. At the first crossing B's weight is 59.9 %. Choose the mixture: K = 1 gives 48.5 % with 33.8 % of the draws between the routes, K = 2 74.0 %, K = 8 71.0 %. Choose the codebook: K = 2 gives 3.0 %, K = 16 76.5 %, K = 64 76.5 %. Return to the draw and narrow the kernel to 0.010: 98.0 %, and the shared neighbourhood shrinks from 16.4 % to 11.9 % of the route; with the mean head at 0.030 the pooled data give 36.5 % and the neighbourhood grows to 21.8 %.

5 · What drawing at every step costs

A draw is independent of the draw before it. At the start pose B holds 38.1 % of the weight, so at each step the head picks B's action with probability about 0.38 and A's with 0.62, and the next step starts again. Count reversals: consecutive steps whose vertical cup velocities are both at least 5 cm/s and of opposite sign. Per course of five forks the mean makes 0.00, one operator drawn from makes 0.12, the heaviest hump 3.17 and two operators drawn from 12.68.

What does a reversal cost? A route's vertical speed is about 9 cm/s at the start and 11.5 cm/s at the first crossing, so a step moves the cup 0.45 to 0.58 cm sideways. If each step's direction is an independent fair coin, the sideways position is a random walk (lesson 2): after n steps its spread is 0.45 √n cm, while a cup that chose a direction and kept it has moved 0.45 n cm. A fork's shared neighbourhood lasts 32 / 5 = 6.4 steps on average (§1), in which the coin-flipper gains 0.45 √6.4 = 1.1 cm sideways and a committed cup 2.9 cm; and a crossing lies only 10 cm, about 17 steps, before the next post, with 5 cm of clearance still to find. The coin here is biased (62 against 38), which adds a drift toward the likelier route and does not remove the spread. This arithmetic says which way every knob pushes, not how many runs end in a post, which depends on the gust and the margin; the tables below are the evidence. Chi et al. (2023) describe the same fault in sequence models: actions drawn independently at consecutive steps can come from different modes, giving jittery actions that alternate between the two valid trajectories.

The collisions follow: on the pooled data the draw ends at a post in 26.5 % of 200 runs, and in 30.3 % averaged over eight datasets (27.7 to 34.0 %). The mean's 46.5 % was a route that cannot work; this is a route that works and is not kept. As B's share of the data grows from 0 to 10 and 50 %, the mean falls from 96.5 to 90.5 and 53.5 %, steadily; the draw loses 14.0 points at once, 87.5 to 71.0 %, and ends at 73.5 %, because drawing between routes costs the same however many frames the second route has; the heaviest hump does not notice (99.0, 98.0, 98.5 %).

One knob moves all of this: the kernel width, which sets how far apart two frames can lie and still be in each other's reach, and so how much of the route is shared. The track has used 0.020 since lesson 1.

kernel width hroute in a shared neighbourhoodone operator, meanpooled, meanpooled, drawpooled, heaviest hump
0.010 rad11.9 %100.0 %92.5 %98.0 %99.0 %
0.020 rad16.4 %96.5 %53.5 %73.5 %98.5 %
0.030 rad21.8 %81.5 %36.5 %35.0 %94.5 %
Road not taken · a sharper kernel
If the damage is a matter of width, narrow the kernel: at 0.010 the pooled mean reaches the mat in 92.5 % of runs and the draw in 98.0 %. On this Bench that is true, and it is a fact about the Bench, a two-number observation covered by 14 000 frames. It is no fix for the fork, where both operators are asked about the same observation and no width separates a pose from itself, and a policy that reads an image has no frame within a kernel that narrow of where it is (lesson 2).
What this lesson did not do
It built each head from the frames in reach of the current observation, by looking them up, and not the heads that learn the distribution by sampling alone. A network that emits a mixture or a codebook (Bishop, 1994; Shafiullah et al., 2022) must learn it from the same frames and has failures a table cannot show, such as components that collapse onto one route and bins that are rarely visited. The two operators are interchangeable, and a person's habits, speed and slips are not; the Bench cannot tell a policy that copies one operator from one that copies both, which is why the most likely action scores best here. It did not say how long a policy should stay on a route once it has drawn one (lesson 5), what happens when the plant's gain is wrong (lessons 5 and 6), or how many layouts the data must cover (lesson 7).

Common mistakes / failure modes

"the best single answer for two good routes is one of them"
Squared error is lowest at the weighted mean: at the start pose the loss at A's route is 1.61 times the loss at the mean, at B's 2.60 times, and no better fit helps (§1, §2).
"drawing is the mean plus noise"
The expectation is the mean, but the mean is always between the routes and a draw never is: 100.0 % against 0.0 % (§3).
"a mixture with one component is a safe default"
One hump sits between the routes: 33.8 % of its draws land there and 48.5 % of runs reach the mat (§4).
"take the most likely action: it scores best"
It does, 98.5 %, and it discards an operator: B's side in 1.0 % of runs at a 30 % share (§3).
"if it draws from the right distribution it behaves like the operators"
It draws again at every step: 12.68 reversals per course against 0.12 for one operator, and 26.5 % of runs end in a post (§5).

Checkpoint exercise

Try it
Every 0.05 s a policy draws the cup's vertical velocity, +9 cm/s or −9 cm/s with equal probability and independently of the steps before, while the cup moves forward at 12 cm/s toward a post 10 cm ahead. (a) How far does the cup move sideways in one step? (b) How many steps does it take to reach the post, and how far from its line has the cup typically strayed by then? (c) How many steps until it has typically strayed 5 cm, and how many would a cup need that chose its direction once and kept it? Answer: (a) 9 × 0.05 = 0.45 cm. (b) 10 / 12 s = 0.83 s, 17 steps. The spread of a sum of n independent steps of ±0.45 cm is 0.45 √n, so 0.45 × √17 = 1.9 cm: at the post the coin-flipper is within 2 cm of its line, a cup that had chosen a direction 17 × 0.45 = 7.65 cm from it. (c) 0.45 √n = 5 gives n = (5 / 0.45)² = 123 steps, 6.2 s; the committed cup needs 5 / 0.45 = 11 steps.

Where this points next

A policy that outputs a distribution keeps two routes two routes: at the start pose none of its draws lands between them (0.0 %) and 35.3 % land on B's route, where the mean answered between them. On the pooled data it reaches the mat in 73.5 % of runs against 53.5 % for the mean, and a mixture of two Gaussians or a codebook does as well without storing a frame. But the draw has no memory. It is made afresh at every step, so near a fork the cup is sent up, then down: 12.68 reversals per course against 0.12 for one operator, a sideways spread that grows as the square root of the steps and not as the steps, and 26.5 % of runs that end in a post (30.3 % averaged over eight datasets). Committing to one route for a while would stop the dithering, at the price of not looking while committed. How long should a policy commit, and what should it commit to?

Takeaway
A regression answers every observation with one action, and where two good actions exist the cheapest single answer is their average, which nobody took: a second operator takes the pooled regression from 96.5 % to 53.5 %, and its error on its own training frames is 29.9 %, a floor and not under-fit. A policy that outputs a distribution with a hump for each route, and acts by drawing from it, keeps the routes apart: no draw lands between them, 35.3 % land on B's route against B's 38.1 % of the weight, and the pooled draw reaches the mat in 73.5 % of runs; a mixture of two Gaussians does the same and a codebook needs 8 to 16 prototypes. The heaviest hump scores best (98.5 %) and discards an operator. What a draw lacks is memory: it is made afresh each step, 12.68 reversals per course, and 26.5 % of runs end in a post.

Interview prompts

Companion reads: World Models · 24 Spreads, not averages (the same move for predictions), World Models · 05 One future is a lie (honest probabilities on separate possibilities) and Reinforcement Learning · 17 Imitation and IRL (behaviour cloning from the MDP side).