People are not functions
Lesson 3 left a clone that reaches the mat on about nineteen runs in twenty when one operator's frames build it (96.5 % on this lesson's draw). Add a second operator, who passes the first post on the other side, pool both sets of frames and fit the same regression, and it reaches the mat on 53.5 %: at the start pose one operator's action takes the cup up and the other's down, and squared error returns their average, an action nobody took, aimed at the post. This lesson changes what a policy outputs: not one action but a distribution over actions, with a hump for each route the data show, drawn from at every step. A draw lands on a route and never between two. But it is made again at every step, so near a fork the arm is sent up, then down, and 26.5 % of runs still end in a post.
New idea: output a distribution over actions, with a hump for each route the data show, and act by drawing from it. The two routes stay two routes, and every step now decides again.
Forces next: A policy that outputs a distribution over actions, and acts by sampling it, recovers both routes instead of their collision. But it draws a new sample at every step, and near the fork each draw can pick a different route: the arm dithers between them and a quarter of the rollouts still end in a post. Committing to one route for a while would stop the dithering, at the price of not looking while committed. How long should a policy commit, and what should it commit to?
1 · A second operator
The expert of the last three lessons passes the first post above it (larger y), the second below it, and so on down the row. Call it operator A. A second operator, B, does the opposite: below the first post, above the second. Both clear every post by the same 5 cm, and their routes are mirror images that cross midway between each pair of posts. There the same pose is asked for opposite actions, A's taking the cup down and B's up. So is the start pose, where both begin. Five posts, five such forks: the start pose and the four crossings.
Each operator supplies lesson 3's data: 20 demonstrations, then four rounds of five runs of their own current clone with every frame labelled by that operator, 7025 frames for A and 7124 for B. Pool the two sets and fit lesson 1's regression, the kernel-weighted mean of the actions stored within three bandwidths (h = 0.02 rad), and score it on 200 runs under the gust.
| data | reaches the mat | 95 % interval | ends at a post |
|---|---|---|---|
| A alone, 20 demonstrations and four rounds of labels | 96.5 % | 93.0 to 98.3 % | 3.5 % |
| A and B pooled | 53.5 % | 46.6 to 60.3 % | 46.5 % |
Pooling takes back 43.0 points, more than the four rounds had added to A's 20 calm demonstrations (63.5 %, lesson 3's six in ten). The runs that end at a post end at all five: 11, 32, 8, 8 and 34 of the 200 at posts 1 to 5.
Look at the first fork. 350 stored frames lie within reach of the start pose. Turn each action into the velocity it gives the cup (forward, up; cm/s) and average by kernel weight, operator by operator: A's frames give (12.1, +8.8), B's (11.7, −9.3). The regression averages over both, (11.9, +1.9), 9.2° above the axis because B holds 38.1 % of the weight in reach. The post is 20 cm ahead and the cup starts level with its centre, so that heading passes 20 × tan 9.2° = 3.2 cm from the centre, inside the 5 cm halo. Nobody took this action: the nearest stored action is 3.9 cm/s away, while one operator's frames scatter by 0.9 cm/s.
Is the pooled regression under-fit? Score it by lesson 1's copy error, on fresh calm demonstrations. One operator's regression errs by 3.5 % on its own operator's frames; the pooled one by 28.3 % on A's and 26.5 % on B's, and by 29.9 % on the very frames it was fitted to. A better fit cannot help: one action cannot equal two labels.
The routes are far apart for most of the course and share a neighbourhood only at the forks. Call a pose shared if each operator holds at least a tenth of the kernel weight there. Shared poses make up 16.4 % of the route's length, and the expert spends 32 of its 186 steps in them. Away from the forks the regression is right; at them it is wrong, and five forks decide the run.
2 · What squared error asks for
Lesson 1 left a promise: at any one input the best single output of a squared-error fit is the mean of the actions the data give there. At a fork that mean is the trouble. Let the frames in reach have actions ai and kernel weights wi, and let c be the answer:
L(c) = Σ wi ‖ai − c‖² / Σ wi, minimised at c* = Σ wi ai / Σ wi, and L(c) = L(c*) + ‖c − c*‖²
Set the derivative to zero and the weighted mean appears; every other answer pays its squared distance from it. Take two routes, A's action aA holding a share p of the weight and B's aB the rest, a distance d apart. The mean costs L(c*) = p(1 − p)d², and an answer on A's route costs (1 − p)d², which is 1/p times as much. The loss prefers the compromise whatever p is. The data agree (a route here is the weighted mean of one operator's frames in reach):
| at the start pose | at the first crossing | |
|---|---|---|
| weight of B's frames within reach | 38.1 % | 59.9 % |
| squared error at A's route ÷ at the mean | 1.61 | 2.49 |
| squared error at B's route ÷ at the mean | 2.60 | 1.67 |
| absolute error at A's route ÷ at the mean | 0.89 | 1.26 |
| absolute error at B's route ÷ at the mean | 1.33 | 0.86 |
At the start pose A holds 1 − 0.381 = 0.619 of the weight and 1/0.619 = 1.62, against 1.61 measured; at the crossing 1/(1 − 0.599) = 2.49, as measured. Squared error is lowest at the mean at both forks. Why is the loss shaped so? Lesson 1 reached it as maximum likelihood under one assumption: the label is the function's value plus Gaussian scatter of one width, so the actions at an observation form one hump. At a fork they form two, and the single hump that fits them best sits between them, on neither. The fault is in the family of outputs, not in the optimiser: a loss that returns one action cannot return two.
Two other single answers deserve a computation. Absolute error. The answer minimising Σ wi |ai − c| is the weighted median of each component, the point with half the weight on each side. One might expect it to stay between the routes. It does not (the last two rows): with unequal weights it goes to the heavier one, A's at the start pose and B's at the first crossing, and it reaches the mat in 93.5 % of runs. The most likely action. Fit two humps to the frames in reach (§4 shows how) and answer with the centre of the heavier. It reaches the mat in 98.5 % of runs, the highest number in this lesson, so there is something to rule out, and §3 does it.
3 · A distribution, drawn from
What the data show at an observation is a set of actions with weights: the frames within reach, weighted by the kernel. Divided by their sum, the weights are a distribution, p(a | o) = Σ wi δ(a − ai) / Σ wi, with mass wi / Σ wi on each stored action. Lesson 1's regression reported its mean. A policy can report the distribution and act by drawing from it: pick a stored frame in reach with probability proportional to its weight and take its action.
A head must (i) put weight on two separate actions at one observation; (ii) draw well inside the 50 ms budget of a step (lesson 1); (iii) train on the frames we have, which carry no tag saying who drove them (a tag would make each person a function again, but someone must choose it at run time); (iv) agree with lesson 1 where the data show one route. The kernel draw visits the same frames as the mean, so it costs the same, and its expectation is the mean: 200 000 draws at the start pose average to within 0.0003 rad/s of it.
| head | what it returns at the start pose | reaches the mat | on B's side of post 1 | reversals per course (§5) |
|---|---|---|---|---|
| mean (squared error) | between the routes, every time | 53.5 % | – | 0.00 |
| median (absolute error) | A's route, the heavier here | 93.5 % | 17.0 % | 3.09 |
| heaviest hump | A's route | 98.5 % | 21.0 % | 3.17 |
| draw one stored frame | B's route in 35.3 % of 400 draws, between in 0.0 % | 73.5 % | 51.5 % | 12.68 |
The draws land on B's route about as often as B's weight in reach (35.3 against 38.1 %), and both routes are in use: B's side of the first post is taken in 51.5 % of the draw's runs, against 17.0 % for the median and 21.0 % for the heaviest hump. The draw reaches the mat in 73.5 % of runs (67.0 to 79.1 %) against 53.5 % for the mean; across eight datasets drawn the same way, 300 runs each, 69.7 % against 55.6 %. A draw is not better in every way: on one operator's data it reaches the mat in 87.5 % of runs where the mean reaches 96.5 % (83.5 against 94.3 % over the eight datasets), because a draw returns one neighbour's action and so adds the spread of the labels in reach to every step. Its virtue is only that it does not average.
4 · Heads that do not store frames
A trained network cannot keep every frame; it must emit the distribution's parameters. Two ways are standard, and both can be built from the frames in reach to see what they can express.
A mixture of Gaussians (Bishop's mixture density network, 1994): the head emits K weights πk, centres μk and spreads σk, p(a | o) = Σ πk N(a; μk, σk² I), is trained by maximum likelihood, and acts by drawing a component by its weight and then a point from it. Here the frames in reach are grouped into K clusters by four rounds of weighted k-means; a cluster's weight is its share of the kernel weight, its centre the weighted mean, its spread the rms distance over √2. K = 1 is lesson 1's answer with a spread added and does no better than the mean: 48.5 % reach the mat, and 33.8 % of its draws land between the routes. K = 2 puts one hump on each route: 74.0 %, no draw between; four and eight components change nothing.
A codebook (RT-1, Brohan et al., 2022, discretises every action dimension into 256 bins; Behavior Transformers, Shafiullah et al., 2022, run k-means over all the actions in the data, predict a bin and add an offset): the head emits a categorical distribution over K prototype actions. Here the prototypes come from k-means over every stored action and the categorical is the kernel-weighted histogram of the frames in reach over them, which is what a classifier trained by cross-entropy converges to. There is no offset, so a prototype is a rounded action.
| K | mixture: reaches the mat | mixture: draws between the routes | codebook: reaches the mat | codebook: runs ending at post 1 |
|---|---|---|---|---|
| 1 | 48.5 % | 33.8 % | – | – |
| 2 | 74.0 % | 0.0 % | 3.0 % | 177 |
| 4 | 68.5 % | 0.0 % | 26.5 % | 51 |
| 8 | 71.0 % | 0.0 % | 70.5 % | 6 |
| 16 | – | – | 76.5 % | 4 |
| 32 | – | – | 75.5 % | 4 |
| 64 | – | – | 76.5 % | 1 |
A codebook is the kernel draw with its action rounded to the nearest prototype: it cannot beat the draw, and equals it once the rounding is finer than the 5 cm margin. Two prototypes cannot thread a post: 177 of 200 runs end at the first. Two other families learn the distribution by sampling alone. Implicit BC (Florence et al., 2021) learns an energy over observation and action and answers with the action that minimises it, found by search; Diffusion Policy (Chi et al., 2023) draws by iterative denoising. Neither is built here; both keep every hump and pay for a search or a chain of denoising steps at every step.
The widget
What to try. Leave the defaults: the mean head, B holding half the weight, kernel 0.020, the start pose. It reaches the mat in 53.5 % of runs and 46.5 % end at a post; the scatter shows two humps of stored frames and the red mean between them, where 100.0 % of its answers lie, with B holding 38.1 % of the weight in reach. Slide the share to 0: one hump is left and the mean reaches the mat in 96.5 % of runs. Choose draw one stored frame at 50 %: 73.5 % reach the mat, 0.0 % of the draws land between the routes and 35.3 % on B's, there are 12.68 reversals per course and run 1 flips between the humps; at a share of 10 % the draw reaches 71.0 %. Choose heaviest hump: 98.5 %, 3.17 reversals, B's side taken in 21.0 % of runs and in 1.0 % at a share of 30. At the first crossing B's weight is 59.9 %. Choose the mixture: K = 1 gives 48.5 % with 33.8 % of the draws between the routes, K = 2 74.0 %, K = 8 71.0 %. Choose the codebook: K = 2 gives 3.0 %, K = 16 76.5 %, K = 64 76.5 %. Return to the draw and narrow the kernel to 0.010: 98.0 %, and the shared neighbourhood shrinks from 16.4 % to 11.9 % of the route; with the mean head at 0.030 the pooled data give 36.5 % and the neighbourhood grows to 21.8 %.
5 · What drawing at every step costs
A draw is independent of the draw before it. At the start pose B holds 38.1 % of the weight, so at each step the head picks B's action with probability about 0.38 and A's with 0.62, and the next step starts again. Count reversals: consecutive steps whose vertical cup velocities are both at least 5 cm/s and of opposite sign. Per course of five forks the mean makes 0.00, one operator drawn from makes 0.12, the heaviest hump 3.17 and two operators drawn from 12.68.
What does a reversal cost? A route's vertical speed is about 9 cm/s at the start and 11.5 cm/s at the first crossing, so a step moves the cup 0.45 to 0.58 cm sideways. If each step's direction is an independent fair coin, the sideways position is a random walk (lesson 2): after n steps its spread is 0.45 √n cm, while a cup that chose a direction and kept it has moved 0.45 n cm. A fork's shared neighbourhood lasts 32 / 5 = 6.4 steps on average (§1), in which the coin-flipper gains 0.45 √6.4 = 1.1 cm sideways and a committed cup 2.9 cm; and a crossing lies only 10 cm, about 17 steps, before the next post, with 5 cm of clearance still to find. The coin here is biased (62 against 38), which adds a drift toward the likelier route and does not remove the spread. This arithmetic says which way every knob pushes, not how many runs end in a post, which depends on the gust and the margin; the tables below are the evidence. Chi et al. (2023) describe the same fault in sequence models: actions drawn independently at consecutive steps can come from different modes, giving jittery actions that alternate between the two valid trajectories.
The collisions follow: on the pooled data the draw ends at a post in 26.5 % of 200 runs, and in 30.3 % averaged over eight datasets (27.7 to 34.0 %). The mean's 46.5 % was a route that cannot work; this is a route that works and is not kept. As B's share of the data grows from 0 to 10 and 50 %, the mean falls from 96.5 to 90.5 and 53.5 %, steadily; the draw loses 14.0 points at once, 87.5 to 71.0 %, and ends at 73.5 %, because drawing between routes costs the same however many frames the second route has; the heaviest hump does not notice (99.0, 98.0, 98.5 %).
One knob moves all of this: the kernel width, which sets how far apart two frames can lie and still be in each other's reach, and so how much of the route is shared. The track has used 0.020 since lesson 1.
| kernel width h | route in a shared neighbourhood | one operator, mean | pooled, mean | pooled, draw | pooled, heaviest hump |
|---|---|---|---|---|---|
| 0.010 rad | 11.9 % | 100.0 % | 92.5 % | 98.0 % | 99.0 % |
| 0.020 rad | 16.4 % | 96.5 % | 53.5 % | 73.5 % | 98.5 % |
| 0.030 rad | 21.8 % | 81.5 % | 36.5 % | 35.0 % | 94.5 % |
Common mistakes / failure modes
Checkpoint exercise
Where this points next
A policy that outputs a distribution keeps two routes two routes: at the start pose none of its draws lands between them (0.0 %) and 35.3 % land on B's route, where the mean answered between them. On the pooled data it reaches the mat in 73.5 % of runs against 53.5 % for the mean, and a mixture of two Gaussians or a codebook does as well without storing a frame. But the draw has no memory. It is made afresh at every step, so near a fork the cup is sent up, then down: 12.68 reversals per course against 0.12 for one operator, a sideways spread that grows as the square root of the steps and not as the steps, and 26.5 % of runs that end in a post (30.3 % averaged over eight datasets). Committing to one route for a while would stop the dithering, at the price of not looking while committed. How long should a policy commit, and what should it commit to?
Interview prompts
- Why does squared error return the average of two routes, and why is the average wrong? (§2 — every answer pays its squared distance from the weighted mean; with weights p and 1 − p at distance d the mean costs p(1 − p)d² and a route (1 − p)d², so the compromise wins and lies where nobody drove.)
- Absolute error is minimised by the median. Does it keep both routes? (§2 — no: it goes to the heavier one, and which that is changes from fork to fork.)
- What does a mixture head with one component output, and why is it not the mean? (§4 — a Gaussian centred on the mean with a spread: the same centre plus noise, a third of whose draws land between the routes.)
- Why does a codebook of two prototypes fail where one of sixteen does not need more? (§4 — prototypes are rounded actions; two cannot thread a 5 cm margin, and from 8 on the rounding is finer than the margin.)
- A policy that draws from the true distribution of two good operators still ends in a post about a quarter of the time. Why? (§5 — each step draws again, so near a fork the cup wanders, a spread of 0.45√n cm against 0.45n cm for a cup that commits, and loses steps it needs to clear the next post.)
- When is the most likely action the right head? (§3 — when either route serves, as here; not when the lighter route is the one needed, since it discards that operator at every fork.)
Companion reads: World Models · 24 Spreads, not averages (the same move for predictions), World Models · 05 One future is a lie (honest probabilities on separate possibilities) and Reinforcement Learning · 17 Imitation and IRL (behaviour cloning from the MDP side).