Targets, not velocities: feedback in the action
A chunk of velocity commands drifts from where it meant the arm to be: 2.14 cm after 16 steps, on a course that leaves 1.55 cm of room, and on a worn arm no chunk length succeeds in more than 41 % of runs. This lesson changes what a chunk holds, from how to move to where to be, with a stiff controller closing the gap every 50 ms: the drift stops growing with the chunk and the course is completed at every chunk length. The price is that the targets are exact about the world they were recorded in, and with the posts shifted 2 cm the tracker walks into them.
New idea: a chunk says where the arm is to be, in the world's coordinates, and a stiff controller closes the gap at every step, so an error loses a share k of itself per step instead of keeping all of it. The plant no longer limits the chunk length, and the output is as exact as the world it was recorded in.
Forces next: A chunk of target positions, tracked by a stiff controller that pulls the arm back to the target at every step, removes the accumulation: on the demonstrated layout the policy completes the course nearly every time at any chunk length. But targets are anchored to where the demonstrations were. Shift the vases by a few centimetres and the policy walks faithfully into them. The policy has to read where the vases are and generalise across layouts, and nobody has said how many layouts that takes. How does success on a new layout grow with the data, and which data buy it?
1 · The drift of a velocity chunk is a sum
Keep one route, so that lesson 5's fork is out of the picture, and reuse lesson 1's 20 calm demonstrations on five posts. The policy looks up the stored frame nearest to the arm, in joint angles, and copies the chunk of commands stored with it, one per control step (dt = 0.05 s); H = 1 looks again at every step, H = 235 replays the whole course. Two arms run it: the new arm of lesson 1 (each joint delivers the velocity it is told to, gain g = 1, under gusts of σ = 0.05 rad/s) and a worn arm that delivers 85 % of it (g = 0.85) under gusts of 0.08 rad/s.
How far can a copy stray? The demonstrated path comes within 6.55 cm of the nearest post's centre and contact is at 5 cm, so it clears the posts by 1.55 cm, the room of lesson 5 (lesson 1 drew a path 10 cm off the centres; the expert cuts the corners). The drift after a chunk is the distance at the cup between where the arm is when a chunk ends and where the chunk meant it to be: its starting pose plus what its commands integrate to. 200 runs per cell.
| chunk length H | new arm: drift | new arm: reaches the mat | worn arm: drift | worn arm: reaches the mat |
|---|---|---|---|---|
| 1 step | 0.20 cm | 76.0 % | 0.35 cm | 31.0 % |
| 4 steps | 0.41 cm | 97.0 % | 0.79 cm | 41.0 % |
| 16 steps | 0.83 cm | 93.0 % | 2.14 cm | 17.5 % |
| 64 steps | 1.54 cm | 67.5 % | 6.43 cm | 0.5 % |
| the whole course | 3.30 cm | 44.0 % | 19.36 cm | 0.0 % |
The drift grows with the chunk: as √H on the new arm (the slope of ln drift against ln H from 4 to 64 steps is 0.49), faster on the worn arm. Success first rises with a few steps of commitment (76.0 % to 97.0 % on the new arm), then falls as the drift reaches the room; on the worn arm the drift is past it from 16 steps on and no length gets above 41.0 %. Even at H = 1, looking at the arm again at every step, it ends 1.36 cm off the path: a lookup returns how the expert moved at the nearest frame, not where the path is, so no loop pulls the arm back (§2 measures that).
Why that size? A joint at angle q, told the velocity u, obeys q ← q + dt · (g u + σ ξ), with ξ ~ N(0, 1) per joint per step, while the pose the chunk means advances by dt · u. The error e between them obeys
et+1 = 1 · et + (g − 1) · dt · ut + dt σ · ξt
Every step adds a term and none takes one away. The share of an error still there a step later, the loop gain J, is 1: a velocity never mentions where the arm is, so nothing can undo what the plant added. At the cup, let c = 0.82 m/rad be how far the cup moves per radian of error in each joint, the two joints added in quadrature (lesson 5's c, here along the one demonstrated path), and l = 0.74 cm how far it moves per step along the path. After n steps the random part is dt σ c √n = 0.33 cm × √n at σ = 0.08 (0.20 at 0.05), the systematic part is (1 − g) times the displacement the chunk asked for, 0.11 cm × n at g = 0.85. In quadrature
drift(n) = √( (dt σ c)² · n + ((1 − g) · l · n)² )
It gives 2.20 cm against the 2.14 measured at 16 steps on the worn arm, and 0.79 against 0.79 at 4: within 3 % up to 16 steps, and an upper bound beyond, since the chord of a long chunk is shorter than its arc on a bending path. Set equal to the room it gives the chunk length H* at which the drift of a velocity chunk has spent it on average: 10 steps on the worn arm and 58 on the new, where the table's drift reaches 1.55 cm (1.54 at 64 steps). Success falls earlier, because a run is lost by its worst chunk and not its average one; lesson 5's budget, which also keeps twice the random part in reserve, is the safer length.
2 · Feed the error back: a target and a tracker
What would break the sum? Three requirements. The output must mention a quantity the plant's noise moves, or no comparison can show a deviation. The correction must be made at every control step, inside the chunk, without asking the policy, which is not consulted again for H steps. And it must not need the plant's gain or delay, which are not in the demonstrations. A rate fails the first; a position meets it. So let the chunk be a list of targets T0, …, TH−1, the pose the arm should have after 1, …, H steps, and let a tracker carry them out. (Targets that nothing tracks are velocities in disguise: executing them blind means commanding (Tj − Tj−1)/dt, the chunk of §1 with J = 1.)
ut = K · (Tj − qt), clipped to ±1.5 rad/s, j = t mod H
K, per second, is the stiffness. The command is now a function of where the arm is. With k = K · dt · g, the share of the gap one step closes, qt+1 = qt + k (T − qt) + dt σ ξt, so against a target that stands still
et+1 = (1 − k) · et + dt σ · ξt, spread = dt σ / √(1 − J²) = dt σ / √( k (2 − k) ) per joint
The loop gain is J = 1 − k. An error of age i has been multiplied by (1 − k)i, so the variance is (dt σ)² Σ (1 − k)2i = (dt σ)² / (k (2 − k)), whatever the length of the chunk. With K = 20 per second and dt = 0.05 s, k = g: 1 on the new arm (J = 0), 0.85 on the worn arm (J = 0.15). The gusts then leave a spread of 0.20 cm at the cup on the new arm and 0.33 cm on the worn arm, a fifth of the room. Measured on the worn arm, the drift after a chunk of targets is between 0.32 and 0.38 cm at every chunk length, and the 19.36 cm of the table is gone.
A shove measures J directly. Push both joints by 0.012 rad (a centimetre at the cup) in the middle of a 16-step chunk on a calm arm and compare with the unshoved run. Under velocities the share of the shove left after one step is 1.00, and after eight (still inside the chunk) 1.00 (0.92 to 1.07 over 48 shoves). Under the tracker (K = 20, worn arm, k = 0.85) it is 0.15 and 0.00. Looking again at every step (H = 1) is not that loop either: twenty steps after the shove, 1.13 times its size is left on average.
3 · What is a target measured from?
A position needs an origin, and there are two honest choices. A target can be measured from the arm (a relative target): integrate the stored commands from where the arm is at the start of the chunk, Tj = q0 + dt Σi≤j ui. It needs nothing new in the data and is the smallest change to lesson 5's chunk: the same velocities, tracked instead of applied. Or it can be measured from the world (an absolute target): the poses the demonstrated arm went through next, whatever the arm did before. The same tracker carries out both; they differ at the start of the next chunk. Worn arm, K = 20, 200 runs per cell:
| chunk length H | 1 | 2 | 4 | 8 | 16 |
|---|---|---|---|---|---|
| targets from the arm: reaches the mat | 31.0 % | 69.5 % | 96.0 % | 100.0 % | 100.0 % |
| targets from the world: reaches the mat | 100.0 % | 100.0 % | 100.0 % | 100.0 % | 100.0 % |
| from the arm: distance from the path at the end | 1.36 cm | 0.64 | 0.36 | 0.26 | 0.20 |
A target measured from the arm is anchored to the arm: whatever error it carries when a chunk ends is the next chunk's starting point, the tracker holds it there, and nothing pulls it back to the path. The errors of successive chunks add, and at H = 1 it is the velocity policy again (31.0 % either way); the longer the chunk, the fewer the starting points that leak, and from 8 steps on the arm finishes within 0.26 cm of the path and completes every run. A target measured from the world is anchored to the demonstrations, and costs nothing they do not already hold, since the poses are the data. The first target of every chunk is on the path whatever the arm did, so the error is the tracker's spread and nothing else: 0.15 to 0.19 cm from the path at every chunk length, and 100 % of the runs complete.
Practice shows the same contrast. ALOHA's policy outputs absolute joint positions that a low-level PID controller tracks, and its authors report degraded performance with delta joint positions (Zhao et al., 2023); Diffusion Policy reports that position control suffers less from compounding error than velocity control (Chi et al., 2023). The Bench shows two mechanisms that could lie behind this, the sum of §1 and the leak above.
4 · What stiffness costs
A larger k closes the gap faster and leaves a smaller spread, until it does not. A joint answers late. The command computed now reaches the motor d steps later (the network, the servo's own loop), so the tracker acts on the position of d steps ago: et+1 = et − k et−d. For d = 0 the loop is stable for k below 2. For d = 1 the characteristic equation is z² − z + k = 0, whose roots have modulus √k, so k must stay below 1. In general the loop is stable only for
k < 2 sin( π / (4d + 2) )
The roots of zd+1 − zd + k and a direct simulation of the loop give the same ceiling (0.618 and 0.618 for d = 2). The Bench shows it in success (targets from the world, 64-step chunks, worn arm):
| delay d | ceiling on k | below the ceiling | above it |
|---|---|---|---|
| 1 step (50 ms) | 1.00 | K = 20: 98.0 % | K = 30: 0.0 % |
| 2 steps (100 ms) | 0.62 | K = 10: 98.5 % | K = 15: 1.0 % |
| 3 steps (150 ms) | 0.45 | K = 6: 99.0 % | K = 12: 0.0 % |
A soft tracker lags. With k < 1 the arm trails its target by about (1 − k)/k steps, and every new chunk starts from where the arm is, so each chunk loses about that many steps of progress. The 20 demonstrations, from jittered starts, average 184 steps of the 235 the course allows, so a chunk can afford to lose about a fifth of its length: it needs H above about 4.6 (1 − k)/k, which is 17, 6 and 3 steps for K = 5, 10 and 15. The Bench agrees (targets from the world, worn arm, no delay): K = 5 completes 13.0 % of the runs at 16 steps and 98.5 % at 32, K = 10 0.0 % at 4 steps and 100.0 % at 8, K = 15 0.0 % at 1 step and 100.0 % at 4. Every failure listed is a timeout, not a collision: the arm is accurate and slow. So stiff means k near 1, below the ceiling the delay sets.
5 · Units and resolution
A velocity is in rad/s and means move at this speed for the next 50 ms. What that does depends on dt, the gain and the delay between command and joint: replayed on another plant the same commands mean another motion (§1). A target is in radians, a unit the plant does not change: its gain changes how fast the gap closes, not where the arm ends up.
How precisely must a target be written? The room is 1.55 cm and a radian of error in each joint is worth c = 0.82 m at the cup, so a target must be right to 0.019 rad, 1.1 degrees. A policy that emits bins (RT-1, Brohan et al., 2022, uses 256 per action dimension) rounds to the middle of a bin, with an error up to half a bin. Over ±π, 256 bins are 0.0245 rad wide, half a bin is 0.0123 rad, and at the cup along the demonstrated path that is up to 1.47 cm, 95 % of the room. Rounding the stored targets to bins (chunks of 8 steps, worn arm, 200 runs): 256 bins reach the mat in 100 % of runs, 128 bins (worst error 3.00 cm) in 86 %, 64 bins (5.53 cm) in 56 %. Bins that span only the range a joint uses buy resolution for free.
The widget
What to try. At the defaults (velocity chunks of 16 steps, worn arm) the drift is 2.14 cm against a clearance of 1.55 cm and 17.5 % of runs reach the mat; at 4 steps 41.0 %, the best a velocity chunk does here, at the whole course 0.0 %. Switch to targets from the world (K = 20): 100.0 % even for the whole course, with a drift of 0.38 cm. Targets from the arm give 31.0 % at 1 step and 100.0 % from 8 steps. Back on targets from the world, set K = 5: 0.0 % at 1 step, 99.0 % at 64. Put K back to 20 and add a delay of 2 steps, with 64-step chunks: K = 20 and K = 15 fail (0.0 % and 1.0 %; the tracker readout says unstable), K = 10 gives 98.5 %. Remove the delay, go back to 16 steps and set the gain to 0.7: targets from the world still reach 100.0 %, velocity chunks 0.0 %. Put the gain back to 0.85 and shift the posts: +1 cm gives 100.0 % with 0.79 cm of clearance, +2 cm gives 0.5 % with −0.02 cm.
6 · Precise, and anchored to the old world
The policy is now a table from joint angles to targets, recorded on one layout. The tracker holds the arm within 0.38 cm of the path the table describes, and that path clears the posts by 1.55 cm. Shift the whole row of posts up by s, the arm still starting where the demonstrations started. The path the arm follows does not move, because the table does not contain the posts, and its clearance falls as the posts come to it (worn arm; targets from the world in chunks of 16 steps):
| shift s | clearance of the demonstrated path | targets from the world: reaches the mat |
|---|---|---|
| −2 cm | 0.06 cm | 5.5 % |
| −1 cm | 0.87 cm | 99.5 % |
| 0 | 1.55 cm | 100.0 % |
| +1 cm | 0.79 cm | 100.0 % |
| +2 cm | −0.02 cm | 0.5 % |
At +2 cm the path the tracker follows touches a post and 99.5 % of the runs touch one, however stiff the tracker. The collapse comes where the clearance reaches zero, at a shift of about 2.0 cm, since 0.79 cm of clearance is lost per centimetre of shift. The accuracy that removed the drift is accuracy about the old world. The layout is in the table and not in the input, as in lessons 1 to 5; a policy that survives a 2 cm shift either reads where the posts are or has seen layouts that differ by less, and what that takes is the next lesson.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
A chunk of targets measured from the world, tracked by a stiff controller, turns the sum of §1 into a loop that forgets: on the demonstrated layout the policy completes the course in 100 % of its runs at every chunk length, on a worn arm where no velocity chunk gets past 41 %. What it buys is exactness about the demonstrations, which is also its limit. Its path clears the posts by 1.55 cm; shifted by 2 cm the clearance is gone and the success with it (0.5 %), because the tracker walks faithfully into the vases. The policy has to read where the vases are and generalise across layouts, and nobody has said how many layouts that takes. How does success on a new layout grow with the data, and which data buy it?
Interview prompts
- Why does a chunk of velocity commands drift more the longer it is, and how fast? (§1 — the plant adds a gust and a gain error at every step and a velocity never refers to position: loop gain 1, a random part growing as √H and a systematic part as H.)
- What is a loop gain, and how would you measure it on a robot? (§2 — the share of an error left a step later; shove the arm mid-chunk and see what remains.)
- Why measure a target from the world and not from the arm? (§3 — from the arm every chunk starts from the last chunk's error, and nothing pulls the arm back to the path.)
- A tracker is made stiffer and the arm starts to oscillate; made softer, it runs out of time. Why? (§4 — a delay makes the correction act on a stale position, stable only for k below 2 sin(π/(4d + 2)); a soft tracker lags and every chunk start loses the lag.)
- How precisely must an action be written, and what does that say about tokenising it? (§5 — to the room over the arm's reach, 0.019 rad here; 256 bins over ±π leave up to 0.0123 rad.)
- Tracked targets complete the demonstrated course every time and fail when the obstacles move 2 cm. Why? (§6 — the targets are the demonstrations' world and the tracker is exact about it.)
Companion reads: Reinforcement Learning · 71 Industrial control (the PID loops real plants already trust).