Commit: action chunks
A policy that draws a fresh action at every step still dithers where two routes fork: the two-route head of this lesson ends in a post in 19.5 % of the courses, lesson 4's sampler of stored frames in 30.5 %. This lesson changes what one draw produces. Draw the route once and play the next H commands of that route without looking, and the policy decides once per fork instead of about 7.9 times: at H = 12 the route changes 0.8 times a course instead of 13.0, and success rises from 80.5 % to 94.5 %. The price is that the policy stops looking, and the plant's noise piles up until it uses the room the course leaves, 1.55 cm: at H = 40 the gain has gone (74.5 %).
New idea: a chunk: draw the route once, by the kernel's mass, then play the next H commands of that route (the kernel-weighted mean of their stored plans) without looking, and draw again after H steps. One decision per fork, paid for with H steps of open loop.
Forces next: Sampling a whole sequence of actions at once and executing it commits the policy to one route, and the dithering stops. But the sequence is a list of velocities, a velocity command is added to the joint angles at every step, and every bit of noise in the plant stays added for the whole chunk: the longer the chunk, the further the arm drifts from where the sequence meant it to be, and the gain from committing is paid back as drift. What should the policy output, so that the plant's noise does not accumulate while it commits?
1 · A fork, counted
Lesson 4's data once more: two operators, A passing the first post above and B below, each with 20 calm demonstrations and four rounds of five rollouts of their own clone under the gust, every state visited labelled by its operator. That is 7025 and 7124 stored frames, five posts, gust 0.05 rad/s. Each frame also carries its route, A or B, because we built the data; a person's recording would not say, and the head below uses the tag only to group frames, as lesson 4's clustering does.
Two heads draw at every step. The plain sampler draws one stored frame near the state, in proportion to its kernel weight, and plays its command. The two-route head draws a route in proportion to the kernel mass of that route's frames and plays that route's kernel-mean command: lesson 4's two-component mixture, with the routes as the components. On 200 courses they end in a post in 30.5 % and 19.5 %; averaged over eight data sets of the same kind, 30.8 % and 21.6 %, which brackets the quarter lesson 4 ended on (a count of 200 courses wanders by about three points; lesson 4 counted 26.5 %). The head wins because it plays the mean of a route, not one frame's command (§3). Drawn at every step, it is the baseline from here on.
Where does it go wrong? Along each operator's own clean route let p be the share of the kernel's mass on route B: 0 or 1 away from the posts, passing through the middle at each fork. Call the steps with 0.05 ≤ p ≤ 0.95 a fork window. The course has 10 of them, 7.4 steps long on average (4 to 10). A rollout of the per-step head visits about 4.7 windows of about 7.9 steps and flips a coin with the kernel's odds in every step of each. Over a course the route changes 13.0 times, in 169 draws.
That is the dithering, and it needs no plant noise. With the gust off the per-step head still fails 10.5 % of the courses, and a route drawn once at the start and followed with the loop closed fails 0 %. Inside a window the commands of the two routes send the cup to opposite sides of a post; a fresh coin every step sends it one way and then the other, and a slow net drift is what is left. With no gust the cup crosses a window sideways at 0.22 cm per step when the route is redrawn every step and at 0.50 cm per step when it is redrawn every 12.
One more number, because the second half of the lesson spends it: the room. Lessons 1 to 4 counted a margin of 5 cm: the path runs 10 cm off each post and a touch is at 5 cm. The expert's cup does not follow the path exactly. It aims at a point 6 cm ahead on it and cuts the corners, so it passes closest to a post centre at 6.55 cm, and a perfect course leaves 6.55 − 5 = 1.55 cm.
2 · What a commitment has to be
The question has two halves, what to commit to and for how long. Four requirements bear on the first; three come from earlier lessons and the fourth is the price. (1) It must come out of the recordings: the data are states and the commands that followed (lesson 1), and nothing in them names the route a frame belongs to, so what is held cannot be a route label. (2) It must keep both routes at the frequencies the data show; a policy that can go one way only has thrown away what lesson 4 bought. (3) It must belong to one route: the mean of the two routes' commands points at the post, and a blend of plans is the same fault in slow motion. (4) It has a price: while it is held the policy is not reading the state. The table runs the candidates on the same 200 courses at gust 0.05; the last column counts how many of the 200 went above the first post, a check that the data's mix of routes survives.
| what is played | courses completed | above the first post |
|---|---|---|
| a fresh draw every step: plain sampler | 69.5 % | 93 |
| a fresh draw every step: two-route head (the baseline) | 80.5 % | 97 |
| always the likelier route (τ = 0) | 96.5 % | 166 |
| route drawn once, state re-read every step | 95.5 % | 114 |
| average of the last 12 plans (temporal ensembling) | 9.0 % | 80 |
| one stored plan, played for 12 steps | 55.0 % | 116 |
| the route's mean plan, played for 12 steps | 94.5 % | 114 |
3 · A chunk is the mean plan of one route
Give every stored frame a plan: the next H commands its operator issues from that state on a perfect plant, ui,1 … ui,H. In a recording these are the next H recorded commands; the Bench recomputes them from the stored state, which also covers frames the clone reached. At a state q the kernel weights are ki = exp(−|q − qi|² / 2h²) with h = 0.02 rad, and the route masses MA and MB are the sums of ki over each route's frames.
The head draws route r with probability Mr / (MA + MB) and plays the kernel-weighted mean plan of that route, û1…H = Σi∈r ki ui,1…H / Σi∈r ki; after H steps it draws again. H = 1 is the per-step head.
Why the mean and not one stored plan? A plan carries a correction: a neighbour sits a little off its route and its plan starts by steering back, which has the wrong size at our state. Take 1000 states along rollouts, away from the forks, and measure how far the cup ends H steps later under the head's plan from the cup under the operator's own plan from the same state, both on a perfect plant. At H = 12 a copied plan misses by 0.88 cm and the route's mean plan by 0.35 cm (at 40: 1.29 and 0.53). Neighbours lie on both sides of the route, so their corrections cancel in the mean: lesson 4's rule again, average within a mode and never across. In the table, 55.0 % against 94.5 %.
A chunk draws once per H steps: per fork visit the head draws 7.9 times at H = 1, 2.0 at 4, 1.0 at 8 and 0.7 at 12 (fewer than one: some chunks begin before the window and pass through it). Published chunks are of the same order: Diffusion Policy predicts 16 steps and executes 8 (Chi et al., 2023), ACT predicts 100 at 50 Hz (Zhao et al., 2023), π0 predicts 50 and executes 16 at 20 Hz (Black et al., 2024). A step is 0.05 s: 12 steps are 0.6 s.
The widget
What to try. At the defaults (H = 1, gust 0.05, gain 1.00) the per-step head completes 80.5 % of 200 courses, changes route 13.0 times a course and draws 169 times. Drag H to 12: 94.5 %, 0.8 route changes, 15 draws, rms drift 0.70 cm; the green area between the curve and its H = 1 level is the gain. Drag on to 40: 74.5 %, drift 1.24 cm, and the area has turned red. Set the gust to 0: the curve reaches 100 % at H = 10 and stays at 99 % or more, and the budget reads none. At gust 0.08 the budget (5.7 steps) is below the fork window and the best chunk gains only 12.0 points. With gust 0 and gain 0.85, step H from 12 (98.0 %) to 20 (0.0 %): a cliff no noise causes. Then switch what is played to one stored plan (55.0 % at 12) and to the average of the last H plans (9.0 %): the table's failed rows.
4 · What committing buys
| chunk length H | route changes per course | courses completed |
|---|---|---|
| 1 | 13.0 | 80.5 % |
| 8 | 1.4 | 94.0 % |
| 12 | 0.8 | 94.5 % |
| 40 | 0.4 | 74.5 % |
The left of the curve is the fork. A chunk shorter than the window (7.4 steps) still draws more than once inside it (2.0 times per visit at H = 4); a chunk that covers it draws about once, route changes fall to about one a course, and success reaches 94.0 %.
5 · What it costs, and where H sits
While the chunk plays the policy is not looking and the plant keeps doing what it does. The gust adds σ times a unit normal to each joint velocity, so a step of dt = 0.05 s adds σ dt times a unit normal to each joint angle, and nothing in the chunk takes it back: a velocity is added to the joint angles, the loop has gain 1, as in lesson 2. A joint error becomes a cup error through the arm's Jacobian; let c = 0.81 m/rad be how far the cup moves per radian of error in each joint, the two joints added in quadrature, along the routes. One step moves the cup off its planned place by c σ dt = 0.20 cm in a random direction, and independent steps add in square: after H steps the rms drift is c σ dt √H.
A plant that does less than it is told adds a second drift. With gain g it executes g u and falls short by a share 1 − g of each planned displacement, l = 0.74 cm per step (15 cm/s is 0.75 cm a step, less where the expert slows). That adds up linearly, (1 − g) l H, so rms drift² = (c σ dt)² H + ((1 − g) l)² H². On the Bench the drift is the distance between the cup when the next draw comes and where the plan, read as joint angles through the arm's kinematics, meant it to be:
| chunk length H | measured rms drift | the law |
|---|---|---|
| 1, gust 0.05 | 0.20 cm | 0.20 cm |
| 12, gust 0.05 | 0.70 cm | 0.70 cm |
| 40, gust 0.05 | 1.24 cm | 1.28 cm |
| 20, no gust, gain 0.85 | 2.15 cm | 2.22 cm |
The drift has to fit in the room, m = 1.55 cm. A chunk is safe when the systematic part plus two standard deviations of the random part stay inside it: (1 − g) l H + 2 c σ dt √H ≤ m. The H that makes this an equality is the drift budget; with no gain error it is (m / 2c σ dt)², 14.6 steps at gust 0.05, where 2c σ dt = 0.405 cm. The table gives it for six plants beside the success of the route's mean plan at five chunk lengths, on the same 200 courses.
| plant | budget | H = 1 | 8 | 12 | 20 | 40 |
|---|---|---|---|---|---|---|
| no gust | none | 89.5 | 99.0 | 99.0 | 100.0 | 99.0 |
| gust 0.05 | 14.6 | 80.5 | 94.0 | 94.5 | 91.5 | 74.5 |
| gust 0.08 | 5.7 | 59.5 | 71.5 | 66.5 | 66.5 | 40.0 |
| gust 0.05, gain 0.95 | 9.0 | 75.5 | 89.0 | 82.5 | 76.5 | 25.0 |
| gust 0.05, gain 0.85 | 5.4 | 74.5 | 70.0 | 69.0 | 11.5 | 0.0 |
| no gust, gain 0.85 | 14.0 | 88.0 | 97.5 | 98.0 | 0.0 | 0.0 |
Read it by rows. No gust, no budget: the curve rises and stays, since without plant noise nothing pays the gain back. At gust 0.05 there is a plateau between the window (7.4) and the budget (14.6), and at 40 steps success, 74.5 %, is back at the level of H = 1 (80.5 %); on 600 courses the three are 78.5, 92.7 and 75.2 %. A louder gust puts the budget (5.7) below the window and the plateau nearly goes: the best chunk gains 12.0 points. A gain error adds a term that grows as H: at 0.95 success peaks at 89.0 % and falls to 25.0 % at 40; at 0.85 the budget (5.4) is below the window and 20 steps complete 11.5 %. The last row is the cleanest. With no gust the drift is deterministic: a 12-step chunk drifts 1.29 cm and completes 98.0 %, a 20-step chunk drifts 2.15 cm, more than the room, and completes 0.0 %. The cliff sits where (1 − g) l H reaches m, at H = 14.0; the per-step policy does not notice such a plant, since its error only slows it down (88.0 % at H = 1). On 600 courses the best chunk of every plant lies within a factor of two of its budget: the budget says where the plateau ends and is not an edge. ACT's chunk-size ablation has the same rise and taper: 1 % at k = 1, 44 % at 100, a slight fall beyond (Zhao et al., 2023).
Common mistakes / failure modes
Checkpoint exercise
Where this points next
A policy that draws a whole chunk decides once per fork: 0.8 route changes a course at H = 12 against 13.0, 94.5 % against 80.5 %. The gust adds a drift that grows as √H and a plant that does less than it is told adds one that grows as H, until the room of 1.55 cm is gone: at 40 steps success is back to 74.5 %, and with a gain of 0.85 a 20-step chunk completes 11.5 %. The chunk is a list of velocities, and a velocity is added to the joint angles, so what the plant adds stays. What should the policy output, so that the plant's noise does not accumulate while it commits?
Interview prompts
- A policy samples a fresh action at every step and the task has two valid routes. Why does it dither, and why does a noise-free plant not cure it? (§1 — near a fork each draw is a coin with the kernel's odds between commands that point to opposite sides of the obstacle; with the gust off it still fails 10.5 % of the courses.)
- Why is playing one stored action sequence not committing to a route? (§3 — a neighbour's plan carries its own correction, wrong at our state; the route's mean plan averages it away: 55.0 % against 94.5 %.)
- Why can averaging overlapping chunks (temporal ensembling) fail on a multimodal task? (§2 — the chunks come from independent draws, so their average is an average over routes: 9.0 % at H = 12.)
- Derive how the open-loop drift of an H-step chunk grows with H for gust noise and for a gain error. (§5 — independent steps add in square, c σ dt √H; a gain error adds (1 − g) l per step, (1 − g) l H.)
- Given a plant's noise and the room a task leaves, how would you choose a chunk length? (§5 — longer than the fork window, shorter than the budget where (1 − g) l H + 2 c σ dt √H reaches the room.)
- Why does a plant that executes 85 % of every command break a 20-step chunk but not a per-step policy? (§5 — the per-step policy re-reads the state and is only slowed; the chunk is blind for 20 steps and the shortfall, 0.15 × 0.74 × 20 = 2.2 cm, exceeds the 1.55 cm room.)
Companion reads: World Models · 09 Planning in latent worlds (play a sequence, then replan: the same trade between looking and committing).