all_lessons/Robot Model Training/05 · Chunkslesson 5 / 24

Commit: action chunks

A policy that draws a fresh action at every step still dithers where two routes fork: the two-route head of this lesson ends in a post in 19.5 % of the courses, lesson 4's sampler of stored frames in 30.5 %. This lesson changes what one draw produces. Draw the route once and play the next H commands of that route without looking, and the policy decides once per fork instead of about 7.9 times: at H = 12 the route changes 0.8 times a course instead of 13.0, and success rises from 80.5 % to 94.5 %. The price is that the policy stops looking, and the plant's noise piles up until it uses the room the course leaves, 1.55 cm: at H = 40 the gain has gone (74.5 %).

The thesis, here
A policy that must choose between two routes should choose once per fork, not once per step. Draw the route, then play the next H commands of that route. While the chunk plays the policy is not looking, so what the plant adds stays: the drift grows as √H and, once it uses up the room, committing costs more than it gains. H lies between the length of a fork (7.4 steps) and the length the drift can afford (14.6 steps at the default noise, fewer when the plant is worse).
Linear position
Forced by: A policy that outputs a distribution over actions, and acts by sampling it, recovers both routes instead of their collision. But it draws a new sample at every step, and near the fork each draw can pick a different route: the arm dithers between them and a quarter of the rollouts still end in a post. Committing to one route for a while would stop the dithering, at the price of not looking while committed. How long should a policy commit, and what should it commit to?
New idea: a chunk: draw the route once, by the kernel's mass, then play the next H commands of that route (the kernel-weighted mean of their stored plans) without looking, and draw again after H steps. One decision per fork, paid for with H steps of open loop.
Forces next: Sampling a whole sequence of actions at once and executing it commits the policy to one route, and the dithering stops. But the sequence is a list of velocities, a velocity command is added to the joint angles at every step, and every bit of noise in the plant stays added for the whole chunk: the longer the chunk, the further the arm drifts from where the sequence meant it to be, and the gain from committing is paid back as drift. What should the policy output, so that the plant's noise does not accumulate while it commits?
The plan
Five moves. (1) Count the dithering: where a fork is, how many draws land in it, how much room the course leaves. (2) State what a commitment must satisfy and measure the candidates, including those that look better on success. (3) Make the commitment a route's mean plan. (4) See what H steps buy. (5) Derive what they cost, the drift of an open-loop chunk against the room, and put H between the two.

1 · A fork, counted

Lesson 4's data once more: two operators, A passing the first post above and B below, each with 20 calm demonstrations and four rounds of five rollouts of their own clone under the gust, every state visited labelled by its operator. That is 7025 and 7124 stored frames, five posts, gust 0.05 rad/s. Each frame also carries its route, A or B, because we built the data; a person's recording would not say, and the head below uses the tag only to group frames, as lesson 4's clustering does.

Two heads draw at every step. The plain sampler draws one stored frame near the state, in proportion to its kernel weight, and plays its command. The two-route head draws a route in proportion to the kernel mass of that route's frames and plays that route's kernel-mean command: lesson 4's two-component mixture, with the routes as the components. On 200 courses they end in a post in 30.5 % and 19.5 %; averaged over eight data sets of the same kind, 30.8 % and 21.6 %, which brackets the quarter lesson 4 ended on (a count of 200 courses wanders by about three points; lesson 4 counted 26.5 %). The head wins because it plays the mean of a route, not one frame's command (§3). Drawn at every step, it is the baseline from here on.

Where does it go wrong? Along each operator's own clean route let p be the share of the kernel's mass on route B: 0 or 1 away from the posts, passing through the middle at each fork. Call the steps with 0.05 ≤ p ≤ 0.95 a fork window. The course has 10 of them, 7.4 steps long on average (4 to 10). A rollout of the per-step head visits about 4.7 windows of about 7.9 steps and flips a coin with the kernel's odds in every step of each. Over a course the route changes 13.0 times, in 169 draws.

That is the dithering, and it needs no plant noise. With the gust off the per-step head still fails 10.5 % of the courses, and a route drawn once at the start and followed with the loop closed fails 0 %. Inside a window the commands of the two routes send the cup to opposite sides of a post; a fresh coin every step sends it one way and then the other, and a slow net drift is what is left. With no gust the cup crosses a window sideways at 0.22 cm per step when the route is redrawn every step and at 0.50 cm per step when it is redrawn every 12.

One more number, because the second half of the lesson spends it: the room. Lessons 1 to 4 counted a margin of 5 cm: the path runs 10 cm off each post and a touch is at 5 cm. The expert's cup does not follow the path exactly. It aims at a point 6 cm ahead on it and cuts the corners, so it passes closest to a post centre at 6.55 cm, and a perfect course leaves 6.55 − 5 = 1.55 cm.

2 · What a commitment has to be

The question has two halves, what to commit to and for how long. Four requirements bear on the first; three come from earlier lessons and the fourth is the price. (1) It must come out of the recordings: the data are states and the commands that followed (lesson 1), and nothing in them names the route a frame belongs to, so what is held cannot be a route label. (2) It must keep both routes at the frequencies the data show; a policy that can go one way only has thrown away what lesson 4 bought. (3) It must belong to one route: the mean of the two routes' commands points at the post, and a blend of plans is the same fault in slow motion. (4) It has a price: while it is held the policy is not reading the state. The table runs the candidates on the same 200 courses at gust 0.05; the last column counts how many of the 200 went above the first post, a check that the data's mix of routes survives.

what is playedcourses completedabove the first post
a fresh draw every step: plain sampler69.5 %93
a fresh draw every step: two-route head (the baseline)80.5 %97
always the likelier route (τ = 0)96.5 %166
route drawn once, state re-read every step95.5 %114
average of the last 12 plans (temporal ensembling)9.0 %80
one stored plan, played for 12 steps55.0 %116
the route's mean plan, played for 12 steps94.5 %114
Road not taken · sharpen the draw
Raise each route's odds to the power 1/τ and renormalise. A small τ, or the argmax (lesson 4's heaviest hump), stops the dithering because it stops drawing, and it succeeds (95.0 % and 96.5 %): on success alone it matches the chunk. It fails requirement 2. At the start state the data put 38.5 % of the mass on route B (38.1 % in lesson 4, which weighted the two operators' frames to count equally); τ = 0.25 turns that into 13.2 %, τ = 0.1 into 0.9 % and the argmax into none. The argmax policy goes above the first post in 166 courses of 200, where a policy that draws once from the data's mix does so in 114: a route is gone, and with it what lesson 4 was for.
Road not taken · hold the route, keep the loop closed
Draw the route once and re-read the state at every step, playing that route's mean command: 95.5 % for the whole course, as good as the chunk, and no drift, because the loop never opens. It fails requirement 1: what it holds is a route label, which the Bench has because it built the data and a person's recordings do not. A commitment cut out of recordings is a stretch of commands.
Road not taken · average overlapping plans
Make a plan at every step and play, for this step, the weighted mean of what the last H plans prescribe, as ACT does (Zhao et al., 2023): 9.0 % at H = 12. Here each plan is a fresh draw, so their average is lesson 4's average of two routes spread over time. The π0 authors report that temporal ensembling hurt and play their chunks open loop (Black et al., 2024).

3 · A chunk is the mean plan of one route

Give every stored frame a plan: the next H commands its operator issues from that state on a perfect plant, ui,1 … ui,H. In a recording these are the next H recorded commands; the Bench recomputes them from the stored state, which also covers frames the clone reached. At a state q the kernel weights are ki = exp(−|q − qi|² / 2h²) with h = 0.02 rad, and the route masses MA and MB are the sums of ki over each route's frames.

The head draws route r with probability Mr / (MA + MB) and plays the kernel-weighted mean plan of that route, û1…H = Σi∈r ki ui,1…H / Σi∈r ki; after H steps it draws again. H = 1 is the per-step head.

Why the mean and not one stored plan? A plan carries a correction: a neighbour sits a little off its route and its plan starts by steering back, which has the wrong size at our state. Take 1000 states along rollouts, away from the forks, and measure how far the cup ends H steps later under the head's plan from the cup under the operator's own plan from the same state, both on a perfect plant. At H = 12 a copied plan misses by 0.88 cm and the route's mean plan by 0.35 cm (at 40: 1.29 and 0.53). Neighbours lie on both sides of the route, so their corrections cancel in the mean: lesson 4's rule again, average within a mode and never across. In the table, 55.0 % against 94.5 %.

A chunk draws once per H steps: per fork visit the head draws 7.9 times at H = 1, 2.0 at 4, 1.0 at 8 and 0.7 at 12 (fewer than one: some chunks begin before the window and pass through it). Published chunks are of the same order: Diffusion Policy predicts 16 steps and executes 8 (Chi et al., 2023), ACT predicts 100 at 50 Hz (Zhao et al., 2023), π0 predicts 50 and executes 16 at 20 Hz (Black et al., 2024). A step is 0.05 s: 12 steps are 0.6 s.

The widget

Commit for H steps
Top: both routes dashed (teal A, purple B) and 12 of the 200 courses coloured by the route in force, a red dot a hit; under it one row per course, fork windows shaded. Bottom left: courses completed against H (log axis), the H = 1 level dashed, green where committing gains and red where the gain is paid back. Bottom right: route changes per course, and the §5 drift (systematic plus twice the random part) against the room.
courses completed
—
95 % interval
—
ended in a post
—
route changes per course
—
draws per course
—
rms drift of a chunk
—
completed at H = 1
—
against H = 1
—
drift budget
—
Show the core JS
CL.meanPlan = function (D, r, H) {
  var sx = new Float64Array(H), sy = new Float64Array(H), tot = 0, i, k;
  for (i = 0; i < D.cnt; i++) {
    var id = D.ids[i]; if (D.route[id] !== r) continue;
    var wj = D.ws[i], base = id * CL.MAXH * 2; tot += wj;
    for (k = 0; k < H; k++) { sx[k] += wj * D.pl[base + 2 * k]; sy[k] += wj * D.pl[base + 2 * k + 1]; }
  }
...
  var seq = []; for (k = 0; k < H; k++) seq.push([sx[k] / tot, sy[k] / tot]); return seq;
};
...
  var r = rng() < CL.pRoute(ms.m, tau) ? 1 : 0, seq = CL.meanPlan(D, r, H);
...
    if (buf.length === 0) {
...
      buf = dr.seq.map(function (u) { return [u[0], u[1]]; }); hist.push([t, dr.route]);
...
    }
    return buf.shift();

What to try. At the defaults (H = 1, gust 0.05, gain 1.00) the per-step head completes 80.5 % of 200 courses, changes route 13.0 times a course and draws 169 times. Drag H to 12: 94.5 %, 0.8 route changes, 15 draws, rms drift 0.70 cm; the green area between the curve and its H = 1 level is the gain. Drag on to 40: 74.5 %, drift 1.24 cm, and the area has turned red. Set the gust to 0: the curve reaches 100 % at H = 10 and stays at 99 % or more, and the budget reads none. At gust 0.08 the budget (5.7 steps) is below the fork window and the best chunk gains only 12.0 points. With gust 0 and gain 0.85, step H from 12 (98.0 %) to 20 (0.0 %): a cliff no noise causes. Then switch what is played to one stored plan (55.0 % at 12) and to the average of the last H plans (9.0 %): the table's failed rows.

4 · What committing buys

chunk length Hroute changes per coursecourses completed
113.080.5 %
81.494.0 %
120.894.5 %
400.474.5 %

The left of the curve is the fork. A chunk shorter than the window (7.4 steps) still draws more than once inside it (2.0 times per visit at H = 4); a chunk that covers it draws about once, route changes fall to about one a course, and success reaches 94.0 %.

5 · What it costs, and where H sits

While the chunk plays the policy is not looking and the plant keeps doing what it does. The gust adds σ times a unit normal to each joint velocity, so a step of dt = 0.05 s adds σ dt times a unit normal to each joint angle, and nothing in the chunk takes it back: a velocity is added to the joint angles, the loop has gain 1, as in lesson 2. A joint error becomes a cup error through the arm's Jacobian; let c = 0.81 m/rad be how far the cup moves per radian of error in each joint, the two joints added in quadrature, along the routes. One step moves the cup off its planned place by c σ dt = 0.20 cm in a random direction, and independent steps add in square: after H steps the rms drift is c σ dt √H.

A plant that does less than it is told adds a second drift. With gain g it executes g u and falls short by a share 1 − g of each planned displacement, l = 0.74 cm per step (15 cm/s is 0.75 cm a step, less where the expert slows). That adds up linearly, (1 − g) l H, so rms drift² = (c σ dt)² H + ((1 − g) l)² H². On the Bench the drift is the distance between the cup when the next draw comes and where the plan, read as joint angles through the arm's kinematics, meant it to be:

chunk length Hmeasured rms driftthe law
1, gust 0.050.20 cm0.20 cm
12, gust 0.050.70 cm0.70 cm
40, gust 0.051.24 cm1.28 cm
20, no gust, gain 0.852.15 cm2.22 cm

The drift has to fit in the room, m = 1.55 cm. A chunk is safe when the systematic part plus two standard deviations of the random part stay inside it: (1 − g) l H + 2 c σ dt √H ≤ m. The H that makes this an equality is the drift budget; with no gain error it is (m / 2c σ dt)², 14.6 steps at gust 0.05, where 2c σ dt = 0.405 cm. The table gives it for six plants beside the success of the route's mean plan at five chunk lengths, on the same 200 courses.

plantbudgetH = 18122040
no gustnone89.599.099.0100.099.0
gust 0.0514.680.594.094.591.574.5
gust 0.085.759.571.566.566.540.0
gust 0.05, gain 0.959.075.589.082.576.525.0
gust 0.05, gain 0.855.474.570.069.011.50.0
no gust, gain 0.8514.088.097.598.00.00.0

Read it by rows. No gust, no budget: the curve rises and stays, since without plant noise nothing pays the gain back. At gust 0.05 there is a plateau between the window (7.4) and the budget (14.6), and at 40 steps success, 74.5 %, is back at the level of H = 1 (80.5 %); on 600 courses the three are 78.5, 92.7 and 75.2 %. A louder gust puts the budget (5.7) below the window and the plateau nearly goes: the best chunk gains 12.0 points. A gain error adds a term that grows as H: at 0.95 success peaks at 89.0 % and falls to 25.0 % at 40; at 0.85 the budget (5.4) is below the window and 20 steps complete 11.5 %. The last row is the cleanest. With no gust the drift is deterministic: a 12-step chunk drifts 1.29 cm and completes 98.0 %, a 20-step chunk drifts 2.15 cm, more than the room, and completes 0.0 %. The cliff sits where (1 − g) l H reaches m, at H = 14.0; the per-step policy does not notice such a plant, since its error only slows it down (88.0 % at H = 1). On 600 courses the best chunk of every plant lies within a factor of two of its budget: the budget says where the plateau ends and is not an edge. ACT's chunk-size ablation has the same rise and taper: 1 % at k = 1, 44 % at 100, a slight fall beyond (Zhao et al., 2023).

What this lesson did not do
It measured one learner, kernel regression over stored frames with the route known, on one course with two routes. A person's recordings carry no route label and the lesson learned none, so the head here is the cleanest case of what a learned head should compute. The drift law is first-order: white noise, a Jacobian constant along the route, one number for a room that differs from post to post; the budget is a rule of thumb, not an edge. Published chunk lengths are tuned on their tasks, not derived from a budget. It did not try to stop the drift: that is lesson 6.

Common mistakes / failure modes

"a chunk is a sequence sampled from the data"
Playing one stored plan completes 55.0 % at H = 12, worse than drawing every step (80.5 %): the plan carries its owner's correction. The route's mean plan completes 94.5 % (§3).
"committing fixes the dithering for free"
It stops looking. Success falls from 94.5 % at H = 12 to 74.5 % at 40, as the rms drift goes from 0.70 to 1.24 cm against a room of 1.55 cm (§5).
"just make the draw greedy"
The argmax completes 96.5 % and sends 166 of 200 courses above the first post, where drawing once from the data's mix sends 114. It stops the dithering by deleting a route (§2).
"with a noise-free plant chunks are free"
Only for noise. A plant that does g = 0.85 of each command, with no noise at all, completes 98.0 % at 12 steps and 0.0 % at 20 (§5).

Checkpoint exercise

Try it
A plant adds a gust of 0.03 rad/s to each joint. Take c = 0.81 m/rad, dt = 0.05 s, l = 0.74 cm per step and a room of 1.55 cm. (a) What is the rms drift of a 16-step chunk? (b) What is the longest chunk whose twice-rms drift fits in the room? (c) The plant also executes only 97 % of each command: what is the budget now? Answer: (a) c σ dt √H = 0.81 × 0.03 × 0.05 × 4 m = 0.49 cm. (b) (m / 2c σ dt)² = (1.55 / 0.243)² = 41 steps. (c) The shortfall is 0.03 × 0.74 = 0.022 cm per step, so 0.022 H + 0.243 √H = 1.55; with x = √H this is 0.022 x² + 0.243 x − 1.55 = 0, x = 4.5 and H = 20 steps: a 3 % gain error halves the budget.

Where this points next

A policy that draws a whole chunk decides once per fork: 0.8 route changes a course at H = 12 against 13.0, 94.5 % against 80.5 %. The gust adds a drift that grows as √H and a plant that does less than it is told adds one that grows as H, until the room of 1.55 cm is gone: at 40 steps success is back to 74.5 %, and with a gain of 0.85 a 20-step chunk completes 11.5 %. The chunk is a list of velocities, and a velocity is added to the joint angles, so what the plant adds stays. What should the policy output, so that the plant's noise does not accumulate while it commits?

Takeaway
Where two routes fork, a policy that draws at every step flips a coin in each of the 7.9 steps of the window. A commitment must come out of the recordings, keep both routes and belong to one of them: the next H commands of a drawn route, averaged over that route's stored plans. A chunk costs looking: the plant adds a random c σ dt = 0.20 cm per step that sums as √H and a gain error (1 − g) l per step that sums as H, against 1.55 cm of room. So H lies between the fork window (7.4 steps) and the budget (14.6 at the default noise, less with a louder gust or a gain error); where the budget falls below the window there is nothing to gain.

Interview prompts

Companion reads: World Models · 09 Planning in latent worlds (play a sequence, then replan: the same trade between looking and committing).