all_lessons/Robot Model Training/12 · Rewardlesson 12 / 24

Learn from a score alone: reinforcement learning

The last lesson made a simulator worth using, and both its repairs assumed someone to copy. For most tasks nobody can demonstrate; there is only a way to tell whether the task was done. This lesson asks what that score can teach and what it costs at 30 s an attempt. The one update a score permits is to try the policy with small random changes and keep what scored. On the Bench, from nothing and with the plain score, which counts falling off the table, it never starts: in 15,000 attempts not one scored. A hand-made progress term rescues it after 1,020 attempts; the lesson-1 clone with a small, bounded correction needs 240. What a correction cannot do is leave the one table it was learned on.

The thesis, here
A score improves a policy only by comparison between attempts. The one update it permits is to try the policy with small random changes and move toward the changes that scored better; that gives one number per attempt for every unknown in the policy, and nothing at all when every attempt scores the same. A policy that starts from nothing meets exactly that case when the score is narrow, as most scores that matter are. A policy that starts nearly right, so that small changes sometimes score, does not; the cheap way to learn from a score is to begin there and to search only for a small correction.
Linear position
Forced by: A simulator is worth using when its error is narrower than the policy can tolerate: randomising the physics widens what a policy survives at the price of a more cautious motion, and calibrating the simulator on a few real rollouts narrows the gap itself. Both assume there is someone in the simulator to copy, a scripted expert or a planner. For most tasks that matter there is no such expert, only a way to tell whether the task was done. What can be learned from a score alone, and how many attempts does it take?
New idea: improve a policy from scored attempts alone by trying small random changes and keeping what scored, and make the first attempts informative by starting from a policy that already succeeds and searching only a small, bounded correction to it. On the Bench that costs 240 attempts, 2.0 robot-hours, where the plain score alone never starts, and the correction it finds belongs to its task.
Forces next: Learning from a score alone works, slowly, on a small task, and works in an afternoon of robot time only when it starts from a policy that is already nearly right and learns a small correction. Every source so far feeds that starting point: demonstrations, other bodies, video, force, simulation. It is one model that must read language and images, feel force, emit chunks, and still answer within one control step. How is such a model built, and what does it cost to run it at the speed of the robot?
The plan
Six moves. (1) Fix the score, the price of an attempt and the policy. (2) Derive the update a score permits, and its noise. (3) See what a score of zero teaches. (4) Start where scores differ: the clone, with a bounded correction. (5) Add up the robot-hours. (6) Find where a correction stops helping.

1 · What counts as done, what an attempt costs, what a policy is

A demonstration told the robot what to do in every frame. A score tells it one thing about a whole attempt: 1 if the cup reaches the mat, touches no post and stays on the table, 0 otherwise. Nobody says how. An attempt takes 30 s of robot time, the reset included: 120 an hour, 480 in a four-hour afternoon, and the 3,000 the widget goes to are 25 hours.

The policy is the plainest function whose size we can choose: a table of joint velocities on a grid of joint angles, read by bilinear interpolation. A spacing of 0.1 rad has 19 × 13 nodes of two numbers each, so d = 494 numbers θ (rad/s), the parameters of the policy; spacings of 0.15, 0.3 and 0.5 rad give 234, 70 and 40. Only the entries a run reads can change its score: the clone's own runs read 101 of the 247 nodes, 202 of the numbers. Course, gust and start jitter are those of lesson 1. A point on a learning curve is the share of 100 fixed runs the current policy completes; those runs are the measurement and cost no robot time, the attempts that drive the learning do.

The table has an edge, 15 cm either side of the row of posts, and it has to count. The expert's path is 10 cm off the row, so its margin to the edge is 5 cm, as to a post, and no policy so far came near it. A learner that explores comes near everything, and a score takes whatever it accepts. Draw the table at random, every entry a Gaussian of standard deviation σ, and run it on five posts. Ignore the edge and 1.0 % of draws at σ = 0.2 rad/s and 7.2 % at 0.4 reach the mat: any detour over the posts scores. Count the edge and none of 6,000 draws does, at 0.1, 0.2 or 0.4. The edge turns a plain score into a narrow target, and most scores that matter are narrow. It is this lesson's addition, and what follows depends on it: ignore it, as lessons 1 to 11 did, and a learner from nothing does learn a detour: at noise 0.2 the median seed completes 90 % of the runs after 180 attempts, 1.5 hours (the widget's "edge ignored"). What decides the cost is whether small random changes to the starting policy sometimes score, and the edge is what makes that false for a policy that does nothing.

2 · Keep what scored

The robot cannot differentiate its world: the plant is a black box and the score is a step, with a jump wherever the cup grazes a post. An update has to be built from comparisons. Smooth the score R(θ) of one attempt (the plus-progress score adds the share of the course covered) by trying the policy with a random change of size σ rad/s, J(θ) = E[R(θ + σε)], ε a vector of d independent unit Gaussians. Write the expectation as an integral over x = θ + σε with the Gaussian density φ(x; θ, σ²I) and use ∇θφ = φ · (x − θ)/σ²:

∇J(θ) = (1/σ) · E[ R(θ + σε) · ε ] (1)

Subtracting a constant from R changes nothing, because E[ε] = 0. Pairing ε with −ε costs two attempts that meet the same gusts, so the luck of the draw cancels:

ĝ = (1/N) · Σk [ R(θ + σεk) − R(θ − σεk) ] · εk / (2σ) (2)

Read (2) as an instruction: keep what scored. Each random change pulls θ toward itself by how much better it scored than its mirror image, and away by how much worse. The learner divides each difference by the spread S of the 2N scores in the batch, so one step size serves the plain score (0 or 1) and the plus-progress score (0 to 2):

θ ← θ + η σ · (1/N) · Σk zk εk, zk = (Rk+ − Rk−) / (2S) (3)

with N = 10 mirrored pairs, 20 attempts, per update and step size η = 2. With the noise in the actions instead of the weights, the same identity is the policy gradient (Reinforcement Learning, lessons 6 and 10); in the weights it is an evolution strategy.

Two checks, because (2) carries the lesson. For a quadratic score it averages to the exact gradient: relative error 0.7 % over 20,000 batches. For a step score, R = 1 when θ·u > 0, whose derivative is zero except at the jump, the smoothed score has slope φ(θ·u/σ)/σ along u (φ the unit Gaussian density), 1.76 at θ·u = σ/2 and σ = 0.2; the estimate measures 1.763 along u and 0.008 across it. Sampling finds a slope where the score has none.

How noisy is one update? If the score is locally linear, R(θ + x) = R(θ) + g·x, each pair contributes (g·ε)ε: the mean of (2) is g and the squared length of its error is about d|g|²/N, so

cos(ĝ, g) ≈ 1 / √(1 + d/N) (4)

Each attempt is one number and the gradient has d unknowns. At N = 10 and d = 494 the cosine is 0.141 (measured 0.138); counting only the 202 numbers the clone's runs read, 0.217 (measured 0.216). A direction within 45° of the gradient needs N = d pairs: for those 202 numbers, 404 attempts, 3.4 robot-hours, for one update. The twenty demonstrations that made the clone are 7,368 numbers; twenty attempts return twenty. So an update is mostly noise, as in stochastic gradient descent: the noise averages out and the small signal adds up, provided there is a signal.

3 · A score of zero teaches nothing

If every score in a batch is the same then S = 0 and every term of (2) vanishes: a policy that never scores never improves. Run it: the scratch learner (θ = 0), the plain score, five posts, five seeds of 3,000 attempts, at noise 0.1, 0.2 and 0.3 rad/s. At each level that is 15,000 attempts, 125 robot-hours, and not one scores; every number of θ is still exactly 0 at the end.

This is the needle. On one post a random table scores 31 times in 60,000 draws at 0.2 rad/s, one attempt in 1,935, and one is all the update needs to start. The five seeds find theirs after 180 to 1,440 attempts and reach 90 % after a median of 1,500, 12.5 hours (range 300 to 1,740). Learning from a score alone works, slowly, on a small task. At 0.1 rad/s even one post never scores, and on two posts no seed scores in 3,000 attempts at 0.2.

optionwhat it costswhat goes wrong
differentiate through the worlda model of the plant and of contactthe score is a step; its slope is zero except at the jump (§2)
search at random on the plain score, from nothingattempts only0 of 6,000 draws score on five posts; the update is exactly zero
add a hand-made progress terma person who knows which way progress liesworks: median 1,020 attempts (8.5 h), range 540 to 1,860, and in 1 seed of five nothing ever scores
Road not taken · shape the reward
The first reflex is to give the score a slope: reward the share of the course covered ("plus progress" in the widget). On the Bench it works, and it is the cheapest start for a learner with no prior. It is not free. It says which way progress lies, here to the right, which a person knows for carrying a cup and not for folding a shirt or seating a connector, whose scores are still 0 or 1. It can trap a learner, and it needs 4.25 times the attempts of the start in §4.

4 · Start where scores already differ

The needle is a problem of the start, not of the update: (3) works as soon as scores differ. So begin with a policy that already scores sometimes. We have one, the clone of lesson 1: twenty calm demonstrations (3,684 frames) and nearest-demo regression, stored as a lookup table so that it costs one read a step. On the 100 runs every learning curve uses it completes 62 %; on 400 fresh runs 58 % (95 % interval 53 to 63; lesson 1 measured 63.5 % on 200 others), with or without the edge. More demonstrations of the same kind do not lift it (lesson 1); labels on the policy's own frames would (lesson 3), and the entry has nobody to give them. Of the first 20 attempts about 12 score, so the first batch has a spread and (3) has something to follow.

Three starts share the learner and a table of d numbers: scratch, u = table(q; θ) with θ = 0; clone, free, u = clone(q) + table(q; θ); and clone, bounded, u = clone(q) + clip(table(q; θ), ±0.1 rad/s). In the last the clone is frozen, only the correction learns, and it cannot exceed 0.1 rad/s a joint, under half the clone's own 0.237 rad/s rms. This is the residual formulation of Johannink et al. (2019), who add a learned correction to a hand-designed controller on a real arm, and of Silver et al. (2018), who write πθ(s) = π(s) + fθ(s) with π fixed and gradients only through fθ.

Why bound it? Exploration noise is added to the correction, and noise of σ on every entry is noise of the order of σ on the action: 0.42 of the clone's rms action at 0.1 rad/s, 0.84 at 0.2, 1.27 at 0.3. Noise sized for a policy with nothing to protect drowns one that has something. At 0.2 the free correction is driven to a median lowest success of 0 % and only 3 of five seeds are back above 90 % after 1,500 attempts; the bounded one dips to 7 % and all 5 recover. At 0.3 only 1 free seed is above 90 % and 2 are exactly where they started, 62 %: every perturbed attempt fails, every batch scores zero, and §3 applies to a policy that began at 62 %. The bounded correction keeps all 5.

The widget

Learn from a score: three starts, five seeds
Above: the table from above, its edge dashed, and 12 runs of the first seed's policy after the chosen attempts (cyan reached the mat, red touched a post or left the table, grey ran out of time). Below: success against attempts for five seeds (grey), their median (teal), range (band), 90 % (dashed) and the chosen attempts (amber). Start: nothing, the clone with a free correction, or with the correction bounded at 0.1 rad/s. Score: plain (the cup must also stay on the table), plus progress, or edge ignored (the score of lessons 1 to 11).
completes the course (median)
—
range over 5 seeds
—
attempts that scored
—
robot time spent
—
attempts to 90 % (median)
—
robot time to 90 %
—
numbers in the table
—
clone alone completes
—
Show the core JS
RL.Learner.prototype.step = function () {
  var c = this.cfg, N = c.N, sig = c.sigma, d = this.d, th = this.th, tp = this.tp, tm = this.tm, R = this.R, i, k;
  for (k = 0; k < N; k++) {
    var e = this.eps[k], sd = 100000 + c.seed * 1000003 + this.att + 2 * k;
    for (i = 0; i < d; i++) { e[i] = RL.randn(this.rng); tp[i] = th[i] + sig * e[i]; tm[i] = th[i] - sig * e[i]; }
    var rp = RL.attempt(c.C, this.actP, sd, c.edge), rm = RL.attempt(c.C, this.actM, sd, c.edge);
    R[2 * k] = RL.score(rp, c.reward); R[2 * k + 1] = RL.score(rm, c.reward); this.wins += (rp.done ? 1 : 0) + (rm.done ? 1 : 0);
  }
  this.att += 2 * N; if (this.first < 0 && this.wins > 0) this.first = this.att;
  var m = 0, v = 0; for (i = 0; i < 2 * N; i++) m += R[i]; m /= 2 * N; for (i = 0; i < 2 * N; i++) v += (R[i] - m) * (R[i] - m);
  var S = sqrt(v / (2 * N)); if (S < 1e-9) return;                       // every attempt scored the same: nothing to keep
  for (k = 0; k < N; k++) { var z = (R[2 * k] - R[2 * k + 1]) / (2 * S), a = c.eta * sig * z / N, ek = this.eps[k]; for (i = 0; i < d; i++) th[i] += a * ek[i]; }
};
...
    if (start === 'residual') { v[0] = max(-rho, min(rho, v[0])); v[1] = max(-rho, min(rho, v[1])); }
    prior.tb.read(prior.th, q, p); u[0] = p[0] + v[0]; u[1] = p[1] + v[1];

What to try. Leave the defaults (scratch, plain score, five posts, 494 numbers, noise 0.1, 3,000 attempts): the policy completes the course in 0 % of runs, 0 of the 3,000 attempts have scored, 25 hours are spent, and dragging attempts back changes nothing. Start: clone, bounded correction. It completes 62 % before any attempt, the median seed passes 90 % after 240 attempts (2.0 h; range 60 to 360) and every seed ends at 100 %; the free correction needs 300. Scratch with the score plain + progress: median 1,020, range 540 to 1,860, one seed never scores. With the clone, raise the noise to 0.3: the bounded correction keeps all five seeds above 90 % by 1,500 attempts, the free one 1. Shrink the table to 40 numbers: progress needs 240 attempts, the bounded correction 1,860. One post, scratch, plain, noise 0.2: median 1,500. Five posts, scratch, edge ignored, noise 0.2: 90 % after 180 attempts, because a detour is enough (§1). Task table moved 8 cm, bounded clone: the clone alone completes 0 %, the median is 240.

5 · The bill

Five posts, noise 0.1 rad/s, 494 numbers, five seeds each:

start and scoreattempts to 90 % (median, range)robot-hours (median)
scratch, plainnot within 3,000: nothing ever scoredmore than 25
scratch, plus progress1,020 (540 to 1,860; one seed never)8.5
clone, free correction300 (120 to 720)2.5
clone, bounded correction240 (60 to 360)2.0

The free clone needs 3.4 times fewer attempts than the progress term, and nobody says which way progress lies; the bound is a further 1.25 times fewer at 0.1 rad/s and saves the run at 0.2 and 0.3. The bounded learner needs 2.0 robot-hours against the 4 of an afternoon. Other settings keep the order: with noise 0.05 or 0.2, step size 1 or 4, or 5 or 15 pairs per update, the progress term's median runs from 840 to 1,560 attempts and the bounded clone's from 240 to 480, and the smallest ratio of the two is 1.9.

The table's size moves the two learners in opposite directions:

numbers in the table4070234494
scratch, plus progress2403607801,020
clone, bounded correction1,860 (2 seeds never)480300240

Scratch pays for every number it has to find, as (4) says. The bounded correction pays less on the larger table and fails on the smallest, probably because nodes 0.5 rad apart are far apart at the cup beside posts 0.2 m apart, so a fix for one post spills onto the next. The two cross between 70 and 234 numbers. Forty numbers hold a smooth detour that scores, a policy for this one course and nothing else; a robot policy has millions of numbers because it serves many tasks, and sits at the right of this table and beyond.

Real systems pay in the same currency. Levine et al. (2016) taught grasping with over 800,000 attempts on between 6 and 14 arms: 6,667 robot-hours at this lesson's 30 s. QT-Opt (Kalashnikov et al., 2018) learned from a binary lift reward on over 580,000 real grasps from seven robots, about 800 robot hours, 5.0 s a grasp if the hours cover those grasps; at 30 s each, 4,833 hours. HIL-SERL (Luo et al., 2024) starts from demonstrations and human corrections and reports near-perfect success within 1 to 2.5 hours of training per task, 6 for the slowest; the same learner with no demonstrations and no corrections scored 0 on every task, the pattern of §3 on real arms. Johannink et al. (2019) lifted a hand-designed controller from 2 successes in 20 to 15 in 20 with 8,000 samples, about three hours.

6 · Where a correction stops

The clone was recorded on one table. Move the table 2 cm and it completes 41 % of 200 runs; 3 cm, 14 %; 4 cm, 0 %: its path clears each post by 5 cm, and a 4 cm shift leaves 1 cm at every second one. Yet the bounded correction repairs each moved table. Noise on the correction shifts the path by a few centimetres, so some perturbed attempts score although the clone alone never does, the first after a median of 20 attempts at 4 cm and 60 at 8 cm, and §3 never applies. The repair costs a median 420 attempts at 4 cm (180 to 660, 3.5 h) and 240 at 8 cm (240 to 420), against 240 on the table the clone knows: 2.5 robot-hours a table on average over the three. "Nearly right" means a score is within reach of the noise, not that the policy already scores, and it depends on the score as much as on the policy (§1).

The correction does not travel. The one learned on the table moved 8 cm completes 99.9 % of runs there and 0.0 % on the table the clone was recorded on; the one learned on the original table completes 99.9 % there and 4.0 % on the moved one. A correction fixes one task. Ten tables are 25 robot-hours, and the eleventh starts from the same clone.

What this lesson did not do
It used one learner, a mirrored-pair evolution strategy on a table. Value functions, replay of old attempts and clipped steps lower the attempts per unit of improvement (Reinforcement Learning, lessons 6, 10 and 31); (4) is why they exist. It never touched a real arm: resets, safety limits, a person who watches and a score that must itself be measured are costs inside the 30 s that it only prices (lesson 3 puts a person's corrections in the loop; lesson 15 prices a verdict). It did not imagine attempts with a model, as DayDreamer (Wu et al., 2022) does. The Bench task is small and the prior's error was varied in one way, by moving the table. The score is narrow because we made it so (§1): on the broad score of lessons 1 to 11 a learner from nothing starts, given enough noise (§1), so the comparison of starts is one for narrow scores. What the starting policy is made of is lesson 13.

Common mistakes / failure modes

"a score is enough to learn from"
Not when every attempt scores 0: 15,000 attempts, not one scored, the table still exactly zero (§3).
"reinforcement learning from nothing never works"
It depends on the score. The table's edge makes it narrow; with the edge ignored the median seed reaches 90 % in 180 attempts at noise 0.2 (§1).
"reinforcement learning needs no teacher"
The teacher moved: into a hint (1,020 attempts, one seed trapped) or into a clone (240) (§3, §4).
"explore as widely from the clone as from scratch"
At 0.3 rad/s only 1 of five free seeds passes 90 % in 1,500 attempts; 2 have not left 62 % (§4).
"a learned correction is a skill"
It completes 99.9 % of runs on its own table and 0.0 % on the next (§6).

Checkpoint exercise

Try it
A task has no demonstrations. A policy that starts from nothing completes it in 1 attempt in 1,000. (a) What is the chance that a first batch of 20 attempts contains no success? (b) How many attempts until the first success, on average, and how long at 30 s each? (c) A clone completes 60 % of attempts: what is the chance that a batch of 20 contains no success? (d) What does (3) do with a batch in which all 20 score 0? Answer: (a) 0.99920 = 0.980. (b) 1,000 attempts, 8.3 hours. (c) 0.420, one chance in 90,949,470; about 12 of the first 20 attempts score. (d) S = 0 and every term vanishes, so the update is exactly zero. The first batch from nothing is very probably wasted, by (a), and the first batch from a clone is not, by (c).

Where this points next

A score alone can improve a policy, slowly on a small task (one post: a median of 1,500 attempts, 12.5 hours), and the cost is set by where the policy starts: from nothing with the plain score on five posts, never in 15,000 attempts; with a hand-made hint 1,020; from the clone with a bounded correction 240, 2.0 robot-hours, inside an afternoon of four. The correction belongs to its table: it completes 0.0 % of runs on the next, and ten tables are 25 hours. What carries from task to task has to be in the starting policy, which every source so far feeds: demonstrations, other bodies, video, force, simulation. In practice it is one pretrained generalist, as in π*0.6 (Physical Intelligence, 2025), a 4-billion-parameter vision-language model with an 860-million-parameter action expert, improved from its own attempts (300 folding trajectories on four robots per iteration) and, for box assembly, from expert corrections as well. It is one model that must read language and images, feel force, emit chunks, and still answer within one control step. How is such a model built, and what does it cost to run it at the speed of the robot?

Takeaway
A score can improve a policy only by comparison: try small random changes, weight each by how much better it scored than its mirror image, and step toward them (2, 3). That gives one number per attempt for every number the policy reads, so at ten pairs the update has a cosine of about 0.22 with the gradient over the 202 numbers a clone's runs read (0.14 over all 494), and it gets nothing when every attempt scores the same. From nothing, with the narrow plain score of §1, 15,000 attempts scored zero times; a hand-made progress term needs 1,020 and traps one seed in five. Starting from the lesson-1 clone and searching a correction bounded at 0.1 rad/s needs 240 attempts, 2.0 robot-hours at 30 s each, and survives noise that wrecks the free version. The correction is a fix for one table (0.0 % on another), so the afternoon is per task and the starting policy must carry what tasks share.

Interview prompts

Companion reads: Reinforcement Learning · 10 Policy gradient (the same identity with the noise in the actions), Reinforcement Learning · 31 PPO (steps that cannot leave the policy they started from), Reinforcement Learning · 75 Robotic grasping (the task of the arm-farm systems of §5, as an MDP) and World Models · 08 Imagination (attempts imagined with a model instead of made on the robot).