Learn from a score alone: reinforcement learning
The last lesson made a simulator worth using, and both its repairs assumed someone to copy. For most tasks nobody can demonstrate; there is only a way to tell whether the task was done. This lesson asks what that score can teach and what it costs at 30 s an attempt. The one update a score permits is to try the policy with small random changes and keep what scored. On the Bench, from nothing and with the plain score, which counts falling off the table, it never starts: in 15,000 attempts not one scored. A hand-made progress term rescues it after 1,020 attempts; the lesson-1 clone with a small, bounded correction needs 240. What a correction cannot do is leave the one table it was learned on.
New idea: improve a policy from scored attempts alone by trying small random changes and keeping what scored, and make the first attempts informative by starting from a policy that already succeeds and searching only a small, bounded correction to it. On the Bench that costs 240 attempts, 2.0 robot-hours, where the plain score alone never starts, and the correction it finds belongs to its task.
Forces next: Learning from a score alone works, slowly, on a small task, and works in an afternoon of robot time only when it starts from a policy that is already nearly right and learns a small correction. Every source so far feeds that starting point: demonstrations, other bodies, video, force, simulation. It is one model that must read language and images, feel force, emit chunks, and still answer within one control step. How is such a model built, and what does it cost to run it at the speed of the robot?
1 · What counts as done, what an attempt costs, what a policy is
A demonstration told the robot what to do in every frame. A score tells it one thing about a whole attempt: 1 if the cup reaches the mat, touches no post and stays on the table, 0 otherwise. Nobody says how. An attempt takes 30 s of robot time, the reset included: 120 an hour, 480 in a four-hour afternoon, and the 3,000 the widget goes to are 25 hours.
The policy is the plainest function whose size we can choose: a table of joint velocities on a grid of joint angles, read by bilinear interpolation. A spacing of 0.1 rad has 19 × 13 nodes of two numbers each, so d = 494 numbers θ (rad/s), the parameters of the policy; spacings of 0.15, 0.3 and 0.5 rad give 234, 70 and 40. Only the entries a run reads can change its score: the clone's own runs read 101 of the 247 nodes, 202 of the numbers. Course, gust and start jitter are those of lesson 1. A point on a learning curve is the share of 100 fixed runs the current policy completes; those runs are the measurement and cost no robot time, the attempts that drive the learning do.
The table has an edge, 15 cm either side of the row of posts, and it has to count. The expert's path is 10 cm off the row, so its margin to the edge is 5 cm, as to a post, and no policy so far came near it. A learner that explores comes near everything, and a score takes whatever it accepts. Draw the table at random, every entry a Gaussian of standard deviation σ, and run it on five posts. Ignore the edge and 1.0 % of draws at σ = 0.2 rad/s and 7.2 % at 0.4 reach the mat: any detour over the posts scores. Count the edge and none of 6,000 draws does, at 0.1, 0.2 or 0.4. The edge turns a plain score into a narrow target, and most scores that matter are narrow. It is this lesson's addition, and what follows depends on it: ignore it, as lessons 1 to 11 did, and a learner from nothing does learn a detour: at noise 0.2 the median seed completes 90 % of the runs after 180 attempts, 1.5 hours (the widget's "edge ignored"). What decides the cost is whether small random changes to the starting policy sometimes score, and the edge is what makes that false for a policy that does nothing.
2 · Keep what scored
The robot cannot differentiate its world: the plant is a black box and the score is a step, with a jump wherever the cup grazes a post. An update has to be built from comparisons. Smooth the score R(θ) of one attempt (the plus-progress score adds the share of the course covered) by trying the policy with a random change of size σ rad/s, J(θ) = E[R(θ + σε)], ε a vector of d independent unit Gaussians. Write the expectation as an integral over x = θ + σε with the Gaussian density φ(x; θ, σ²I) and use ∇θφ = φ · (x − θ)/σ²:
∇J(θ) = (1/σ) · E[ R(θ + σε) · ε ] (1)
Subtracting a constant from R changes nothing, because E[ε] = 0. Pairing ε with −ε costs two attempts that meet the same gusts, so the luck of the draw cancels:
ĝ = (1/N) · Σk [ R(θ + σεk) − R(θ − σεk) ] · εk / (2σ) (2)
Read (2) as an instruction: keep what scored. Each random change pulls θ toward itself by how much better it scored than its mirror image, and away by how much worse. The learner divides each difference by the spread S of the 2N scores in the batch, so one step size serves the plain score (0 or 1) and the plus-progress score (0 to 2):
θ ← θ + η σ · (1/N) · Σk zk εk, zk = (Rk+ − Rk−) / (2S) (3)
with N = 10 mirrored pairs, 20 attempts, per update and step size η = 2. With the noise in the actions instead of the weights, the same identity is the policy gradient (Reinforcement Learning, lessons 6 and 10); in the weights it is an evolution strategy.
Two checks, because (2) carries the lesson. For a quadratic score it averages to the exact gradient: relative error 0.7 % over 20,000 batches. For a step score, R = 1 when θ·u > 0, whose derivative is zero except at the jump, the smoothed score has slope φ(θ·u/σ)/σ along u (φ the unit Gaussian density), 1.76 at θ·u = σ/2 and σ = 0.2; the estimate measures 1.763 along u and 0.008 across it. Sampling finds a slope where the score has none.
How noisy is one update? If the score is locally linear, R(θ + x) = R(θ) + g·x, each pair contributes (g·ε)ε: the mean of (2) is g and the squared length of its error is about d|g|²/N, so
cos(ĝ, g) ≈ 1 / √(1 + d/N) (4)
Each attempt is one number and the gradient has d unknowns. At N = 10 and d = 494 the cosine is 0.141 (measured 0.138); counting only the 202 numbers the clone's runs read, 0.217 (measured 0.216). A direction within 45° of the gradient needs N = d pairs: for those 202 numbers, 404 attempts, 3.4 robot-hours, for one update. The twenty demonstrations that made the clone are 7,368 numbers; twenty attempts return twenty. So an update is mostly noise, as in stochastic gradient descent: the noise averages out and the small signal adds up, provided there is a signal.
3 · A score of zero teaches nothing
If every score in a batch is the same then S = 0 and every term of (2) vanishes: a policy that never scores never improves. Run it: the scratch learner (θ = 0), the plain score, five posts, five seeds of 3,000 attempts, at noise 0.1, 0.2 and 0.3 rad/s. At each level that is 15,000 attempts, 125 robot-hours, and not one scores; every number of θ is still exactly 0 at the end.
This is the needle. On one post a random table scores 31 times in 60,000 draws at 0.2 rad/s, one attempt in 1,935, and one is all the update needs to start. The five seeds find theirs after 180 to 1,440 attempts and reach 90 % after a median of 1,500, 12.5 hours (range 300 to 1,740). Learning from a score alone works, slowly, on a small task. At 0.1 rad/s even one post never scores, and on two posts no seed scores in 3,000 attempts at 0.2.
| option | what it costs | what goes wrong |
|---|---|---|
| differentiate through the world | a model of the plant and of contact | the score is a step; its slope is zero except at the jump (§2) |
| search at random on the plain score, from nothing | attempts only | 0 of 6,000 draws score on five posts; the update is exactly zero |
| add a hand-made progress term | a person who knows which way progress lies | works: median 1,020 attempts (8.5 h), range 540 to 1,860, and in 1 seed of five nothing ever scores |
4 · Start where scores already differ
The needle is a problem of the start, not of the update: (3) works as soon as scores differ. So begin with a policy that already scores sometimes. We have one, the clone of lesson 1: twenty calm demonstrations (3,684 frames) and nearest-demo regression, stored as a lookup table so that it costs one read a step. On the 100 runs every learning curve uses it completes 62 %; on 400 fresh runs 58 % (95 % interval 53 to 63; lesson 1 measured 63.5 % on 200 others), with or without the edge. More demonstrations of the same kind do not lift it (lesson 1); labels on the policy's own frames would (lesson 3), and the entry has nobody to give them. Of the first 20 attempts about 12 score, so the first batch has a spread and (3) has something to follow.
Three starts share the learner and a table of d numbers: scratch, u = table(q; θ) with θ = 0; clone, free, u = clone(q) + table(q; θ); and clone, bounded, u = clone(q) + clip(table(q; θ), ±0.1 rad/s). In the last the clone is frozen, only the correction learns, and it cannot exceed 0.1 rad/s a joint, under half the clone's own 0.237 rad/s rms. This is the residual formulation of Johannink et al. (2019), who add a learned correction to a hand-designed controller on a real arm, and of Silver et al. (2018), who write πθ(s) = π(s) + fθ(s) with π fixed and gradients only through fθ.
Why bound it? Exploration noise is added to the correction, and noise of σ on every entry is noise of the order of σ on the action: 0.42 of the clone's rms action at 0.1 rad/s, 0.84 at 0.2, 1.27 at 0.3. Noise sized for a policy with nothing to protect drowns one that has something. At 0.2 the free correction is driven to a median lowest success of 0 % and only 3 of five seeds are back above 90 % after 1,500 attempts; the bounded one dips to 7 % and all 5 recover. At 0.3 only 1 free seed is above 90 % and 2 are exactly where they started, 62 %: every perturbed attempt fails, every batch scores zero, and §3 applies to a policy that began at 62 %. The bounded correction keeps all 5.
The widget
What to try. Leave the defaults (scratch, plain score, five posts, 494 numbers, noise 0.1, 3,000 attempts): the policy completes the course in 0 % of runs, 0 of the 3,000 attempts have scored, 25 hours are spent, and dragging attempts back changes nothing. Start: clone, bounded correction. It completes 62 % before any attempt, the median seed passes 90 % after 240 attempts (2.0 h; range 60 to 360) and every seed ends at 100 %; the free correction needs 300. Scratch with the score plain + progress: median 1,020, range 540 to 1,860, one seed never scores. With the clone, raise the noise to 0.3: the bounded correction keeps all five seeds above 90 % by 1,500 attempts, the free one 1. Shrink the table to 40 numbers: progress needs 240 attempts, the bounded correction 1,860. One post, scratch, plain, noise 0.2: median 1,500. Five posts, scratch, edge ignored, noise 0.2: 90 % after 180 attempts, because a detour is enough (§1). Task table moved 8 cm, bounded clone: the clone alone completes 0 %, the median is 240.
5 · The bill
Five posts, noise 0.1 rad/s, 494 numbers, five seeds each:
| start and score | attempts to 90 % (median, range) | robot-hours (median) |
|---|---|---|
| scratch, plain | not within 3,000: nothing ever scored | more than 25 |
| scratch, plus progress | 1,020 (540 to 1,860; one seed never) | 8.5 |
| clone, free correction | 300 (120 to 720) | 2.5 |
| clone, bounded correction | 240 (60 to 360) | 2.0 |
The free clone needs 3.4 times fewer attempts than the progress term, and nobody says which way progress lies; the bound is a further 1.25 times fewer at 0.1 rad/s and saves the run at 0.2 and 0.3. The bounded learner needs 2.0 robot-hours against the 4 of an afternoon. Other settings keep the order: with noise 0.05 or 0.2, step size 1 or 4, or 5 or 15 pairs per update, the progress term's median runs from 840 to 1,560 attempts and the bounded clone's from 240 to 480, and the smallest ratio of the two is 1.9.
The table's size moves the two learners in opposite directions:
| numbers in the table | 40 | 70 | 234 | 494 |
|---|---|---|---|---|
| scratch, plus progress | 240 | 360 | 780 | 1,020 |
| clone, bounded correction | 1,860 (2 seeds never) | 480 | 300 | 240 |
Scratch pays for every number it has to find, as (4) says. The bounded correction pays less on the larger table and fails on the smallest, probably because nodes 0.5 rad apart are far apart at the cup beside posts 0.2 m apart, so a fix for one post spills onto the next. The two cross between 70 and 234 numbers. Forty numbers hold a smooth detour that scores, a policy for this one course and nothing else; a robot policy has millions of numbers because it serves many tasks, and sits at the right of this table and beyond.
Real systems pay in the same currency. Levine et al. (2016) taught grasping with over 800,000 attempts on between 6 and 14 arms: 6,667 robot-hours at this lesson's 30 s. QT-Opt (Kalashnikov et al., 2018) learned from a binary lift reward on over 580,000 real grasps from seven robots, about 800 robot hours, 5.0 s a grasp if the hours cover those grasps; at 30 s each, 4,833 hours. HIL-SERL (Luo et al., 2024) starts from demonstrations and human corrections and reports near-perfect success within 1 to 2.5 hours of training per task, 6 for the slowest; the same learner with no demonstrations and no corrections scored 0 on every task, the pattern of §3 on real arms. Johannink et al. (2019) lifted a hand-designed controller from 2 successes in 20 to 15 in 20 with 8,000 samples, about three hours.
6 · Where a correction stops
The clone was recorded on one table. Move the table 2 cm and it completes 41 % of 200 runs; 3 cm, 14 %; 4 cm, 0 %: its path clears each post by 5 cm, and a 4 cm shift leaves 1 cm at every second one. Yet the bounded correction repairs each moved table. Noise on the correction shifts the path by a few centimetres, so some perturbed attempts score although the clone alone never does, the first after a median of 20 attempts at 4 cm and 60 at 8 cm, and §3 never applies. The repair costs a median 420 attempts at 4 cm (180 to 660, 3.5 h) and 240 at 8 cm (240 to 420), against 240 on the table the clone knows: 2.5 robot-hours a table on average over the three. "Nearly right" means a score is within reach of the noise, not that the policy already scores, and it depends on the score as much as on the policy (§1).
The correction does not travel. The one learned on the table moved 8 cm completes 99.9 % of runs there and 0.0 % on the table the clone was recorded on; the one learned on the original table completes 99.9 % there and 4.0 % on the moved one. A correction fixes one task. Ten tables are 25 robot-hours, and the eleventh starts from the same clone.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
A score alone can improve a policy, slowly on a small task (one post: a median of 1,500 attempts, 12.5 hours), and the cost is set by where the policy starts: from nothing with the plain score on five posts, never in 15,000 attempts; with a hand-made hint 1,020; from the clone with a bounded correction 240, 2.0 robot-hours, inside an afternoon of four. The correction belongs to its table: it completes 0.0 % of runs on the next, and ten tables are 25 hours. What carries from task to task has to be in the starting policy, which every source so far feeds: demonstrations, other bodies, video, force, simulation. In practice it is one pretrained generalist, as in π*0.6 (Physical Intelligence, 2025), a 4-billion-parameter vision-language model with an 860-million-parameter action expert, improved from its own attempts (300 folding trajectories on four robots per iteration) and, for box assembly, from expert corrections as well. It is one model that must read language and images, feel force, emit chunks, and still answer within one control step. How is such a model built, and what does it cost to run it at the speed of the robot?
Interview prompts
- Why can a reward not be differentiated through, and what is computed instead? (§2 — the score is a step; the estimator weights random parameter changes by the score difference of each mirrored pair and recovers the smoothed gradient.)
- Why does reinforcement learning from scratch with a sparse reward fail, and what are the ways out? (§3, §4 — every batch scores zero, so the update is zero; shape the reward, which costs knowledge, or start from a policy that scores, which costs a prior.)
- What is a residual policy and why bound the correction? (§4 — a frozen base plus a learned correction; noise sized for scratch drowns a good base.)
- How does the cost of the update grow with the number of parameters? (§2, §5 — the cosine with the gradient is 1/√(1 + d/N); from scratch, 240 attempts at 40 numbers, 1,020 at 494.)
- A correction learned on one layout is run on another. What happens, and where must generality live? (§6 — it completes about 0 % of runs; generality has to be in the starting policy.)
- At 30 s an attempt, what are 580,000 attempts? (§1, §5 — 4,833 robot-hours, about 201 days; QT-Opt's own 800 robot hours imply about 5 s a grasp.)
Companion reads: Reinforcement Learning · 10 Policy gradient (the same identity with the noise in the actions), Reinforcement Learning · 31 PPO (steps that cannot leave the policy they started from), Reinforcement Learning · 75 Robotic grasping (the task of the arm-farm systems of §5, as an MDP) and World Models · 08 Imagination (attempts imagined with a model instead of made on the robot).