Unlimited data, wrong physics: simulate, randomise, calibrate
Lesson 10 ended with a policy that needs force in its data, which only a real robot gives, slowly and with wear. A simulator makes actions and demonstrations without limit, from a physics that is not quite the robot's. This lesson measures how wrong that can be, on a plant with a gain, a lag and a delay. A clone trained in a simulator with no lag reaches the mat in 94 % of its runs there, in 43.5 % on a robot whose lag is 0.1 s and in 0 % at 0.2 s, where its expert still passes every run. Two repairs follow. Draw a different plant in every simulated run, which puts the robot's states into the clone's table at a price that grows with the width; and measure the robot for a few seconds, which shrinks the width needed. Both assume an expert inside the simulator.
New idea: a policy trained in a simulator survives the plants the simulator drew, so draw the plants the robot may be, as widely as the measurement leaves them unknown and no wider. The width is a price, and how well the robot has been measured sets it.
Forces next: A simulator is worth using when its error is narrower than the policy can tolerate: randomising the physics widens what a policy survives at the price of a more cautious motion, and calibrating the simulator on a few real rollouts narrows the gap itself. Both assume there is someone in the simulator to copy, a scripted expert or a planner. For most tasks that matter there is no such expert, only a way to tell whether the task was done. What can be learned from a score alone, and how many attempts does it take?
1 · The plant, written down
Lesson 1 gave the arm a plant that does what it is told plus a gust. A real arm does what it is told late, and by a different amount. Three numbers describe that well enough here. With the step Δt = 0.05 s, a joint's velocity v follows the command u through a first-order lag, the command may reach the arm d steps late, and the arm may move by another amount than it was told:
v ← v + k (g · ut−d − v), k = Δt / (τ + Δt), q ← q + Δt (v + gust)
| symbol | meaning | unit |
|---|---|---|
| g | gain: motion per unit of command (1 = as commanded) | |
| τ | lag: time constant of the velocity's response to the command | s |
| d | delay: steps between the command and its effect | steps |
The simulator so far, the new arm of lesson 1, has g = 1, τ = 0, d = 0. A lag of 0.2 s is 4 steps: after four steps the velocity has reached 59 % of a constant command, 1 − (τ / (τ + Δt))n after n steps. Data from such a simulator are cheap: Rudin et al. (2022) trained a walking policy on 4096 simulated robots at once, about 800 robot-hours of experience (1,500 updates of 98,304 steps at 50 Hz) in under 20 minutes on one workstation GPU. The expert is a feedback law of the joint angles, so a late or sluggish arm changes where it ends up and the expert steers back from there. How far that holds is the next measurement.
2 · The gap, as lost success
Train a clone in the plain simulator as lesson 3 would: 20 demonstrations, then four rounds in which the clone drives five runs and the expert relabels every frame it visited. That is 40 simulated runs and 7079 stored frames, and the clone reaches the mat in 94 % of 200 runs in its own simulator (gust 0.05 rad/s, as always). Now run it, unchanged, on robots that differ from the simulator in one number, and run the expert on the same robots:
| the robot differs from the simulator in | clone reaches the mat | expert reaches the mat |
|---|---|---|
| nothing (the simulator itself) | 94 % | 100 % |
| lag 0.05 s | 82.5 % | 100 % |
| lag 0.1 s | 43.5 % | 100 % |
| lag 0.2 s | 0 % | 100 % |
| lag 0.3 s | 0 % | 99 % |
| delay 2 steps | 44.5 % | 100 % |
| delay 4 steps | 0.5 % | 100 % |
| delay 6 steps | 0 % | 100 % |
| gain 1.5 | 99 % | 100 % |
| gain 0.85 | 86 % | 100 % |
| gain 0.7 | 0 % | 0 % |
The gap is the clone's, not the expert's. The expert passes at least 99 % of its runs up to a lag of 0.3 s and a delay of 6 steps, while the clone has lost more than half its runs at 0.1 s of lag or two steps of delay and all of them at 0.2 s. It opens early. A twentieth of a second of lag, one step, is where the clone starts to lose runs, and every one of its failures at a lag or a delay is a collision (100 % of the runs at 0.2 s). Gain is different. A gain of 1.5 costs nothing, and 0.85 leaves 86 % with every failure a collision; at 0.7 the clone and the expert both fail, and 80 % of the clone's runs by the clock: feedback absorbs a gain error until the arm is too slow for the time allowed. Nor is it a seed accident: over six training seeds the clone's success at 0.1 s of lag runs from 43.5 to 77 %, and at 0.2 s from 0 to 5 %.
The reason is lesson 1's. The clone is a table, and a table answers where it holds frames. In its own simulator every frame the clone reaches has a stored frame within three bandwidths; on the robot with 0.2 s of lag 10.7 % of its steps do not (0 % in the simulator), and the nearest-frame copy that replaces the lookup is a command for somewhere else. The labels were right. The states were the simulator's.
3 · Two repairs that do not work
More simulated data. It is lesson 1's reflex, and a simulator can supply it. With four and eight times the runs, three training seeds each, the mean success on the plain simulator and on two robots is:
| simulated runs (stored frames) | the simulator | lag 0.1 s | lag 0.2 s |
|---|---|---|---|
| 40 (7079) | 95.7 % | 57.2 % | 1.8 % |
| 160 (29343) | 95.8 % | 59.8 % | 3.8 % |
| 320 (58512) | 95.7 % | 64.2 % | 5 % |
Eight times the runs buys 7 points at 0.1 s of lag and 3 at 0.2 s. The new frames are more of the same band: a simulator cannot visit states its own plant does not produce.
A better simulator. Everything that can be written down is already in it. What is left is a number nobody knows, and the only way to learn it is to measure the robot. A measurement has an error, section 6 computes how it falls with the effort, and it never reaches zero. The contact stiffness and friction that lesson 10's simulator got wrong are in the same position.
4 · Draw the plant at random
The clone fails because the robot makes it visit states its table does not hold, and the simulator can put them there. Let every simulated run draw its own lag, uniformly from a range of half-width w around the value the simulator believes, and change nothing else: still 40 runs, the same expert labels. The clone cannot tell which lag a frame came from, and does not need to: the label is the expert's command at that state, which does not depend on the plant. Here the simulator believes there is no lag, so the range is 0 to w s. The robots have a lag of 0 to 0.4 s and gain 1:
| range 0 to w s | robot lag 0 | 0.1 s | 0.2 s | 0.3 s | 0.4 s |
|---|---|---|---|---|---|
| w = 0 | 94 % | 43.5 % | 0 % | 0 % | 0 % |
| w = 0.1 | 94 % | 99 % | 50 % | 0 % | 0 % |
| w = 0.2 | 86.5 % | 99.5 % | 74 % | 5.5 % | 0 % |
| w = 0.3 | 81.5 % | 99 % | 97.5 % | 57 % | 2 % |
| w = 0.45 | 90 % | 99.5 % | 99.5 % | 86 % | 23 % |
Each row is high on the lags it was trained on and a little beyond: widening from 0 to 0.3 lifts the robot with 0.2 s of lag from 0 % to 97.5 %. The edge of a range is weaker than its middle (at w = 0.2 the robot at the edge, 0.2 s, gets 74 %, while the robot at 0.1 s gets 99.5 %). No width rescues the robot with 0.4 s of lag; its best is 23 %. The rows also take something away on the left: the robot with no lag falls from 94 % to 81.5 % as the range grows to 0.3 (one training seed here; section 5 averages three). That is the first price.
5 · What the width costs
1 · Precision where the simulator was right. Take a robot whose lag really is 0 (gain 0.9, which the simulator knows) and widen the range around it, three training seeds each, 200 runs per clone:
| half-width w | success, 40 runs | success, 160 runs | copy error on fresh expert frames, 40 runs |
|---|---|---|---|
| 0 | 93.8 % | 93.8 % | 4.8 % |
| 0.15 | 83.3 % | 87.5 % | 5.3 % |
| 0.3 | 76.5 % | 86.2 % | 6 % |
| 0.6 | 69.8 % | 83 % | 7.9 % |
The widest range costs 24 points, and four times the runs cut that to 10.8. The copy error on fresh expert frames, lesson 1's first score, rises from 4.8 to 7.9 %: the kernel's neighbourhood now holds frames from states farther off the path, whose labels point back toward it, and they bend the commands given on the path. A simulator that was exact pays for the ones that were not.
2 · The expert's own limit. On the clock of lesson 1 the expert passes 99, 90.5, 65 and 3.5 % of its runs at lags of 0.3, 0.35, 0.4 and 0.5 s (gain 1). A range that reaches past that contains runs that fail whatever the labels. For a robot at 0.35 s, with the simulator exactly right, success stays between 56 and 80 % at every width, below the expert's 90.5 %.
3 · Caution. A slower expert reaches further, if the clock lets it. With 1.6 times the time (372 steps instead of 235) the expert passes 8.5 % of its runs at a lag of 0.5 s at 15 cm/s, 62 % at 12 cm/s and 85.5 % at 10 cm/s (at 0.4 s: 64.5, 91 and 97.5 %), for 1.25 and 1.5 times as long per task. On the clock of lesson 1 the slowest pace that still passes 95 % of runs with no lag is 12.25 cm/s, and any lag uses the slack: there is no room to be careful. A range wide enough to contain a plant the expert cannot pass at full speed therefore forces the demonstrations to slow down. That is the more cautious motion that width buys. This lesson keeps the clock and the pace, so its widths stay inside the expert's limit.
The large systems pay the same prices. OpenAI et al. (2018) randomised the physics of a robot hand in simulation and needed about 100 years of simulated experience, against about 3 years without randomisation; on the real hand the median run managed 13 consecutive goals with every randomisation, 2 without the physics ones and 0 with none. Peng et al. (2018) found for a puck-pushing arm that fixing only the action timestep, one of the randomised parameters, cut real success from 0.89 (28 trials) to 0.29 (17 trials).
6 · Measure the robot
Draw less widely by knowing more. Command gentle random steps in free space for 3 s (each joint's command redrawn every 5 steps, uniform in ±0.15 rad/s) and log the commands and the measured joint velocities yt = (qt+1 − qt) / Δt. Under the plant of section 1, yt = g ft(τ, d) + gust, where ft is what a plant with gain 1, lag τ and delay d does to the logged commands. For any candidate (τ, d) the best gain is a least-squares scale and the residual follows:
g = Σ f y / Σ f², residual = Σ (y − g f)²
Search τ from 0 to 0.6 s in 5 ms steps and d from 0 to 6 steps, and keep the smallest residual (the four robots below have no delay, so their fits search τ only). A probe of 3 s is 120 velocity samples (two joints) with gust noise 0.05 rad/s. Samples from separate probes are independent, so the information adds and the error of the fit falls as 1/√n in the number of probes n. Six hundred repeated calibrations of a robot with gain 0.9 and lag 0.25 s give:
| probes n | robot time | scatter of the fitted lag | 90 % of fits within | smallest scatter any unbiased fit can have |
|---|---|---|---|---|
| 1 | 3 s | 0.070 s | 0.105 s | 0.064 s |
| 3 | 9 s | 0.037 s | 0.060 s | 0.036 s |
| 10 | 30 s | 0.019 s | 0.030 s | 0.019 s |
| 30 | 90 s | 0.011 s | 0.015 s | 0.011 s |
The log–log slope of the scatter against n is −0.55, where the law says −0.5 (the fit from a single probe is the least efficient), and the scatter sits within 8 % of the smallest an unbiased fit can have (the bound from the Fisher information of these probes). The fit is also unbiased to a few milliseconds. A delay is found the same way: a robot with a delay of 3 steps and no lag is recovered with exactly 3 steps in 100 of 100 fits from three probes. What the measurement leaves is the tail: one fit in ten misses by more than 0.105 s with one probe, 0.060 s with three and 0.030 s with ten. Set that against what the plain clone tolerates, about 0.05 s of lag (section 2): the tail after one probe is outside it and needs a range, after three it is at the edge, after ten it is inside. That is the sense in which a simulator is worth using when its error is narrower than the policy can tolerate, and the width worth drawing is the part of the tail that is not.
SimOpt (Chebotar et al., 2019) works on the same economy: it matched the distribution of a simulator's parameters to 3 real roll-outs per iteration, and its swing-peg policy passed in 90 % of 20 real trials after two iterations, about six roll-outs. In their simulated study, naive wide randomisation failed once the cabinet's offset reached 10 cm, and SimOpt handled 15 cm in three iterations.
7 · Both together, on four robots
The widget puts them together. Four robots have lags of 0, 0.1, 0.2 and 0.25 s (gain 0.9 for all). For each number of probes the calibration is repeated five times with different probes; the simulator's range is centred on each fit, with the fitted gain; a clone is trained in each (40 simulated runs) and run on the robot, 100 runs per point.
What to try. Leave the defaults: the robot with 0.2 s of lag, three probes, half-width 0.15 s. The mean success of its five clones is 96 %; at width 0 it is 88.2 % and at 0.3 s 97.8 %, a plateau, and the narrowest width within 2 points of the best is 0.15 s. Set the probes to 0: nobody measured, the simulator says no lag, and success is 2 % at width 0 and 95 % at 0.3 s, so the knee is 0.3 s. One probe (knee 0.3 s again) shows the tail: the mean at width 0 is 74.2 %, but the worst of the five calibrations gets 44 %, rising to 91 % at 0.3 s. Ten probes need no randomisation: 96 % at width 0 and a knee of 0 s. Now change the robot. With lag 0 and ten probes, width 0 gives 96.8 % and width 0.3 s gives 81 %: the price of section 5, paid by the robot the simulator had right. At 0.1 s with no probes, success is 45 % at width 0 and 98 % at 0.1 s. At 0.25 s, near the expert's limit, ten probes give 79.6 % at width 0 and 88.4 % at 0.3 s: close to the limit, width buys margin even after the measurement.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
A simulator labels any state for free, and a clone copied from it is only as good as the overlap between the states it learned from and the states the robot makes it visit. Drawing the plant at random widens that overlap at a price (the robot the simulator had right loses up to 24 points, and plants past the expert's limit need a slower, more cautious expert), and measuring the robot narrows what has to be drawn: the lag is known to 0.070 s from one probe and 0.011 s from thirty. A simulator is worth using when what is left is narrower than the policy tolerates. Every number above had someone to copy. In the simulator the expert labelled all 7079 frames of 40 runs and scored them for free. For most tasks that matter there is no such expert, only a way to tell whether the task was done, and that is one bit per attempt. Telling a policy that succeeds 96 % of the time from one that succeeds 88.2 % of the time takes about 188 attempts of each (for a 5 % chance of a false alarm and an 80 % chance of seeing a real difference; lesson 15 derives it), and a 3-point difference, 96 against 93 %, about 907: at 30 s an attempt, 1.6 and 7.6 hours of robot per candidate. What can be learned from a score alone, and how many attempts does it take?
Interview prompts
- A policy trained in simulation scores 95 % there and fails on the robot. Is the simulator's data wrong? (§2 — no: the expert passes 100 % of runs on the robot's plant, but the clone's table has no frames where the robot's plant takes it, and 10.7 % of its steps are outside the table's reach at 0.2 s of lag.)
- Why does a clone fail at a lag the expert it copied handles? (§2 — the expert is a feedback law that steers back from any state; the clone is a lookup that answers only where it holds frames.)
- What does drawing the plant at random do to the training data, and why can a policy that cannot see the plant still use it? (§4 — it adds the states each plant produces; the label at a state is the expert's command, which does not depend on the plant.)
- Name three prices of a wide range. (§5 — precision for the robot the simulator had right, the expert's own limit, and a slower, more cautious expert bounded by the clock; large systems add simulated experience, about 33 times as much for OpenAI et al. (2018).)
- How would you calibrate a simulator's gain, lag and delay, and how many rollouts does it take? (§6 — gentle random steps in free space and a least-squares fit; the scatter falls as 1/√n, 0.070 s for one probe and 0.019 s for ten here, and SimOpt used 3 roll-outs per iteration.)
- How wide should the randomisation be after calibrating? (§6, §7 — about the tail of what the fit leaves, 0.105, 0.060 and 0.030 s at the 90th percentile for 1, 3 and 10 probes; nearer the expert's limit more, since width also buys margin.)
- Why is the simulator route closed for most tasks that matter? (§7, §2 — it needs an expert or a planner to label states, and most tasks have none, only a score of whole attempts.)
Companion reads: Synthetic Vision · 02 Where the gap lives (the same gap for rendered images, priced stage by stage), Reinforcement Learning · 04 Reward and simulation (what a simulator is for in RL) and World Models · 31 Capstone (the track this one follows).