all_lessons/Robot Model Training/11 · Simulationlesson 11 / 24

Unlimited data, wrong physics: simulate, randomise, calibrate

Lesson 10 ended with a policy that needs force in its data, which only a real robot gives, slowly and with wear. A simulator makes actions and demonstrations without limit, from a physics that is not quite the robot's. This lesson measures how wrong that can be, on a plant with a gain, a lag and a delay. A clone trained in a simulator with no lag reaches the mat in 94 % of its runs there, in 43.5 % on a robot whose lag is 0.1 s and in 0 % at 0.2 s, where its expert still passes every run. Two repairs follow. Draw a different plant in every simulated run, which puts the robot's states into the clone's table at a price that grows with the width; and measure the robot for a few seconds, which shrinks the width needed. Both assume an expert inside the simulator.

The thesis, here
A simulator labels any state for free, and the expert's label does not depend on the plant, but the states it labels are the ones the simulator's plant visits. A policy copied from it knows what to do in those states, and a robot with another plant visits others. Drawing the plant at random in each simulated run widens the set of states the policy knows, and costs precision where the simulator was already right. Measuring the plant on a few short rollouts shrinks the width that has to be drawn, with an error that falls as 1/√n. A simulator is worth using when the error left after measuring is narrower than what the policy tolerates.
Linear position
Forced by: A policy that feels force and yields to it, instead of holding a position stiffly, inserts the peg where position control jams, and the force it needs is in the data only if someone recorded it with a force sensor in the loop, on a real robot, slowly, with wear. A simulator produces force, actions and unlimited demonstrations, at the price of a physics that is not quite the real one. How wrong can a simulator be before a policy trained in it fails on the robot, and what buys back the difference?
New idea: a policy trained in a simulator survives the plants the simulator drew, so draw the plants the robot may be, as widely as the measurement leaves them unknown and no wider. The width is a price, and how well the robot has been measured sets it.
Forces next: A simulator is worth using when its error is narrower than the policy can tolerate: randomising the physics widens what a policy survives at the price of a more cautious motion, and calibrating the simulator on a few real rollouts narrows the gap itself. Both assume there is someone in the simulator to copy, a scripted expert or a planner. For most tasks that matter there is no such expert, only a way to tell whether the task was done. What can be learned from a score alone, and how many attempts does it take?
The plan
Seven moves. (1) Write the plant down: gain, lag, delay. (2) Train a clone in the simulator and measure the gap as lost success. (3) Rule out two cheap repairs: more simulated data and a better simulator. (4) Draw the plant at random in every simulated run. (5) Read what the width costs. (6) Measure the robot: least squares on a few seconds of motion. (7) Put both together on four robots, and see what the combination still needs.

1 · The plant, written down

Lesson 1 gave the arm a plant that does what it is told plus a gust. A real arm does what it is told late, and by a different amount. Three numbers describe that well enough here. With the step Δt = 0.05 s, a joint's velocity v follows the command u through a first-order lag, the command may reach the arm d steps late, and the arm may move by another amount than it was told:

v ← v + k (g · ut−d − v), k = Δt / (τ + Δt), q ← q + Δt (v + gust)

symbolmeaningunit
ggain: motion per unit of command (1 = as commanded)
τlag: time constant of the velocity's response to the commands
ddelay: steps between the command and its effectsteps

The simulator so far, the new arm of lesson 1, has g = 1, τ = 0, d = 0. A lag of 0.2 s is 4 steps: after four steps the velocity has reached 59 % of a constant command, 1 − (τ / (τ + Δt))n after n steps. Data from such a simulator are cheap: Rudin et al. (2022) trained a walking policy on 4096 simulated robots at once, about 800 robot-hours of experience (1,500 updates of 98,304 steps at 50 Hz) in under 20 minutes on one workstation GPU. The expert is a feedback law of the joint angles, so a late or sluggish arm changes where it ends up and the expert steers back from there. How far that holds is the next measurement.

2 · The gap, as lost success

Train a clone in the plain simulator as lesson 3 would: 20 demonstrations, then four rounds in which the clone drives five runs and the expert relabels every frame it visited. That is 40 simulated runs and 7079 stored frames, and the clone reaches the mat in 94 % of 200 runs in its own simulator (gust 0.05 rad/s, as always). Now run it, unchanged, on robots that differ from the simulator in one number, and run the expert on the same robots:

the robot differs from the simulator inclone reaches the matexpert reaches the mat
nothing (the simulator itself)94 %100 %
lag 0.05 s82.5 %100 %
lag 0.1 s43.5 %100 %
lag 0.2 s0 %100 %
lag 0.3 s0 %99 %
delay 2 steps44.5 %100 %
delay 4 steps0.5 %100 %
delay 6 steps0 %100 %
gain 1.599 %100 %
gain 0.8586 %100 %
gain 0.70 %0 %

The gap is the clone's, not the expert's. The expert passes at least 99 % of its runs up to a lag of 0.3 s and a delay of 6 steps, while the clone has lost more than half its runs at 0.1 s of lag or two steps of delay and all of them at 0.2 s. It opens early. A twentieth of a second of lag, one step, is where the clone starts to lose runs, and every one of its failures at a lag or a delay is a collision (100 % of the runs at 0.2 s). Gain is different. A gain of 1.5 costs nothing, and 0.85 leaves 86 % with every failure a collision; at 0.7 the clone and the expert both fail, and 80 % of the clone's runs by the clock: feedback absorbs a gain error until the arm is too slow for the time allowed. Nor is it a seed accident: over six training seeds the clone's success at 0.1 s of lag runs from 43.5 to 77 %, and at 0.2 s from 0 to 5 %.

The reason is lesson 1's. The clone is a table, and a table answers where it holds frames. In its own simulator every frame the clone reaches has a stored frame within three bandwidths; on the robot with 0.2 s of lag 10.7 % of its steps do not (0 % in the simulator), and the nearest-frame copy that replaces the lookup is a command for somewhere else. The labels were right. The states were the simulator's.

3 · Two repairs that do not work

More simulated data. It is lesson 1's reflex, and a simulator can supply it. With four and eight times the runs, three training seeds each, the mean success on the plain simulator and on two robots is:

simulated runs (stored frames)the simulatorlag 0.1 slag 0.2 s
40 (7079)95.7 %57.2 %1.8 %
160 (29343)95.8 %59.8 %3.8 %
320 (58512)95.7 %64.2 %5 %

Eight times the runs buys 7 points at 0.1 s of lag and 3 at 0.2 s. The new frames are more of the same band: a simulator cannot visit states its own plant does not produce.

A better simulator. Everything that can be written down is already in it. What is left is a number nobody knows, and the only way to learn it is to measure the robot. A measurement has an error, section 6 computes how it falls with the effort, and it never reaches zero. The contact stiffness and friction that lesson 10's simulator got wrong are in the same position.

4 · Draw the plant at random

The clone fails because the robot makes it visit states its table does not hold, and the simulator can put them there. Let every simulated run draw its own lag, uniformly from a range of half-width w around the value the simulator believes, and change nothing else: still 40 runs, the same expert labels. The clone cannot tell which lag a frame came from, and does not need to: the label is the expert's command at that state, which does not depend on the plant. Here the simulator believes there is no lag, so the range is 0 to w s. The robots have a lag of 0 to 0.4 s and gain 1:

range 0 to w srobot lag 00.1 s0.2 s0.3 s0.4 s
w = 094 %43.5 %0 %0 %0 %
w = 0.194 %99 %50 %0 %0 %
w = 0.286.5 %99.5 %74 %5.5 %0 %
w = 0.381.5 %99 %97.5 %57 %2 %
w = 0.4590 %99.5 %99.5 %86 %23 %

Each row is high on the lags it was trained on and a little beyond: widening from 0 to 0.3 lifts the robot with 0.2 s of lag from 0 % to 97.5 %. The edge of a range is weaker than its middle (at w = 0.2 the robot at the edge, 0.2 s, gets 74 %, while the robot at 0.1 s gets 99.5 %). No width rescues the robot with 0.4 s of lag; its best is 23 %. The rows also take something away on the left: the robot with no lag falls from 94 % to 81.5 % as the range grows to 0.3 (one training seed here; section 5 averages three). That is the first price.

Road not taken · randomise as widely as the simulator allows
It is what the large systems do and it is attractive because it needs no measurement: every robot is inside the range. Here it costs the robot the simulator had right up to 24 points (section 5), four times the runs halve that, and past the expert's own limit it buys nothing. OpenAI et al. (2019) automate the opposite discipline: a parameter's range widens while the policy does well at its edge and narrows when it does badly there, so the width follows what the policy tolerates.

5 · What the width costs

1 · Precision where the simulator was right. Take a robot whose lag really is 0 (gain 0.9, which the simulator knows) and widen the range around it, three training seeds each, 200 runs per clone:

half-width wsuccess, 40 runssuccess, 160 runscopy error on fresh expert frames, 40 runs
093.8 %93.8 %4.8 %
0.1583.3 %87.5 %5.3 %
0.376.5 %86.2 %6 %
0.669.8 %83 %7.9 %

The widest range costs 24 points, and four times the runs cut that to 10.8. The copy error on fresh expert frames, lesson 1's first score, rises from 4.8 to 7.9 %: the kernel's neighbourhood now holds frames from states farther off the path, whose labels point back toward it, and they bend the commands given on the path. A simulator that was exact pays for the ones that were not.

2 · The expert's own limit. On the clock of lesson 1 the expert passes 99, 90.5, 65 and 3.5 % of its runs at lags of 0.3, 0.35, 0.4 and 0.5 s (gain 1). A range that reaches past that contains runs that fail whatever the labels. For a robot at 0.35 s, with the simulator exactly right, success stays between 56 and 80 % at every width, below the expert's 90.5 %.

3 · Caution. A slower expert reaches further, if the clock lets it. With 1.6 times the time (372 steps instead of 235) the expert passes 8.5 % of its runs at a lag of 0.5 s at 15 cm/s, 62 % at 12 cm/s and 85.5 % at 10 cm/s (at 0.4 s: 64.5, 91 and 97.5 %), for 1.25 and 1.5 times as long per task. On the clock of lesson 1 the slowest pace that still passes 95 % of runs with no lag is 12.25 cm/s, and any lag uses the slack: there is no room to be careful. A range wide enough to contain a plant the expert cannot pass at full speed therefore forces the demonstrations to slow down. That is the more cautious motion that width buys. This lesson keeps the clock and the pace, so its widths stay inside the expert's limit.

The large systems pay the same prices. OpenAI et al. (2018) randomised the physics of a robot hand in simulation and needed about 100 years of simulated experience, against about 3 years without randomisation; on the real hand the median run managed 13 consecutive goals with every randomisation, 2 without the physics ones and 0 with none. Peng et al. (2018) found for a puck-pushing arm that fixing only the action timestep, one of the randomised parameters, cut real success from 0.89 (28 trials) to 0.29 (17 trials).

6 · Measure the robot

Draw less widely by knowing more. Command gentle random steps in free space for 3 s (each joint's command redrawn every 5 steps, uniform in ±0.15 rad/s) and log the commands and the measured joint velocities yt = (qt+1 − qt) / Δt. Under the plant of section 1, yt = g ft(τ, d) + gust, where ft is what a plant with gain 1, lag τ and delay d does to the logged commands. For any candidate (τ, d) the best gain is a least-squares scale and the residual follows:

g = Σ f y / Σ f², residual = Σ (y − g f)²

Search τ from 0 to 0.6 s in 5 ms steps and d from 0 to 6 steps, and keep the smallest residual (the four robots below have no delay, so their fits search τ only). A probe of 3 s is 120 velocity samples (two joints) with gust noise 0.05 rad/s. Samples from separate probes are independent, so the information adds and the error of the fit falls as 1/√n in the number of probes n. Six hundred repeated calibrations of a robot with gain 0.9 and lag 0.25 s give:

probes nrobot timescatter of the fitted lag90 % of fits withinsmallest scatter any unbiased fit can have
13 s0.070 s0.105 s0.064 s
39 s0.037 s0.060 s0.036 s
1030 s0.019 s0.030 s0.019 s
3090 s0.011 s0.015 s0.011 s

The log–log slope of the scatter against n is −0.55, where the law says −0.5 (the fit from a single probe is the least efficient), and the scatter sits within 8 % of the smallest an unbiased fit can have (the bound from the Fisher information of these probes). The fit is also unbiased to a few milliseconds. A delay is found the same way: a robot with a delay of 3 steps and no lag is recovered with exactly 3 steps in 100 of 100 fits from three probes. What the measurement leaves is the tail: one fit in ten misses by more than 0.105 s with one probe, 0.060 s with three and 0.030 s with ten. Set that against what the plain clone tolerates, about 0.05 s of lag (section 2): the tail after one probe is outside it and needs a range, after three it is at the edge, after ten it is inside. That is the sense in which a simulator is worth using when its error is narrower than the policy can tolerate, and the width worth drawing is the part of the tail that is not.

SimOpt (Chebotar et al., 2019) works on the same economy: it matched the distribution of a simulator's parameters to 3 real roll-outs per iteration, and its swing-peg policy passed in 90 % of 20 real trials after two iterations, about six roll-outs. In their simulated study, naive wide randomisation failed once the cabinet's offset reached 10 cm, and SimOpt handled 15 cm in three iterations.

7 · Both together, on four robots

The widget puts them together. Four robots have lags of 0, 0.1, 0.2 and 0.25 s (gain 0.9 for all). For each number of probes the calibration is repeated five times with different probes; the simulator's range is centred on each fit, with the fitted gain; a clone is trained in each (40 simulated runs) and run on the robot, 100 runs per point.

Randomise around a measured plant: success on the robot against the width
Top: the lag axis for the chosen robot (gain 0.9). The amber line is its true lag; each row is one of five repeated calibrations: a bar for the lags its simulator draws from (teal if it contains the true lag, red if not), a tick at the fitted lag, and the success of its clone. Pink: where the expert itself fails. Bottom: success on the robot against the half-width, one thin line per calibration, the mean in bold; amber is the slider, the green ring the narrowest width within 2 points of the best. With 0 probes nobody measured and the simulator keeps lag 0 and gain 1.
success on the robot, mean of the calibrations
—
worst calibration
—
ranges containing the true lag
—
fit error, median of the five
—
scatter of the fit, 20 repeats
—
mean at width 0
—
narrowest width within 2 points of the best
—
best mean on the curve
—
Show the core JS
GL.drawPlant = function (range, rng) {
  return GL.plant({ gain: range.gain, delay: range.delay, lag: range.lo + (range.hi - range.lo) * rng() });
};
...
  for (i = 0; i < b.demos; i++) store(BN.rollout(w, BN.expertPolicy(w), rng, { plant: GL.drawPlant(range, rng), jit: GL.JIT, T: T }), false);
  for (r = 0; r < b.rounds; r++)
    for (i = 0; i < b.per; i++) store(BN.rollout(w, function (q) { return GL.predict(nw, q); }, rng, { plant: GL.drawPlant(range, rng), jit: GL.JIT, T: T }), true);
...
GL.response = function (U, j, tau, d) {
  var k = GL.DT / (tau + GL.DT), f = new Float64Array(U.length), v = 0, t, c;
  for (t = 0; t < U.length; t++) { c = t >= d ? Math.max(-1.5, Math.min(1.5, U[t - d][j])) : 0; v = tau > 0 ? v + k * (c - v) : c; f[t] = v; }
  return f;
};
...
      for (t = 0; t < f.length; t++) { sfy += f[t] * runs[ri].Y[t][j]; sff += f[t] * f[t]; syy += runs[ri].Y[t][j] * runs[ri].Y[t][j]; }
...
    g = sfy / sff; sse = syy - g * sfy;
    if (!best || sse < best.sse) best = { gain: g, lag: tau, delay: d, sse: sse };

What to try. Leave the defaults: the robot with 0.2 s of lag, three probes, half-width 0.15 s. The mean success of its five clones is 96 %; at width 0 it is 88.2 % and at 0.3 s 97.8 %, a plateau, and the narrowest width within 2 points of the best is 0.15 s. Set the probes to 0: nobody measured, the simulator says no lag, and success is 2 % at width 0 and 95 % at 0.3 s, so the knee is 0.3 s. One probe (knee 0.3 s again) shows the tail: the mean at width 0 is 74.2 %, but the worst of the five calibrations gets 44 %, rising to 91 % at 0.3 s. Ten probes need no randomisation: 96 % at width 0 and a knee of 0 s. Now change the robot. With lag 0 and ten probes, width 0 gives 96.8 % and width 0.3 s gives 81 %: the price of section 5, paid by the robot the simulator had right. At 0.1 s with no probes, success is 45 % at width 0 and 98 % at 0.1 s. At 0.25 s, near the expert's limit, ten probes give 79.6 % at width 0 and 88.4 % at 0.3 s: close to the limit, width buys margin even after the measurement.

Road not taken · correct the policy on the robot
Lesson 3's labels on the policy's own states, applied on the real arm: it fixes exactly the states the robot makes the clone visit, and it needs a person's label for each of them, thousands of frames (lesson 3 paid for 3,228). A calibration is three seconds of motion that moves every state of the simulator at once. The robot's own attempts return in lesson 12.
What this lesson did not do
The policy is a table that cannot tell which plant it is on; the better real-arm result of Peng et al. (2018), 0.89 against 0.67 for a feedforward policy, came from a recurrent one, which is not built here. Only the plant differed. The rendered images of Tobin et al. (2017) and the contact stiffness and friction of lesson 10 differ from the real thing in the same way, with the same arithmetic and other numbers, and contact cannot be probed in free space. The expert labelled every state for free, and lesson 12 asks what is left when there is none. Caution was measured and not used: the widths stayed inside the expert's limit and the clock stayed at that of lesson 1. Lesson 15 owns how many runs an evaluation needs.

Common mistakes / failure modes

"the simulator's labels are wrong, so the policy is wrong"
The expert passes 100 % of runs at 0.2 s of lag, where the clone passes 0 %. The labels were right and the states were not (§2).
"more simulated data closes the gap"
Eight times the runs: 57.2 to 64.2 % at 0.1 s of lag, 1.8 to 5 % at 0.2 s (§3).
"randomise everything as widely as you can"
The robot the simulator had right loses 24 points at the widest range; four times the runs cut that to 10.8 (§5).
"a wide enough range covers any robot"
At a lag of 0.35 s no width gets past 80 %, and at 0.5 s the expert itself passes 3.5 % (§5).
"calibrate once, then randomisation is unnecessary"
One probe leaves a tail: 1 fit in 10 is off by more than 0.105 s, and the worst of five calibrations gets 44 % at width 0 (§6, §7).

Checkpoint exercise

Try it
One probe fits the lag to a scatter of 0.070 s, and the clone trained with width 0 tolerates about 0.05 s. (a) If the scatter falls as 1/√n, how many probes bring it to 0.02 s? (b) How much robot time is that at 3 s a probe? (c) After three probes, 90 % of fits are within 0.060 s; which width in the widget's grid first covers that? Answer: (a) n = (0.070 / 0.02)² = 12.25, so 13 probes; the table reaches 0.019 s already at ten, because a single probe is the least efficient fit. (b) 13 × 3 s = 39 s of motion. (c) The grid has 0.05 and 0.1 s, so 0.1 s covers it; the widget's knee for the robot at 0.2 s after three probes is 0.15 s, a little wider, because the edge of a range is weaker than its middle (section 4).

Where this points next

A simulator labels any state for free, and a clone copied from it is only as good as the overlap between the states it learned from and the states the robot makes it visit. Drawing the plant at random widens that overlap at a price (the robot the simulator had right loses up to 24 points, and plants past the expert's limit need a slower, more cautious expert), and measuring the robot narrows what has to be drawn: the lag is known to 0.070 s from one probe and 0.011 s from thirty. A simulator is worth using when what is left is narrower than the policy tolerates. Every number above had someone to copy. In the simulator the expert labelled all 7079 frames of 40 runs and scored them for free. For most tasks that matter there is no such expert, only a way to tell whether the task was done, and that is one bit per attempt. Telling a policy that succeeds 96 % of the time from one that succeeds 88.2 % of the time takes about 188 attempts of each (for a 5 % chance of a false alarm and an 80 % chance of seeing a real difference; lesson 15 derives it), and a 3-point difference, 96 against 93 %, about 907: at 30 s an attempt, 1.6 and 7.6 hours of robot per candidate. What can be learned from a score alone, and how many attempts does it take?

Takeaway
A policy cloned in a simulator knows what to do in the states the simulator visited, and a robot with another plant visits others: with no lag in the simulator the clone passes 94 % of its runs there and 0 % on a robot with 0.2 s of lag that the expert passes 100 %. More simulated data does not help (64.2 % at 0.1 s after eight times the runs). Drawing the lag at random in every simulated run puts the robot's states into the table, up to the limit of the expert itself, at a price in precision (24 points for the robot the simulator had right at the widest range) and, past that limit, in a slower expert. Measuring the robot with a few seconds of gentle motion fits gain, lag and delay by least squares, with a scatter that falls as 1/√n (0.070 s for one probe, 0.019 s for ten). The width to draw is what the measurement leaves unknown: the narrowest useful width falls from 0.3 s with no probes to 0.15 s with three and to zero with ten. All of it needs an expert in the simulator, and without one a score costs hundreds of attempts per comparison.

Interview prompts

Companion reads: Synthetic Vision · 02 Where the gap lives (the same gap for rendered images, priced stage by stage), Reinforcement Learning · 04 Reward and simulation (what a simulator is for in RL) and World Models · 31 Capstone (the track this one follows).