all_lessons/Robot Model Training/10 · Contactlesson 10 / 24

What cameras cannot see: contact and force

Lesson 9 turned footage into demonstrations and stopped at what footage cannot hold: how hard something is pushed. This lesson meets the task where that decides everything, a peg going into a hole with a millimetre to spare. A policy that holds a position inserts about two thirds of its attempts at one millimetre; the rest is decided by the error of the camera and the arm, not by anything the policy can learn. The arm's own stiffness turns a stretch of a few micrometres into newtons, so a force sensor sees what the camera cannot. The move is to feed that force back and let the action give way to it. The same demonstrations without the force channel teach nothing of this, and a rule made in a simulator two and a half times too soft loses most of its advantage on the robot.

The thesis, here
A task decided below the resolution of the sensors cannot be done by choosing positions, because the chooser never sees what decides it. Contact makes it visible: an arm of stiffness K turns a stretch x into a force K x, so a misalignment of micrometres is newtons, and a funnel says which way to move. The policy therefore needs the force in what it senses and an action that moves its own target along that force instead of holding it. The rule that results can be learned only from demonstrations that contain force, and its gains belong to the stiffness of the world it was made in.
Linear position
Forced by: A model that infers the missing actions from a few hours of labelled motion can turn a large pile of video into demonstrations, and what it learns is what to do next, not how hard to push. Video records where things moved, never the force that moved them: an insertion with a millimetre of clearance, watched through a camera that cannot resolve a millimetre, looks the same when it succeeds and when it jams. What must a policy sense, and how must it act, when the task is decided by something the camera cannot see?
New idea: give the policy the force as an observation and make its action yield to it, by moving the target along the force it feels. A misalignment that no camera can see is then removed by the contact itself; the price is that the rule exists only where force was recorded and its gains are tied to the stiffness of its world.
Forces next: A policy that feels force and yields to it, instead of holding a position stiffly, inserts the peg where position control jams, and the force it needs is in the data only if someone recorded it with a force sensor in the loop, on a real robot, slowly, with wear. A simulator produces force, actions and unlimited demonstrations, at the price of a physics that is not quite the real one. How wrong can a simulator be before a policy trained in it fails on the robot, and what buys back the difference?
The plan
Six moves. (1) Put the peg on the Bench and measure what holding a position reaches. (2) Try the three repairs that stay with position. (3) Read the contact as a spring and find the signal the task depends on. (4) Make the action yield to it and set its gains. (5) Learn it from demonstrations, with and without the force. (6) Price the force data, and see what a simulator's physics does to the rule.

1 · The peg on the Bench, and what holding a position reaches

A peg is lowered into a hole. Work in the frame of the hole: x sideways and z up, both in millimetres, with z = 0 at the rim. A peg of width W in a hole of width W + c has c/2 of slack on each side, so the peg's axis can drop only while |x| ≤ c/2, and a chamfer 1 mm wide at the mouth, its face tilted 30° from the horizontal, guides a peg that arrives a little off the axis. Real tasks are in this range: IndustReal (Tang et al., 2023) and FORGE (Noseworthy et al., 2024) insert pegs and gears with about 0.5 mm of clearance, and Lee et al. (2019) use about 2 mm. The widget sweeps c from 0.2 to 3 mm.

The arm is reduced to what matters here: a tip that a tracker pulls toward a target with a spring of stiffness K = 100 N/mm (lesson 6's stiff tracker). Every 0.05 s the policy commands u, the speed of the target in mm/s. The target starts 2 mm above the rim and the peg is in when its tip is 4 mm below it. What the policy cannot know is where the hole is. The camera places it to 0.5 mm, better than the wrist camera quoted in lesson 9 (6.1 mm), and the arm repeats to 0.3 mm (one standard deviation each), so the hole lies at an error e from where the arm aims, with standard deviation σ = 0.58 mm. A cable and an unbalanced tool also push the tip sideways with a load of 5 N (one standard deviation). A force sensor on the wrist reads the sideways and upward force on the tool, Fx and Fz, with 0.15 N of noise, and a protective stop ends the attempt if the contact force passes 20 N: we call that a jam.

A policy that holds a position does what every policy so far has done: it says where to be. Here that is a target that descends at 1 mm/s over the point where it believes the hole is. The widget below runs every behaviour of this lesson on the same 200 holes. At 1 mm of clearance holding a position inserts 63.5 % of them and ends on the 20 N stop in 36.5 %.

That share is not luck. The peg drops only if the aim error is inside the slack plus whatever the arm can give. At the stop the chamfer pushes the tip sideways with at most 20 N × (sin 30° − 0.15 cos 30°) = 20 N × 0.37 (the second term is friction, 0.15 on the Bench), and the spring gives by that force over its stiffness, g = 7.4 N / 100 N/mm = 0.074 mm. So

P(insert) = P(|e| < c/2 + g) = erf( (c/2 + g) / (σ √2) )

where erf is the Gaussian error function (an error of standard deviation σ stays within ±h with probability erf(h/(σ √2))). It gives 67.5 % at 1 mm (66.1 % measured on 2000 holes; 200 holes scatter by about 3.3 points, which is why 63.5 is not 67.5). No learning enters this formula. It is a ceiling set by the slack and by σ, whatever the policy does with its data.

Why the camera cannot lift it is also a number. A rule that reads the aim error off the camera (0.5 mm of error of its own) and says "inserts" when the reading is small enough is right on 74.9 % of attempts, and saying "inserts" without looking is right on 67.9 %. Attempts that insert have a mean aim error of 0.27 mm and attempts that jam 0.88 mm: the 0.6 mm between them is about the camera's own 0.5 mm of error, which is why a reading cannot tell the two apart. Footage of the approach looks alike for the two, and a policy trained on footage (lesson 9) inherits the ceiling.

2 · Three repairs that stay with position

A policy that only chooses where to be has three knobs. Each is rejected by a number.

repairwhat it costswhat the Bench says
a better camera or arm (smaller σ)a sensor several times betterThe ceiling moves only when σ falls against the half-width c/2 + g. For 99 % the aim error must be below (c/2 + g)/2.576: 0.223 mm at 1 mm (2.6 times better than σ), 0.126 mm at 0.5 mm (4.6 times), 0.087 mm at 0.3 mm (6.7 times). Halving σ lifts 1 mm from 67.5 to 95.1 % and 0.3 mm from 29.9 to 55.8 %.
a stiffer armforce on the partThe give g = 7.4 N / K shrinks. At 0.5 mm the formula falls from 42.2 % at K = 100 N/mm to 37.7 % at 200 and 34.1 % at 1000; the 200 holes give 44, 39 and 37.5 %, while the mean peak force climbs to 13.5, 15.0 and 19.4 N.
a softer arm, with no force input (passive compliance)precision against every other loadThe give grows as 1/K, so a soft arm slides down the chamfer by itself. But the same softness lets the 5 N load shift the tip by 1.7 mm at 3 N/mm, against 0.05 mm at 100. With the best stiffness for each clearance (5 N/mm at 1 mm) it inserts 69 % at 0.3 mm and 80 % at 1 mm. With no load a 2 N/mm arm reaches 94.5 and 98.5 %.

Each repair moves the number and none removes its cause: the camera error and the arm's softness are two ways of not knowing where the hole is. What removes the cause is a signal that does know.

Road not taken · passive compliance
A compliant wrist, or a soft tracker, centres the peg in the funnel with no sensor at all, and on the Bench it is the best of the three repairs. It is soft all the time. The approach, which needs the arm stiff so that loads do not move it, and the contact, which needs it soft, share one stiffness, so with the 5 N load it inserts 80 % at 1 mm against 98.5 %. Nothing in it can be told to push less. It returns below as a policy that is stiff in free space and soft only along the force it feels.

3 · Contact is a spring: the signal the task depends on

The tracker is a spring between where the arm was told to be and where the world lets it be. Whatever holds the tip away from its target stretches the spring, and the spring pushes back with F = K × stretch. The wall of the hole is a second spring (1000 N/mm on the Bench), in series with the arm: 100 and 1000 give 91 N/mm, sideways or straight down. A stretch of 0.1 mm is therefore 9.1 N, which a wrist sensor with 0.15 N of noise resolves to 1.65 µm, 303 times finer than the camera's 0.5 mm. The camera cannot place the hole to better than half a millimetre; the force says, to a micrometre, how far the world holds the tip from where it was told to be.

The force also says which way. A push N on the chamfer is returned along the face's normal, tilted 30° from the vertical: a sideways part N sin 30° toward the axis, less friction's 0.15 N cos 30°, in all 0.37 N per newton. The sign of the sideways force names the side the hole is on, and the upward force says how hard the arm is pressing.

And it says why the stiff arm jams. The arm does not give: its target keeps coming at 1 mm/s, the spring stretches further, and the force climbs at about 91 N per millimetre of target until the stop. The stiffness that made the arm precise is the stiffness that makes it press, and what the contact offers, a push toward the axis, is spent on a spring that will not move.

The constraints on a repair follow, with no taste in them. (i) The signal must be force, sideways and up, in what the policy senses. (ii) The action must let the target move along the force. (iii) It must stay stiff where nothing touches, so that loads do not move it. (iv) It must keep the force below the stop.

4 · The move: an action that yields

Give the policy the two forces each step and let the command depend on them, one rule per axis. Upward, hold a push: uz = clip(−Cz (Fpush − Fz), −v, v), with Fpush = 10 N (half the stop), v = 1 mm/s and Cz = v / Fpush = 0.1 mm/s per N: with nothing pushing back the target descends at v, with 10 N of push it stops. Sideways, move the target along the force: ux = Cx · dead(Fx), with Cx = 0.35 mm/s per N, where dead(·) sets forces under 0.4 N (2.7 times the sensor's noise) to zero. The tracker stays at 100 N/mm: the arm is stiff against the load and yields only along the force it feels. The Bench hardly needs the dead band, since any width from 0 to 2 N inserts the same share at 1 mm, and the sideways gain is forgiving: from 0.02 to 1 mm/s per N the share inserted at 0.3 mm is the same 95 % and only the time moves, 8.7 s to 6.3 s; at 3 the target outruns the contact and 30 % end on the stop.

The gains are not free. The target moves by uz Δt each step and the force changes by the vertical stiffness of arm and wall in series, k = 90.9 N/mm (a 30° face lowers it by 2.9 %, which the formulas ignore), times that, so the force error obeys en+1 = (1 − γ) en with γ = Cz k Δt = 0.45. The share of the error left after a step, the loop gain of lessons 2 and 6, is 0.55 (0.57 measured on a flat rim), and the loop is stable while γ < 2.

The approach speed has its own limit. The tip follows its target with a lag τ = 0.1 s (the arm's damping, 10 N s/mm, over its stiffness), so when it meets the surface the target is already v τ past it and moves another v Δt before the first reading arrives. The contact spring is compressed by v (τ + Δt), and the first reading can be as large as Ffirst = k v (Δt + τ) = 13.6 N (measured on a flat rim: 12.4 N on average, 12.7 N at most). The stop is 20 N, so 1 mm/s leaves a margin of one and a half; the speed that would use half the stop is 0.73 mm/s. This is why an insertion takes 6.3 s.

What it does is small. At 1 mm, 35.8 % of attempts touch the chamfer; each moves its target 0.39 mm toward the axis in 0.47 s, and drops in.

The widget

Insert the peg: hold a position, or yield to the force
Left: the hole in section (mm) with 12 of 200 attempts, the path of the peg's axis (teal inserted, red ended on the 20 N stop, amber timed out); purple is one standard deviation of where the hole is. Right: their force against time, and the share inserted against clearance for the chosen behaviour (grey: ceiling of holding a position, green: ceiling of the funnel, teal: the chosen behaviour at each clearance, amber: its share at the chosen clearance with its 95 % interval). Clearance is the first control; stiffness acts on "holds a position", demonstrations on the learned behaviours, and world stiffness on all of them.
inserted
—
95 % interval
—
ended on the 20 N stop
—
timed out
—
mean peak force
—
time to insert
—
ceiling, holds a position
—
ceiling, the funnel
—
Show the core JS
IL.yielding = function () {
  var P = IL.P;
  return { K: P.K, act: function (ob) {
    var uz = -P.Cz * (P.Fpush - ob.fz); uz = uz > P.v ? P.v : uz < -P.v ? -P.v : uz;
    return [P.Cx * IL.dead(ob.fx, P.fdead), uz];
  } };
};
...
      var c = IL.contact(W, x, z, vx, vz); cx = c[0]; cz = c[1];
      var x1 = (x + a * (Kt * (e0 + xc) + FL + cx)) / (1 + a * Kt), z1 = (z + a * (Kt * zc + cz)) / (1 + a * Kt);   // the spring of the arm is taken implicitly
      vx = (x1 - x) / h; vz = (z1 - z) / h; x = x1; z = z1;

What to try. Leave the defaults, "holds a position" at 1 mm and 100 N/mm: it inserts 63.5 % (95 % interval 56.6 to 69.9) against the formula's 67.5 %, ends on the stop in 36.5 % and presses with a mean peak of 8.7 N. Slide the clearance to 0.3 mm: 32 % against 29.9. At 1 mm soften the arm to 5 N/mm: 80 %. Choose "yields to the force": at 1 mm 98.5 %, mean peak 3.9 N, 6.3 s to insert; at 0.2 mm 93 % against 23.5 % for holding, with 7 % timed out, the holes beyond the funnel. Choose "learned, with force" at 1 mm: with 1 demonstration 82.5 %, with 5 the same 82.5 %, with 10 98.5 %: this draw of demonstrations touches the funnel on one side up to 5 demonstrations and on both from 10. Switch to "learned, without force": 10 demonstrations give 64 % and 40 give 63 %. Back on "yields to the force", raise the world stiffness: 2 times still gives 98.5 %, 2.5 times 69 %, 4 times 64 %.

The behaviours on the same 200 holes, with the formula for the first:

clearanceceiling of holding a position (formula)holds a position (100 N/mm)best soft arm (N/mm)yields to the force
0.3 mm29.9 %32 %69 % (5)95 %
1 mm67.5 %63.5 %80 % (5)98.5 %
2 mm93.5 %93 %95.5 % (20)100 %

Yielding reaches the funnel's own ceiling, erf((c/2 + 1 mm)/(σ √2)): 99.0 % at 1 mm, 95.1 % at 0.3. What is left are the holes more than a chamfer off the axis: the arm lands on the flat rim, feels only a vertical push, and times out. The force says that it is pressing, not which way to go.

5 · The force has to be in the data

So far the yielding rule was written by hand. A learned policy gets it only from demonstrations, and a demonstration is what was sensed and what was done. Record m demonstrations of the yielding rule at 1 mm, with the frames (depth, sideways target offset, Fx, Fz, command) thinned to one per half bandwidth, and fit lesson 1's nearest-demo learner (bandwidths 1.5 mm, 1.0 mm, 0.6 N, 1.5 N). Then do the same with the force columns deleted: the same attempts, the same commands, and a learner that sees only depth and target.

Without the force the clone cannot yield. When contact starts, the depth and the target offset are the same whether the hole is on the left or the right, and the command is +u in one case and −u in the other; the best squared-error answer is their mean, about zero. The clone descends, holds its target and jams like the position policy: 63.5 % with 5 demonstrations and 63.4 % with 40, averaged over 24 independent draws of the demonstrations, against 63.5 % for holding a position on the same holes. Eight times the data does nothing, because the cause is not in it. Footage cannot supply it either: the whole slide is 0.39 mm on average, smaller than the camera's 0.5 mm of error, so labels inferred from footage as in lesson 9 cannot tell it from noise.

With force the clone learns the rule, but only on the sides it has seen. A hole on the right is answered only if some demonstration touched the funnel on the right, and a given side is touched by a fraction q = 0.19 of attempts. Over 24 draws and six sizes of set, a clone that saw neither side inserts 63.5 %, which is the position policy; one that saw one side 81.0 %; one that saw both 98.3 %. So the expected success after m demonstrations is 63.5 + 34.9 (1 − (1 − q)m), and it follows the measured mean: for m = 1, 5, 10 and 20, 71.5, 86.0, 93.4 and 98.0 % measured against 70.2, 86.3, 94.2 and 97.8 % from the formula. About 15 demonstrations make each side 95 % likely to have been seen. This is lesson 7's coverage law in miniature: the demonstrator does not choose which holes come. The rule itself does not depend on the clearance: demonstrations at 1 mm give 95.5 % at 0.3 mm, level with the funnel's ceiling of 95.1 % within the scatter of 200 holes.

Force as an observation is half of the move. Hou et al. (2024), on an item-flipping task, compared a diffusion policy with a stiff controller and force input (14 % success), one with a uniformly soft controller (23 %), and one that also outputs its own stiffness (96 %). ForceMimic (Liu et al., 2024) found a diffusion policy with force input but no force-position control worse than the one without it (10 % against 55 % on its strict peeling criterion, 10 trials for the first) and a hybrid force-position policy at 85 %.

6 · What force data cost, and what a simulator's physics does to the rule

The force has to be recorded with the sensor in the loop, on the arm that will use it, and slowly. A yielding insertion takes 6.3 s at the speed section 4 allowed, an attempt that times out takes 30 s, and the Bench prices a reset at 20 s, so a demonstration cycle averages 26.6 s and 20 demonstrations of one arm and one part take 8.9 minutes. Force costs time even for people: ForceMimic measured 2.9 minutes to peel a zucchini with a bare hand, 4.5 with a handheld force-capture device and 13 with force-feedback teleoperation. Gathering force data by trying with position control costs the hardware instead: at 0.5 mm, 56 % of attempts end on the 20 N stop, which at 149.2 attempts an hour is 83.5 stops an hour.

A simulator gives force, commands and unlimited demonstrations. What it cannot promise is the stiffness, and the rule is tuned to one. The approach speed was chosen so that the first reading after contact, Ffirst = k v (Δt + τ), stays under the stop. If the world is ρ times stiffer, k becomes ρ k and the lag shrinks to τ/ρ, so Ffirst(ρ) = k v (ρ Δt + τ) with k the Bench's: 13.6 N at ρ = 1, 18.2 N at 2 and 22.7 N at 3. The stop is reached at ρ = 2.40.

Run the rule, and a clone made from it, in a world whose arm and part are ρ times stiffer than the one they were made in. At ρ = 2 both insert 98.5 %. At 2.5 they insert 69 % and 65 %, and at 4 they insert 64 % and 64 % while holding a position gets 58.5 %. The success passes the midpoint between its two plateaus at ρ = 2.38, against 2.40 from the formula. Every attempt that touches the funnel ends on the stop, and far past it: the jams end at 53 N on average. What was gained over holding a position is mostly gone (69 % against 62 % at 2.5 times). The rule had encoded the stiffness of its world.

What this lesson did not do
The Bench's peg is a point: a real peg also tilts and wedges, which no point can do, and it has two sideways axes, not one. Holes beyond the funnel (1.5 % at 1 mm, 7 % at 0.2 mm) time out; the force can guide a search for them, which this lesson does not build. The force is read with no delay; a real wrist adds some, and delay tightens the gain bound as in lesson 6. Force is not always the missing signal: FORGE (Noseworthy et al., 2024) scored 0.84 on an 8 mm peg with its force-aware method and 0.82 for its no-force variant, and 0.69 against 0.40 on a nut; IndustReal (Tang et al., 2023) inserted pegs with 0.5 to 0.6 mm of clearance with no force sensor, at 76.7 % on its real pegs. How wrong a simulator may be, and what buys the difference back, is lesson 11; one network that reads, feels and emits chunks inside a control step is lesson 13.

Common mistakes / failure modes

"a better camera is the fix"
The ceiling moves only when σ falls against c/2 + g: for 99 % the error must fall 4.6 times at 0.5 mm and 6.7 times at 0.3 mm (§2).
"a stiffer arm is more precise, so it inserts more"
At 0.5 mm, 100 to 1000 N/mm takes insertion from 44 to 37.5 % and the mean peak force to 19.4 N (§2).
"a soft arm is the same as a force-aware one"
With the sideways load the best soft arm inserts 80 % at 1 mm against 98.5 %; with no load it reaches 98.5 % (§2, §4).
"force in the observation is enough"
Only if the action uses it: holding the target ignores force (63.5 %), and a stiff policy with force input scored 14 % against 96 % in Hou et al. (2024) (§4, §5).
"force can be added to demonstrations later"
Without the force columns the clone inserts 63.5 % at 5 demonstrations and 63.4 % at 40 (§5).
"a policy made in a simulator transfers if the geometry is right"
98.5 % at the stiffness it was made in or twice it, 69 % at 2.5 times (§6).

Checkpoint exercise

Try it
A hole has 0.4 mm of clearance, the camera and arm together place it to σ = 0.5 mm, and the arm gives 0.07 mm at its stop. (a) What share of insertions can holding a position reach? (b) How many times better must σ be for 99 %? (c) A wrist sensor with 0.2 N of noise sits on an arm and wall that stretch like an 80 N/mm spring: what stretch can it resolve, and how many times finer is that than the camera's 0.5 mm? Answer: (a) The half-width is 0.4/2 + 0.07 = 0.27 mm, so the share is erf(0.27/(0.5 √2)) = 41.1 %. (b) For 99 % the half-width must be 2.576 σ, so σ = 0.27/2.576 = 0.105 mm: 4.8 times better. (c) 0.2 N / 80 N/mm = 2.5 µm, 200 times finer.

Where this points next

A policy that feels the force and yields to it inserts 98.5 % of its attempts at 1 mm where holding a position inserts 63.5 %, and 93 % at 0.2 mm where holding a position can reach at most 23.5 %. The force is in the data only if it was recorded: without it a clone is the position policy (63.4 % after 40 demonstrations), and with it a clone needs both sides of the funnel (about 15 demonstrations for a 95 % chance of each). Recording it costs robot time (20 demonstrations of one part take 8.9 minutes) and presses the hardware (83.5 stops an hour at 0.5 mm with a stiff arm), and a simulator's force carries the simulator's stiffness: a rule made at one stiffness inserts 98.5 % at twice that and 69 % at two and a half times. How wrong can a simulator be before a policy trained in it fails on the robot, and what buys back the difference?

Takeaway
A policy that only chooses positions has a ceiling on a task decided below the camera's resolution: at 1 mm of clearance with 0.58 mm of error it is 67.5 %, and it moves only with σ against the half-width c/2 + g. Stiffer arms press harder and insert less; softer arms slide in and lose to every sideways load (80 %). Contact turns a stretch into a force 303 times finer than the camera sees, and a funnel names the side. A policy that keeps a stiff tracker but moves its target along the force it feels inserts 98.5 % at 1 mm, with the push loop stable for γ < 2 and the approach speed bounded by the first reading after contact. It learns this only from demonstrations that contain force (63.4 % without, 98.3 % with both sides seen), and the rule carries the stiffness of its world: 69 % at 2.5 times.

Interview prompts

Companion reads: World Models · 15 Bodies, contact, language (contact as a switch, and force as what a camera does not measure), Computer Vision 3D · 07 What a pixel measures (what a camera records, and so what it cannot) and Reinforcement Learning · 75 Robotic grasping (grasping, where force and position are separate channels).