What cameras cannot see: contact and force
Lesson 9 turned footage into demonstrations and stopped at what footage cannot hold: how hard something is pushed. This lesson meets the task where that decides everything, a peg going into a hole with a millimetre to spare. A policy that holds a position inserts about two thirds of its attempts at one millimetre; the rest is decided by the error of the camera and the arm, not by anything the policy can learn. The arm's own stiffness turns a stretch of a few micrometres into newtons, so a force sensor sees what the camera cannot. The move is to feed that force back and let the action give way to it. The same demonstrations without the force channel teach nothing of this, and a rule made in a simulator two and a half times too soft loses most of its advantage on the robot.
New idea: give the policy the force as an observation and make its action yield to it, by moving the target along the force it feels. A misalignment that no camera can see is then removed by the contact itself; the price is that the rule exists only where force was recorded and its gains are tied to the stiffness of its world.
Forces next: A policy that feels force and yields to it, instead of holding a position stiffly, inserts the peg where position control jams, and the force it needs is in the data only if someone recorded it with a force sensor in the loop, on a real robot, slowly, with wear. A simulator produces force, actions and unlimited demonstrations, at the price of a physics that is not quite the real one. How wrong can a simulator be before a policy trained in it fails on the robot, and what buys back the difference?
1 · The peg on the Bench, and what holding a position reaches
A peg is lowered into a hole. Work in the frame of the hole: x sideways and z up, both in millimetres, with z = 0 at the rim. A peg of width W in a hole of width W + c has c/2 of slack on each side, so the peg's axis can drop only while |x| ≤ c/2, and a chamfer 1 mm wide at the mouth, its face tilted 30° from the horizontal, guides a peg that arrives a little off the axis. Real tasks are in this range: IndustReal (Tang et al., 2023) and FORGE (Noseworthy et al., 2024) insert pegs and gears with about 0.5 mm of clearance, and Lee et al. (2019) use about 2 mm. The widget sweeps c from 0.2 to 3 mm.
The arm is reduced to what matters here: a tip that a tracker pulls toward a target with a spring of stiffness K = 100 N/mm (lesson 6's stiff tracker). Every 0.05 s the policy commands u, the speed of the target in mm/s. The target starts 2 mm above the rim and the peg is in when its tip is 4 mm below it. What the policy cannot know is where the hole is. The camera places it to 0.5 mm, better than the wrist camera quoted in lesson 9 (6.1 mm), and the arm repeats to 0.3 mm (one standard deviation each), so the hole lies at an error e from where the arm aims, with standard deviation σ = 0.58 mm. A cable and an unbalanced tool also push the tip sideways with a load of 5 N (one standard deviation). A force sensor on the wrist reads the sideways and upward force on the tool, Fx and Fz, with 0.15 N of noise, and a protective stop ends the attempt if the contact force passes 20 N: we call that a jam.
A policy that holds a position does what every policy so far has done: it says where to be. Here that is a target that descends at 1 mm/s over the point where it believes the hole is. The widget below runs every behaviour of this lesson on the same 200 holes. At 1 mm of clearance holding a position inserts 63.5 % of them and ends on the 20 N stop in 36.5 %.
That share is not luck. The peg drops only if the aim error is inside the slack plus whatever the arm can give. At the stop the chamfer pushes the tip sideways with at most 20 N × (sin 30° − 0.15 cos 30°) = 20 N × 0.37 (the second term is friction, 0.15 on the Bench), and the spring gives by that force over its stiffness, g = 7.4 N / 100 N/mm = 0.074 mm. So
P(insert) = P(|e| < c/2 + g) = erf( (c/2 + g) / (σ √2) )
where erf is the Gaussian error function (an error of standard deviation σ stays within ±h with probability erf(h/(σ √2))). It gives 67.5 % at 1 mm (66.1 % measured on 2000 holes; 200 holes scatter by about 3.3 points, which is why 63.5 is not 67.5). No learning enters this formula. It is a ceiling set by the slack and by σ, whatever the policy does with its data.
Why the camera cannot lift it is also a number. A rule that reads the aim error off the camera (0.5 mm of error of its own) and says "inserts" when the reading is small enough is right on 74.9 % of attempts, and saying "inserts" without looking is right on 67.9 %. Attempts that insert have a mean aim error of 0.27 mm and attempts that jam 0.88 mm: the 0.6 mm between them is about the camera's own 0.5 mm of error, which is why a reading cannot tell the two apart. Footage of the approach looks alike for the two, and a policy trained on footage (lesson 9) inherits the ceiling.
2 · Three repairs that stay with position
A policy that only chooses where to be has three knobs. Each is rejected by a number.
| repair | what it costs | what the Bench says |
|---|---|---|
| a better camera or arm (smaller σ) | a sensor several times better | The ceiling moves only when σ falls against the half-width c/2 + g. For 99 % the aim error must be below (c/2 + g)/2.576: 0.223 mm at 1 mm (2.6 times better than σ), 0.126 mm at 0.5 mm (4.6 times), 0.087 mm at 0.3 mm (6.7 times). Halving σ lifts 1 mm from 67.5 to 95.1 % and 0.3 mm from 29.9 to 55.8 %. |
| a stiffer arm | force on the part | The give g = 7.4 N / K shrinks. At 0.5 mm the formula falls from 42.2 % at K = 100 N/mm to 37.7 % at 200 and 34.1 % at 1000; the 200 holes give 44, 39 and 37.5 %, while the mean peak force climbs to 13.5, 15.0 and 19.4 N. |
| a softer arm, with no force input (passive compliance) | precision against every other load | The give grows as 1/K, so a soft arm slides down the chamfer by itself. But the same softness lets the 5 N load shift the tip by 1.7 mm at 3 N/mm, against 0.05 mm at 100. With the best stiffness for each clearance (5 N/mm at 1 mm) it inserts 69 % at 0.3 mm and 80 % at 1 mm. With no load a 2 N/mm arm reaches 94.5 and 98.5 %. |
Each repair moves the number and none removes its cause: the camera error and the arm's softness are two ways of not knowing where the hole is. What removes the cause is a signal that does know.
3 · Contact is a spring: the signal the task depends on
The tracker is a spring between where the arm was told to be and where the world lets it be. Whatever holds the tip away from its target stretches the spring, and the spring pushes back with F = K × stretch. The wall of the hole is a second spring (1000 N/mm on the Bench), in series with the arm: 100 and 1000 give 91 N/mm, sideways or straight down. A stretch of 0.1 mm is therefore 9.1 N, which a wrist sensor with 0.15 N of noise resolves to 1.65 µm, 303 times finer than the camera's 0.5 mm. The camera cannot place the hole to better than half a millimetre; the force says, to a micrometre, how far the world holds the tip from where it was told to be.
The force also says which way. A push N on the chamfer is returned along the face's normal, tilted 30° from the vertical: a sideways part N sin 30° toward the axis, less friction's 0.15 N cos 30°, in all 0.37 N per newton. The sign of the sideways force names the side the hole is on, and the upward force says how hard the arm is pressing.
And it says why the stiff arm jams. The arm does not give: its target keeps coming at 1 mm/s, the spring stretches further, and the force climbs at about 91 N per millimetre of target until the stop. The stiffness that made the arm precise is the stiffness that makes it press, and what the contact offers, a push toward the axis, is spent on a spring that will not move.
The constraints on a repair follow, with no taste in them. (i) The signal must be force, sideways and up, in what the policy senses. (ii) The action must let the target move along the force. (iii) It must stay stiff where nothing touches, so that loads do not move it. (iv) It must keep the force below the stop.
4 · The move: an action that yields
Give the policy the two forces each step and let the command depend on them, one rule per axis. Upward, hold a push: uz = clip(−Cz (Fpush − Fz), −v, v), with Fpush = 10 N (half the stop), v = 1 mm/s and Cz = v / Fpush = 0.1 mm/s per N: with nothing pushing back the target descends at v, with 10 N of push it stops. Sideways, move the target along the force: ux = Cx · dead(Fx), with Cx = 0.35 mm/s per N, where dead(·) sets forces under 0.4 N (2.7 times the sensor's noise) to zero. The tracker stays at 100 N/mm: the arm is stiff against the load and yields only along the force it feels. The Bench hardly needs the dead band, since any width from 0 to 2 N inserts the same share at 1 mm, and the sideways gain is forgiving: from 0.02 to 1 mm/s per N the share inserted at 0.3 mm is the same 95 % and only the time moves, 8.7 s to 6.3 s; at 3 the target outruns the contact and 30 % end on the stop.
The gains are not free. The target moves by uz Δt each step and the force changes by the vertical stiffness of arm and wall in series, k = 90.9 N/mm (a 30° face lowers it by 2.9 %, which the formulas ignore), times that, so the force error obeys en+1 = (1 − γ) en with γ = Cz k Δt = 0.45. The share of the error left after a step, the loop gain of lessons 2 and 6, is 0.55 (0.57 measured on a flat rim), and the loop is stable while γ < 2.
The approach speed has its own limit. The tip follows its target with a lag τ = 0.1 s (the arm's damping, 10 N s/mm, over its stiffness), so when it meets the surface the target is already v τ past it and moves another v Δt before the first reading arrives. The contact spring is compressed by v (τ + Δt), and the first reading can be as large as Ffirst = k v (Δt + τ) = 13.6 N (measured on a flat rim: 12.4 N on average, 12.7 N at most). The stop is 20 N, so 1 mm/s leaves a margin of one and a half; the speed that would use half the stop is 0.73 mm/s. This is why an insertion takes 6.3 s.
What it does is small. At 1 mm, 35.8 % of attempts touch the chamfer; each moves its target 0.39 mm toward the axis in 0.47 s, and drops in.
The widget
What to try. Leave the defaults, "holds a position" at 1 mm and 100 N/mm: it inserts 63.5 % (95 % interval 56.6 to 69.9) against the formula's 67.5 %, ends on the stop in 36.5 % and presses with a mean peak of 8.7 N. Slide the clearance to 0.3 mm: 32 % against 29.9. At 1 mm soften the arm to 5 N/mm: 80 %. Choose "yields to the force": at 1 mm 98.5 %, mean peak 3.9 N, 6.3 s to insert; at 0.2 mm 93 % against 23.5 % for holding, with 7 % timed out, the holes beyond the funnel. Choose "learned, with force" at 1 mm: with 1 demonstration 82.5 %, with 5 the same 82.5 %, with 10 98.5 %: this draw of demonstrations touches the funnel on one side up to 5 demonstrations and on both from 10. Switch to "learned, without force": 10 demonstrations give 64 % and 40 give 63 %. Back on "yields to the force", raise the world stiffness: 2 times still gives 98.5 %, 2.5 times 69 %, 4 times 64 %.
The behaviours on the same 200 holes, with the formula for the first:
| clearance | ceiling of holding a position (formula) | holds a position (100 N/mm) | best soft arm (N/mm) | yields to the force |
|---|---|---|---|---|
| 0.3 mm | 29.9 % | 32 % | 69 % (5) | 95 % |
| 1 mm | 67.5 % | 63.5 % | 80 % (5) | 98.5 % |
| 2 mm | 93.5 % | 93 % | 95.5 % (20) | 100 % |
Yielding reaches the funnel's own ceiling, erf((c/2 + 1 mm)/(σ √2)): 99.0 % at 1 mm, 95.1 % at 0.3. What is left are the holes more than a chamfer off the axis: the arm lands on the flat rim, feels only a vertical push, and times out. The force says that it is pressing, not which way to go.
5 · The force has to be in the data
So far the yielding rule was written by hand. A learned policy gets it only from demonstrations, and a demonstration is what was sensed and what was done. Record m demonstrations of the yielding rule at 1 mm, with the frames (depth, sideways target offset, Fx, Fz, command) thinned to one per half bandwidth, and fit lesson 1's nearest-demo learner (bandwidths 1.5 mm, 1.0 mm, 0.6 N, 1.5 N). Then do the same with the force columns deleted: the same attempts, the same commands, and a learner that sees only depth and target.
Without the force the clone cannot yield. When contact starts, the depth and the target offset are the same whether the hole is on the left or the right, and the command is +u in one case and −u in the other; the best squared-error answer is their mean, about zero. The clone descends, holds its target and jams like the position policy: 63.5 % with 5 demonstrations and 63.4 % with 40, averaged over 24 independent draws of the demonstrations, against 63.5 % for holding a position on the same holes. Eight times the data does nothing, because the cause is not in it. Footage cannot supply it either: the whole slide is 0.39 mm on average, smaller than the camera's 0.5 mm of error, so labels inferred from footage as in lesson 9 cannot tell it from noise.
With force the clone learns the rule, but only on the sides it has seen. A hole on the right is answered only if some demonstration touched the funnel on the right, and a given side is touched by a fraction q = 0.19 of attempts. Over 24 draws and six sizes of set, a clone that saw neither side inserts 63.5 %, which is the position policy; one that saw one side 81.0 %; one that saw both 98.3 %. So the expected success after m demonstrations is 63.5 + 34.9 (1 − (1 − q)m), and it follows the measured mean: for m = 1, 5, 10 and 20, 71.5, 86.0, 93.4 and 98.0 % measured against 70.2, 86.3, 94.2 and 97.8 % from the formula. About 15 demonstrations make each side 95 % likely to have been seen. This is lesson 7's coverage law in miniature: the demonstrator does not choose which holes come. The rule itself does not depend on the clearance: demonstrations at 1 mm give 95.5 % at 0.3 mm, level with the funnel's ceiling of 95.1 % within the scatter of 200 holes.
Force as an observation is half of the move. Hou et al. (2024), on an item-flipping task, compared a diffusion policy with a stiff controller and force input (14 % success), one with a uniformly soft controller (23 %), and one that also outputs its own stiffness (96 %). ForceMimic (Liu et al., 2024) found a diffusion policy with force input but no force-position control worse than the one without it (10 % against 55 % on its strict peeling criterion, 10 trials for the first) and a hybrid force-position policy at 85 %.
6 · What force data cost, and what a simulator's physics does to the rule
The force has to be recorded with the sensor in the loop, on the arm that will use it, and slowly. A yielding insertion takes 6.3 s at the speed section 4 allowed, an attempt that times out takes 30 s, and the Bench prices a reset at 20 s, so a demonstration cycle averages 26.6 s and 20 demonstrations of one arm and one part take 8.9 minutes. Force costs time even for people: ForceMimic measured 2.9 minutes to peel a zucchini with a bare hand, 4.5 with a handheld force-capture device and 13 with force-feedback teleoperation. Gathering force data by trying with position control costs the hardware instead: at 0.5 mm, 56 % of attempts end on the 20 N stop, which at 149.2 attempts an hour is 83.5 stops an hour.
A simulator gives force, commands and unlimited demonstrations. What it cannot promise is the stiffness, and the rule is tuned to one. The approach speed was chosen so that the first reading after contact, Ffirst = k v (Δt + τ), stays under the stop. If the world is ρ times stiffer, k becomes ρ k and the lag shrinks to τ/ρ, so Ffirst(ρ) = k v (ρ Δt + τ) with k the Bench's: 13.6 N at ρ = 1, 18.2 N at 2 and 22.7 N at 3. The stop is reached at ρ = 2.40.
Run the rule, and a clone made from it, in a world whose arm and part are ρ times stiffer than the one they were made in. At ρ = 2 both insert 98.5 %. At 2.5 they insert 69 % and 65 %, and at 4 they insert 64 % and 64 % while holding a position gets 58.5 %. The success passes the midpoint between its two plateaus at ρ = 2.38, against 2.40 from the formula. Every attempt that touches the funnel ends on the stop, and far past it: the jams end at 53 N on average. What was gained over holding a position is mostly gone (69 % against 62 % at 2.5 times). The rule had encoded the stiffness of its world.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
A policy that feels the force and yields to it inserts 98.5 % of its attempts at 1 mm where holding a position inserts 63.5 %, and 93 % at 0.2 mm where holding a position can reach at most 23.5 %. The force is in the data only if it was recorded: without it a clone is the position policy (63.4 % after 40 demonstrations), and with it a clone needs both sides of the funnel (about 15 demonstrations for a 95 % chance of each). Recording it costs robot time (20 demonstrations of one part take 8.9 minutes) and presses the hardware (83.5 stops an hour at 0.5 mm with a stiff arm), and a simulator's force carries the simulator's stiffness: a rule made at one stiffness inserts 98.5 % at twice that and 69 % at two and a half times. How wrong can a simulator be before a policy trained in it fails on the robot, and what buys back the difference?
Interview prompts
- Why does a policy that only chooses positions have a ceiling on insertion, and what does it depend on? (§1 — the peg drops only if the aim error is inside the slack plus the arm's give: erf((c/2 + g)/(σ√2)), a function of σ, c and g and not of the data.)
- What does a wrist force sensor see that a camera cannot? (§3 — the arm's spring turns a micrometre stretch into newtons, about 300 times finer than a 0.5 mm camera error, and the chamfer's tilt names the side.)
- How do you make a policy yield, and what bounds its gains? (§4 — move the target along the dead-banded force and regulate the push; the push loop is stable for Cz k Δt < 2, and the approach speed is bounded by the first reading after contact.)
- Passive compliance already centres a peg. Why add a sensor? (§2, §4 — one stiffness for the whole move is shifted by every sideways load, 80 % against 98.5 % at 1 mm; the sensor lets the arm stay stiff and yield only along the force.)
- Why must demonstrations include force? (§5 — the command depends on a cause the stripped observation lacks, so the clone is the position policy; with force it learns only the sides of the funnel it has seen.)
- What goes wrong when a force-aware policy is trained in a simulator with the wrong stiffness? (§6 — the first reading after contact scales with stiffness, so a rule tuned at one stiffness jams at 2.5 times it.)
Companion reads: World Models · 15 Bodies, contact, language (contact as a switch, and force as what a camera does not measure), Computer Vision 3D · 07 What a pixel measures (what a camera records, and so what it cannot) and Reinforcement Learning · 75 Robotic grasping (grasping, where force and position are separate channels).