all_lessons/World Models/15 · Bodies and contactlesson 15 / 31

Bodies, contact, and words

Lesson 14 ended at a wall: its smooth model, good in free flight, was many times worse at the steps that touch one, even when it had trained on bounces. A body is touched, not only seen, and a touch is a jump. A smooth network trained on bounces predicts a wall that slows the ball gradually, so nearly all its error sits in the steps next to a wall. This lesson shows why a smooth model must blur a jump and replaces it with a hybrid one: a mode, a simple law inside each mode, and a switch. On a table the switch is friction, a force compared with a threshold, and a camera cannot read force. Words then enter as costs over what the model predicts. The lesson ends with a model that scores better on the usual test and decides worse.

The thesis, here
Contact is a switch, not a slope. A model of a body carries a discrete mode (free or touching, stuck or sliding), a simple law inside each mode, and a rule that picks the mode by comparing a quantity with a threshold. A smooth network can only blur that rule, and the blur sits where decisions are made. The quantity is often a force, which a camera does not measure and a touch channel does. A goal in words is a cost over what the model predicts, so a cost such as "gently" can steer a planner to the very place where the model is wrong.
Linear position
Forced by: Actions can be inferred from video and the model can run in real time, for a world seen through a screen. A robot's world is touched, not only seen: contact is discontinuous, force is invisible to a camera, and the goal arrives in words. What changes when the agent has a body, and the model must predict what it will feel?
New idea: a model of a body is hybrid: a discrete mode chosen by comparing a measured quantity with a threshold, and a smooth law inside each mode. A smooth network can only blur the switch. Once the switch is one comparison, the laws inside the modes are, here, linear; the quantity compared, a force, has to be measured.
Forces next: We now have every piece: state, dynamics, uncertainty, rollouts, causality, practice, planning, abstraction, pixels, memory, objects, actions and a body, each justified by an exam the previous model failed. How do we know that a particular model passes the exam that matters, a better decision in the world, and not only the ones that look good?
The plan
Seven moves. (1) Measure where a smooth model fails at a wall. (2) Show that it must blur a jump, and that training harder does not remove the blur. (3) Split the model into modes, laws and a switch. (4) Derive stick and slip, a switch on a force. (5) Show what a camera cannot see and a touch channel can. (6) Turn words into costs. (7) Score the models twice, by one step and by decisions.

1 · The wall is a step

Lesson 14 ended at the walls; look at one wall, on one axis. A free ball obeys p′ = p + G·v and v′ = D·v (primes: one step later), where D = e−γ·dt = 0.9656 is the velocity kept after one 0.1 s step and G = (1 − D)/γ = 9.83 cm per m/s is how far a step carries a unit of speed (lesson 14 wrote these d and g; here d is a distance and g is gravity). A wall acts inside the step: a ball that would end it beyond the wall comes back with v′ = −e·D·v, e = 0.9. So the next velocity, as a function of the distance d to the wall ahead, is a step with its edge at d* = G·v. At v = 2 m/s the edge is 19.7 cm from the wall and the step is h = 3.67 m/s tall, from 1.93 m/s toward the wall to 1.74 m/s away from it.

Train a smooth model on it: a network like lesson 6's, two hidden layers of 12 tanh units, trained with Adam for 120 epochs to predict the next position and velocity. Its log has lesson 6's size, 60 launches of 93 steps, but each launch gets one random nudge (any direction, up to 3 m/s) in place of lesson 6's goal-aimed operator: 5,580 one-step transitions per axis, only 27 of them (0.48 %) bounces. On 27,900 held-out steps, the RMS error of the next velocity by distance to the wall ahead (balls faster than 0.3 m/s; one network, seed 3):

distance to the wall0–10 cm10–2020–3030–5050–100beyond 1 m
RMS error (m/s)0.5510.6780.6450.4530.0390.0075

The 2.7 % of steps that start within half a metre of the wall ahead carry 96 % of the squared error (95 to 97 % over five network seeds), and the error there is 75 times the error beyond a metre (48 to 75 times; this network is the top of that range). Lesson 14 scored a random-feature model by wall steps against the rest, this is a tanh network scored by distance to the wall: two readings of one failure, not one number. Rolled forward, a ball 1 m from the wall at 2 m/s bounces, and ten steps later it is 62 cm from the wall. In four smooth networks it is 10 to 19 cm from it: the ball slowed at the wall instead of bouncing off it, and left slowly.

2 · A smooth model must blur a step

This is not a training accident. A network with finite weights has a bounded slope, |∂v′/∂d| ≤ L. To fall by h it needs a run of at least w = h/L, and the best fit to a step under that limit is a ramp of width w centred on the edge, whose error ε is a sawtooth of height ±h/2 over that width. Its squared error integrated over distance is

∫ ε(d)² dd = h²·w/12

The widget measures the ramp as the distance over which the prediction travels from 10 % to 90 % of the way to the bounce value. At 2 m/s it is 22.8 cm (21 to 24 cm over five network seeds, 19 to 40 cm over five logs), about as wide as the edge is far from the wall. A straight ramp with that 10 to 90 % width has w = 28.5 cm and would cost 0.32 (m/s)²·m; measured, with a ramp that is not straight, the blur costs 0.46.

Road not taken · make the network better
The obvious repair is more of everything. Same task, measured at 2 m/s:
variantramp widthRMS error within 0.5 m
the network above (120 epochs)22.8 cm0.561
six times the epochs (720)16.2 cm0.434
hidden layers of 32 units21.2 cm0.557
four times the launches (60 epochs)17.4 cm0.414
bounce steps repeated twenty times12.0 cm0.679
Longer training, more launches and oversampling narrow the ramp, width does not (21.2 cm is inside the 21 to 24 cm of five seeds of the first network), and none removes it. The narrowest makes the error near the wall worse: its 12.0 cm ramp is centred 6.2 cm past the true edge, the first network's 4.0 cm, and a step in the wrong place costs more than a gentle ramp.

3 · Modes, laws and a switch

If the jump cannot live inside a smooth function, take it out. Lesson 10 met the same fold from the other side: a jump of many steps skips bounces, and a count of them named the piece of the map. Give every transition a mode: free, upper wall, lower wall. In the log a step is a wall step when its velocity change is more than friction explains, and the side of the push says which wall (these labels are exact; a robot would take them from a contact sensor). Inside a mode the law is plain, so fit one linear law per mode by least squares. Then learn only the switch: a logistic regression of each wall mode on (1, p, v), with a ridge of 10−5 because the classes are separable, so the mode is whichever score is positive. The true edge, p = pwall − G·v, is a straight line in the (p, v) plane, so a linear score can hold it.

The fitted laws are the physics, because the log has no noise: the free law keeps 0.9656 of the velocity, the true D, and the wall law returns 0.869 of it reversed, the true e·D. The switch puts the edge 3.0 mm from the true one at 2 m/s (1.5 to 11.5 mm over five logs) and names the wrong mode in 10 of the 27,900 held-out steps (5 to 15 over five logs). The RMS error over all steps is 0.0213 m/s against 0.0939 for the smooth network, and on each of five logs the hybrid is the lower, 0.020 to 0.060 against 0.075 to 0.102.

The two prices can be compared. A ramp costs h²·w/12; a switch placed δ from the edge costs h²·δ, because the whole jump is wrong over a length δ. The hybrid wins when δ < w/12, here 23.8 mm against 3.0 mm. Integrated over the approach at 2 m/s its squared error is 0.0404, against 0.0398 by the formula. It is not exact, and it stands on that one fit: its few mistakes cost a whole jump each, and a heavier ridge (10−3) moves the switch 11 mm and puts 41 steps in the wrong mode.

Road not taken · soft contact
Replace the rigid wall by a stiff spring, F = k·max(0, penetration), and the jump becomes a hinge a network can fit; lesson 13 took this road for the contact between balls. The price is time: for an assumed k = 105 N/m and a 1 kg ball the contact lasts π√(m/k) = 9.9 ms, so the integrator needs a step near 1 ms, 100 inner steps for each 0.1 s step of the Courtyard. And the mode is still there: "in contact or not" is a comparison of a depth with zero.

4 · Stick and slip: a switch on a force

A wall is a mode that happens at a place; friction is a mode that happens at a force. A block of mass m rests on a table that pushes up with N = m·g (g = 9.81 m/s²), and a finger pushes it with a force F. Coulomb's rule gives the friction f, with a static coefficient μs = 0.5 above the kinetic one μk = 0.3:

modewhenwhat happens
stuckv = 0 and F ≤ μsNf = F; nothing moves
slippingv > 0, or v = 0 and F > μsNf = μkN against the motion, so m·a = F − μkN; stuck again when v would reverse

Push for T = 0.3 s (six steps of 0.05 s) and release. Below μsN nothing happens. Above it the block accelerates at a = (F − μkN)/m, then slides to rest, and it travels ½aT² + (aT)²/(2μkg). At F = μsN the acceleration is (μs − μk)g = 1.96 m/s² whatever the mass, so the final distance jumps from 0 to 14.7 cm for both blocks. The jump is at 4.905 N for block A (1 kg) and 9.81 N for block B (2 kg), and the blocks look alike.

The mode is part of the state and is not in the position: at 4 N block A at rest stays put, while the same block moving at 1 mm/s speeds up to 0.0539 m/s in one step. The hybrid block model has the wall model's three parts: the mode (read from v; zero means stuck), a switch that lets a stuck block slip when a logistic function of the push is at least 0.9, and one linear law for slipping, Δv = a0 + a1v + a2q. The push is q = F/N when the model is given N (section 5) and the raw F when it is not. Fitted with N on 2,400 steps of pushing, it recovers μk = 0.300, and its switch is 50 % sure at μs = 0.496 (0.496 to 0.510 over five logs).

5 · What a camera cannot see

The switch needs F/N, and N is not in a picture, which is lesson 3's point again: a picture need not keep what decides. A model that sees only the speed, from the camera, and the push, its own command, gets the same input for both blocks and must give the same answer, although one lets go at 4.905 N and the other at 9.81 N. Nor does the motion give N away in time: a block that has not moved shows nothing, whatever it weighs. Once it moves, the toy camera reads position every Δt = 50 ms with an error σ = 2 mm, so a velocity from two readings is off by √2·σ/Δt = 0.057 m/s and an acceleration from three readings by √6·σ/Δt² = 1.96 m/s². That equals the friction drop (μs − μk)g it would have to detect; as a force, m·a, it is 1.96 N for A and 3.92 N for B, 40 % of either block's breakaway force. More readings average it down, but only after the first push has been decided. An ideal touch channel, a load cell that reports N exactly, has none of this trouble, and it gives something to predict: the friction force a sensor would read. One step ahead, on a held-out log, the hybrid with touch gets it to 0.04 to 0.12 N RMS over five logs, the smooth network with touch to 0.34 to 0.43 N over eight seeds, the camera-only models to 0.6 to 0.7 N.

Cross the two repairs. Each cell is the mean error of the final distance over a sweep of pushes from 0 to 12 N, both blocks, for the widget's networks:

camera onlycamera + touch
smooth network22.1 cm4.59 cm
hybrid24.8 cm0.11 cm

The widget's smooth network is the largest of its eight seeds in both columns, and its hybrid with touch the smallest of five logs. Over the eight seeds the smooth cells are 21 to 22 cm and 1.8 to 4.6 cm; over the five logs the hybrid cells are 24 to 25 cm and 0.11 to 0.44 cm. Touch removes an information floor (the smooth network's error falls 4.8 to 12 times) and the hybrid removes the blur floor (another 4.1 to 41 times, over every pairing of seed and log).

Touch alone helps; the hybrid alone does not, because without N the floor is information, not blur. One law and one switch must serve two blocks that differ in mass: a push of 12 N sends A 1.66 m and B 0.28 m, and a camera-only model must say the same for both. Its switch is 50 % sure at 9.4 N, between their thresholds, and it lets a block slip only when 90 % sure, from 12.0 N up (11.8 to 13.6 N over five logs); at 50 % its cell would read 21.1 cm instead of 24.8. On a first push above 3 N (144 pushes of a held-out log) it chooses between slipping and staying rightly 53 % of the time, against 99.3 % with touch. The smooth network with touch has a signature failure, creep: its blurred switch lets a block that cannot move slide. At 4.6 N, below block A's threshold, the eight seeds move A by 4.2 to 10.2 cm (the widget's, the most), and by more than a centimetre already at 1.8 to 4.0 N (the widget's, 1.8); the truth and the hybrid say 0.

The widget

A wall, two blocks and a sentence
Left: the next velocity of a ball driving at the wall against its distance from the wall (dots are training steps, red where the wall pushed). Middle and right: where block A (1 kg) and block B (2 kg) stop after a push of force F for 0.3 s; the cyan line is your slider. Bottom: the chosen block's position, and the friction force a force sensor would read. Black is the truth, purple the smooth network, amber the hybrid. "Camera only" gives the block models the speed and the push; "camera + touch" adds the normal force N. Train adds 120 epochs to the wall network. Plan and the nudge test belong to sections 6 and 7; the nudge test always uses the first 120 epochs of the first network seed.
bounce edge d*
—
smooth ramp width
—
smooth error within 0.5 m
—
epochs trained
—
hybrid edge error
—
hybrid wrong mode
—
block ends: truth
—
block ends: smooth
—
block ends: hybrid
—
sweep error, smooth
—
sweep error, hybrid
—
hybrid's switch
—
plans on the mark, smooth
—
plans on the mark, hybrid
—
nudge test, friction only
—
nudge test, smooth net (seed 3)
—
nudge test, hybrid
—
Show the core JS
L15.blockStep = function (m, x, v, F) {
  var N = m * G0, Fs = MUS * N, Fk = MUK * N, a;
  if (v === 0) {
    if (F <= Fs) return [x, 0, F];  // stuck: the table holds
    a = (F - Fk) / m; return [x + 0.5 * a * DTB * DTB, a * DTB, Fk];  // breakaway
  }
  ...  // moving: friction mu_k N against the motion, and it stops inside the step if v would reverse
};
...
L15.WallHybrid.prototype.mode = function (p, v) {
  var a = L15.AX[this.ax], z0 = (p - a.mid) / a.half, z1 = v / 3, u = this.wUp, d = this.wDn;
  if (u[0] + u[1] * z0 + u[2] * z1 > 0) return 1;  // the switch: a positive score means the upper wall
  return d[0] + d[1] * z0 + d[2] * z1 > 0 ? -1 : 0;
};
L15.WallHybrid.prototype.pred = function (p, v) {
  var L = this.law[this.mode(p, v)] || this.law[0];  // the law of the chosen mode
  return [L[0] + L[2] * p + L[4] * v, L[1] + L[3] * p + L[5] * v];
};

What to try. Leave the start state (7 N, block A, camera only). The edge is 19.7 cm from the wall, the smooth ramp 22.8 cm wide and the hybrid's edge 3.0 mm off. Block A truly ends 43.4 cm away. The camera-only smooth model says 8.7 cm and the hybrid 0.0 cm, and each says the same for block B, which truly stays put (switch tile: 9.4 N at 50 %, 12.0 N at 90 %). Choose camera + touch: the hybrid now says 43.2 cm for A and 0.0 cm for B, the smooth network 37.4 and 2.2. Slide the push to 4.6 N, below A's threshold: the truth is 0, the hybrid 0, the smooth network 10.2 cm. Set the ball to 1 and to 3 m/s: the edge moves to 9.8 and 29.5 cm, the hybrid follows to within 2.2 and 3.7 mm, and the smooth ramp is 18 and 38 cm wide. With the ball back at 2 m/s, press train: one press gives a ramp of 20.2 cm, five presses 16.2 cm.

6 · Words are costs

A goal arrives as a sentence. A model of consequences answers "what happens if I do this?" and a sentence is not an action, so in this model it enters elsewhere: a planner minimises a cost, and a cost is a function of what the model predicts, so the sentence is compiled into one. "To the mark" is the miss in units of 5 cm. "Gently" adds 2F/Fmax, a price on force; "briefly" adds 2Tp/Tmax, a price on the duration of the push (Fmax = 12 N, Tmax = 10 steps). The planner is lesson 9's: CEM over the push (F, Tp), 80 candidates for 8 iterations, each scored by rolling the model 60 steps, which is 38,400 model steps per plan. The plan then runs in the true world and counts if the block stops within 5 cm of the mark. Three marks (15, 25 and 35 cm) and two blocks make six plans per model: choose "camera + touch" and press plan under each phrase.

Plans that land, out of six, with touch: the hybrid lands 6 under every phrase and every one of five planner seeds; the smooth network lands 3 to 6 for plain "to the mark", 4 to 6 for "briefly" and 0 to 3 for "gently" over eight seeds, and the widget's, seed 3, lands 3, 6 and 0. A price on force sends the planner to the smallest push that still moves the block, which is the threshold, which is where a blurred switch is most wrong. For the 25 cm mark the hybrid's gentle plans sit 0.84 N (block A) and 0.68 N (block B) above μsN; for the 15 cm mark the widget's smooth model plans 4.46 N for block A, below its threshold, imagines the block reaching the mark, and the real block does not move. Without touch no model lands more than 1 of six. This is lesson 8's optimiser's curse with a sentence in it: a cost that says "gently" makes the planner look for the place where the model is most wrong.

Real systems put words in at two seats. A policy maps (observation, words) to an action: RT-2 (2023), which coined "vision-language-action", OpenVLA (2024) and π0 (2024), a flow-matching action model on a pretrained VLM; none of the three abstracts calls its model a world model. A world model maps (state, action) to a consequence, and words reach it as a cost over its predictions, as above; as goal images, as in V-JEPA 2-AC (2025), which plans with CEM to minimise the L1 distance between imagined and goal-image representations, 16 s per action where a Cosmos baseline with a tenth of the samples took 4 minutes; or as an input that conditions it, as in 1X (2026), which conditions a 14 B-parameter video model on text and uses it as the policy, an inverse dynamics model (lesson 14) turning the generated video into actions. A world-action model (WAM) does both: DreamZero (2026), "Unlike VLAs, WAMs learn physical dynamics by predicting future world states and actions", and Cosmos Policy (2026), which makes actions, future states and values latent frames of one diffusion process, so planning is possible at test time. Hogan and Rodriguez (2016) controlled a modelled pusher–slider by hybrid MPC over its contact modes, and Hogan, Grau and Rodriguez (2018) learned the mode sequences offline: modes apart from laws, for control, not prediction.

7 · Scored by one step, judged by decisions

The model now has a body, and nothing yet says how to know it is good. Test it twice on lesson 1's nudge task, with the true state given: the ball is launched, one impulse is allowed at step 12, and the ball must be in the goal 80 steps later. The agent imagines 128 candidate nudges, keeps the one whose imagined ending is nearest the goal, and executes it in the true world. Sixty launches, with a model for each axis (the widget's last button):

modelone-step error, both axes (m/s)nudge successmiss, imagined and real (m)picks that bounce
hybrid0.0152100 %0.23, 0.2392 %
friction only (no walls)0.18341.7 %0.57, 0.570 %
smooth networks, six seeds0.121 to 0.1350 to 10 %0.77, 1.53 (means)5 to 53 %

Rank the models by held-out one-step error and every smooth network beats the model that knows nothing about walls. Rank them by what they decide and every smooth network loses to it, by at least 30 points; on other launch sets the friction-only model scores 35 to 50 % and the widget's smooth network, the first of the six and the best of them on both counts, at most 8 %. The friction-only model is exact in free flight and its picks never touch a wall, so what it imagines is what happens. A smooth network is a little wrong everywhere and very wrong near walls, and a planner that takes the best of 128 imagined nudges selects for those errors: the smooth networks imagine landing 0.77 m from the goal and land 1.53 m away.

What this lesson did not do
The wall is one axis of the Courtyard (its walls are axis-aligned, so the axes do not interact); the block on its table is a separate one-dimensional toy. Touch is an ideal load cell that reports N exactly, the wall's mode labels come exactly from the log, the block models are handed the exact speed, and no log has noise. The switch is a linear score and the laws inside the modes are linear because the toy's edge and laws are; real contact has curved edges, curved laws and many modes. Real touch is noisy and costly (Robot Model Training · 17 prices the data), and learning modes without labels is not attempted. The words were three hand-written phrases. Two exams were run, one step and one decision, without saying which to trust or how many trials a decision exam needs: that is lesson 16.

Common mistakes / failure modes

"a bigger network, or a longer run, will learn the bounce"
Six times the epochs narrows the ramp only from 22.8 to 16.2 cm (§2).
"the hybrid model is exact"
It names the wrong mode in 10 of 27,900 steps, and each mistake costs a whole jump (§3).
"a hybrid model, or touch, is enough alone"
Hybrid without touch: 24 to 25 cm of error; smooth with touch: 1.8 to 4.6 cm; both: 0.11 to 0.44 cm (§5).
"a cost that says gently is a safe cost"
It sends the planner to the smallest push that still moves the block, the threshold, where a blurred switch is most wrong: over eight seeds the smooth network lands 0 to 3 of six plans (§6).
"the lower the one-step error, the better the model"
Smooth networks score 0.121 to 0.135 against 0.183 for a model without walls, and decide worse (§7).

Checkpoint exercise

Try it
A block of mass 3 kg has μs = 0.4 and μk = 0.25. (a) What push breaks it away, and what friction acts while it slides? (b) How much does the friction force drop at breakaway? (c) A camera has the acceleration error of section 5, 1.96 m/s². What is the error of its estimate of the force F = m·a, and how does it compare with the drop in (b)? Answer: (a) μsmg = 11.8 N, then μkmg = 7.36 N. (b) 4.41 N. (c) 3 × 1.96 = 5.88 N, which is 1.33 times the drop, so one camera estimate cannot see the transition; a load cell that reads N turns it into the comparison of F/N with 0.4.

Where this points next

A model can now represent a body: a mode, a law inside it, a switch on a force that a touch channel supplies, and a cost compiled from a sentence. And two exams now disagree. By the usual score, the held-out one-step error, the smooth networks beat a model that has never met a wall, 0.121 to 0.135 m/s against 0.183, and by decisions they lose to it, 0 to 10 % success against 41.7 %. How do we know that a particular model passes the exam that matters, a better decision in the world, and not only the ones that look good?

Takeaway
A bounce is a step in the next-velocity map, with its edge at d* = G·v. A smooth network replaces it by a ramp about as wide as the edge is far from the wall, which costs h²·w/12 for a ramp of width w; training longer or on more data narrows the ramp, a wider network does not, and nothing removes it. A hybrid model keeps a mode, a linear law per mode and a switch fitted to contact labels, and wins when the switch is within w/12 of the edge (3.0 mm against 23.8 mm). Friction is a switch on a force, F against μsN, and a camera cannot supply the threshold, since N is not in a picture, nor read the transition from one step's motion, since the friction drop is as large as the acceleration noise. A sentence is a cost over predicted quantities, so "gently" can steer a planner onto the discontinuity. And a model can win the one-step score and lose the decision.

Interview prompts

Companion reads: Robot Model Training · 10 What cameras cannot see: contact and force (the same missing column, for policies), Robot Model Training · 17 What an hour costs: prices and the ledger (what force data costs), Lesson 22 · How the handle gets in (how an action or a sentence enters a model) and Lesson 17 · Two seats, one tuple (the two seats).