Bodies, contact, and words
Lesson 14 ended at a wall: its smooth model, good in free flight, was many times worse at the steps that touch one, even when it had trained on bounces. A body is touched, not only seen, and a touch is a jump. A smooth network trained on bounces predicts a wall that slows the ball gradually, so nearly all its error sits in the steps next to a wall. This lesson shows why a smooth model must blur a jump and replaces it with a hybrid one: a mode, a simple law inside each mode, and a switch. On a table the switch is friction, a force compared with a threshold, and a camera cannot read force. Words then enter as costs over what the model predicts. The lesson ends with a model that scores better on the usual test and decides worse.
New idea: a model of a body is hybrid: a discrete mode chosen by comparing a measured quantity with a threshold, and a smooth law inside each mode. A smooth network can only blur the switch. Once the switch is one comparison, the laws inside the modes are, here, linear; the quantity compared, a force, has to be measured.
Forces next: We now have every piece: state, dynamics, uncertainty, rollouts, causality, practice, planning, abstraction, pixels, memory, objects, actions and a body, each justified by an exam the previous model failed. How do we know that a particular model passes the exam that matters, a better decision in the world, and not only the ones that look good?
1 · The wall is a step
Lesson 14 ended at the walls; look at one wall, on one axis. A free ball obeys p′ = p + G·v and v′ = D·v (primes: one step later), where D = e−γ·dt = 0.9656 is the velocity kept after one 0.1 s step and G = (1 − D)/γ = 9.83 cm per m/s is how far a step carries a unit of speed (lesson 14 wrote these d and g; here d is a distance and g is gravity). A wall acts inside the step: a ball that would end it beyond the wall comes back with v′ = −e·D·v, e = 0.9. So the next velocity, as a function of the distance d to the wall ahead, is a step with its edge at d* = G·v. At v = 2 m/s the edge is 19.7 cm from the wall and the step is h = 3.67 m/s tall, from 1.93 m/s toward the wall to 1.74 m/s away from it.
Train a smooth model on it: a network like lesson 6's, two hidden layers of 12 tanh units, trained with Adam for 120 epochs to predict the next position and velocity. Its log has lesson 6's size, 60 launches of 93 steps, but each launch gets one random nudge (any direction, up to 3 m/s) in place of lesson 6's goal-aimed operator: 5,580 one-step transitions per axis, only 27 of them (0.48 %) bounces. On 27,900 held-out steps, the RMS error of the next velocity by distance to the wall ahead (balls faster than 0.3 m/s; one network, seed 3):
| distance to the wall | 0–10 cm | 10–20 | 20–30 | 30–50 | 50–100 | beyond 1 m |
|---|---|---|---|---|---|---|
| RMS error (m/s) | 0.551 | 0.678 | 0.645 | 0.453 | 0.039 | 0.0075 |
The 2.7 % of steps that start within half a metre of the wall ahead carry 96 % of the squared error (95 to 97 % over five network seeds), and the error there is 75 times the error beyond a metre (48 to 75 times; this network is the top of that range). Lesson 14 scored a random-feature model by wall steps against the rest, this is a tanh network scored by distance to the wall: two readings of one failure, not one number. Rolled forward, a ball 1 m from the wall at 2 m/s bounces, and ten steps later it is 62 cm from the wall. In four smooth networks it is 10 to 19 cm from it: the ball slowed at the wall instead of bouncing off it, and left slowly.
2 · A smooth model must blur a step
This is not a training accident. A network with finite weights has a bounded slope, |∂v′/∂d| ≤ L. To fall by h it needs a run of at least w = h/L, and the best fit to a step under that limit is a ramp of width w centred on the edge, whose error ε is a sawtooth of height ±h/2 over that width. Its squared error integrated over distance is
∫ ε(d)² dd = h²·w/12
The widget measures the ramp as the distance over which the prediction travels from 10 % to 90 % of the way to the bounce value. At 2 m/s it is 22.8 cm (21 to 24 cm over five network seeds, 19 to 40 cm over five logs), about as wide as the edge is far from the wall. A straight ramp with that 10 to 90 % width has w = 28.5 cm and would cost 0.32 (m/s)²·m; measured, with a ramp that is not straight, the blur costs 0.46.
| variant | ramp width | RMS error within 0.5 m |
|---|---|---|
| the network above (120 epochs) | 22.8 cm | 0.561 |
| six times the epochs (720) | 16.2 cm | 0.434 |
| hidden layers of 32 units | 21.2 cm | 0.557 |
| four times the launches (60 epochs) | 17.4 cm | 0.414 |
| bounce steps repeated twenty times | 12.0 cm | 0.679 |
3 · Modes, laws and a switch
If the jump cannot live inside a smooth function, take it out. Lesson 10 met the same fold from the other side: a jump of many steps skips bounces, and a count of them named the piece of the map. Give every transition a mode: free, upper wall, lower wall. In the log a step is a wall step when its velocity change is more than friction explains, and the side of the push says which wall (these labels are exact; a robot would take them from a contact sensor). Inside a mode the law is plain, so fit one linear law per mode by least squares. Then learn only the switch: a logistic regression of each wall mode on (1, p, v), with a ridge of 10−5 because the classes are separable, so the mode is whichever score is positive. The true edge, p = pwall − G·v, is a straight line in the (p, v) plane, so a linear score can hold it.
The fitted laws are the physics, because the log has no noise: the free law keeps 0.9656 of the velocity, the true D, and the wall law returns 0.869 of it reversed, the true e·D. The switch puts the edge 3.0 mm from the true one at 2 m/s (1.5 to 11.5 mm over five logs) and names the wrong mode in 10 of the 27,900 held-out steps (5 to 15 over five logs). The RMS error over all steps is 0.0213 m/s against 0.0939 for the smooth network, and on each of five logs the hybrid is the lower, 0.020 to 0.060 against 0.075 to 0.102.
The two prices can be compared. A ramp costs h²·w/12; a switch placed δ from the edge costs h²·δ, because the whole jump is wrong over a length δ. The hybrid wins when δ < w/12, here 23.8 mm against 3.0 mm. Integrated over the approach at 2 m/s its squared error is 0.0404, against 0.0398 by the formula. It is not exact, and it stands on that one fit: its few mistakes cost a whole jump each, and a heavier ridge (10−3) moves the switch 11 mm and puts 41 steps in the wrong mode.
4 · Stick and slip: a switch on a force
A wall is a mode that happens at a place; friction is a mode that happens at a force. A block of mass m rests on a table that pushes up with N = m·g (g = 9.81 m/s²), and a finger pushes it with a force F. Coulomb's rule gives the friction f, with a static coefficient μs = 0.5 above the kinetic one μk = 0.3:
| mode | when | what happens |
|---|---|---|
| stuck | v = 0 and F ≤ μsN | f = F; nothing moves |
| slipping | v > 0, or v = 0 and F > μsN | f = μkN against the motion, so m·a = F − μkN; stuck again when v would reverse |
Push for T = 0.3 s (six steps of 0.05 s) and release. Below μsN nothing happens. Above it the block accelerates at a = (F − μkN)/m, then slides to rest, and it travels ½aT² + (aT)²/(2μkg). At F = μsN the acceleration is (μs − μk)g = 1.96 m/s² whatever the mass, so the final distance jumps from 0 to 14.7 cm for both blocks. The jump is at 4.905 N for block A (1 kg) and 9.81 N for block B (2 kg), and the blocks look alike.
The mode is part of the state and is not in the position: at 4 N block A at rest stays put, while the same block moving at 1 mm/s speeds up to 0.0539 m/s in one step. The hybrid block model has the wall model's three parts: the mode (read from v; zero means stuck), a switch that lets a stuck block slip when a logistic function of the push is at least 0.9, and one linear law for slipping, Δv = a0 + a1v + a2q. The push is q = F/N when the model is given N (section 5) and the raw F when it is not. Fitted with N on 2,400 steps of pushing, it recovers μk = 0.300, and its switch is 50 % sure at μs = 0.496 (0.496 to 0.510 over five logs).
5 · What a camera cannot see
The switch needs F/N, and N is not in a picture, which is lesson 3's point again: a picture need not keep what decides. A model that sees only the speed, from the camera, and the push, its own command, gets the same input for both blocks and must give the same answer, although one lets go at 4.905 N and the other at 9.81 N. Nor does the motion give N away in time: a block that has not moved shows nothing, whatever it weighs. Once it moves, the toy camera reads position every Δt = 50 ms with an error σ = 2 mm, so a velocity from two readings is off by √2·σ/Δt = 0.057 m/s and an acceleration from three readings by √6·σ/Δt² = 1.96 m/s². That equals the friction drop (μs − μk)g it would have to detect; as a force, m·a, it is 1.96 N for A and 3.92 N for B, 40 % of either block's breakaway force. More readings average it down, but only after the first push has been decided. An ideal touch channel, a load cell that reports N exactly, has none of this trouble, and it gives something to predict: the friction force a sensor would read. One step ahead, on a held-out log, the hybrid with touch gets it to 0.04 to 0.12 N RMS over five logs, the smooth network with touch to 0.34 to 0.43 N over eight seeds, the camera-only models to 0.6 to 0.7 N.
Cross the two repairs. Each cell is the mean error of the final distance over a sweep of pushes from 0 to 12 N, both blocks, for the widget's networks:
| camera only | camera + touch | |
|---|---|---|
| smooth network | 22.1 cm | 4.59 cm |
| hybrid | 24.8 cm | 0.11 cm |
The widget's smooth network is the largest of its eight seeds in both columns, and its hybrid with touch the smallest of five logs. Over the eight seeds the smooth cells are 21 to 22 cm and 1.8 to 4.6 cm; over the five logs the hybrid cells are 24 to 25 cm and 0.11 to 0.44 cm. Touch removes an information floor (the smooth network's error falls 4.8 to 12 times) and the hybrid removes the blur floor (another 4.1 to 41 times, over every pairing of seed and log).
Touch alone helps; the hybrid alone does not, because without N the floor is information, not blur. One law and one switch must serve two blocks that differ in mass: a push of 12 N sends A 1.66 m and B 0.28 m, and a camera-only model must say the same for both. Its switch is 50 % sure at 9.4 N, between their thresholds, and it lets a block slip only when 90 % sure, from 12.0 N up (11.8 to 13.6 N over five logs); at 50 % its cell would read 21.1 cm instead of 24.8. On a first push above 3 N (144 pushes of a held-out log) it chooses between slipping and staying rightly 53 % of the time, against 99.3 % with touch. The smooth network with touch has a signature failure, creep: its blurred switch lets a block that cannot move slide. At 4.6 N, below block A's threshold, the eight seeds move A by 4.2 to 10.2 cm (the widget's, the most), and by more than a centimetre already at 1.8 to 4.0 N (the widget's, 1.8); the truth and the hybrid say 0.
The widget
What to try. Leave the start state (7 N, block A, camera only). The edge is 19.7 cm from the wall, the smooth ramp 22.8 cm wide and the hybrid's edge 3.0 mm off. Block A truly ends 43.4 cm away. The camera-only smooth model says 8.7 cm and the hybrid 0.0 cm, and each says the same for block B, which truly stays put (switch tile: 9.4 N at 50 %, 12.0 N at 90 %). Choose camera + touch: the hybrid now says 43.2 cm for A and 0.0 cm for B, the smooth network 37.4 and 2.2. Slide the push to 4.6 N, below A's threshold: the truth is 0, the hybrid 0, the smooth network 10.2 cm. Set the ball to 1 and to 3 m/s: the edge moves to 9.8 and 29.5 cm, the hybrid follows to within 2.2 and 3.7 mm, and the smooth ramp is 18 and 38 cm wide. With the ball back at 2 m/s, press train: one press gives a ramp of 20.2 cm, five presses 16.2 cm.
6 · Words are costs
A goal arrives as a sentence. A model of consequences answers "what happens if I do this?" and a sentence is not an action, so in this model it enters elsewhere: a planner minimises a cost, and a cost is a function of what the model predicts, so the sentence is compiled into one. "To the mark" is the miss in units of 5 cm. "Gently" adds 2F/Fmax, a price on force; "briefly" adds 2Tp/Tmax, a price on the duration of the push (Fmax = 12 N, Tmax = 10 steps). The planner is lesson 9's: CEM over the push (F, Tp), 80 candidates for 8 iterations, each scored by rolling the model 60 steps, which is 38,400 model steps per plan. The plan then runs in the true world and counts if the block stops within 5 cm of the mark. Three marks (15, 25 and 35 cm) and two blocks make six plans per model: choose "camera + touch" and press plan under each phrase.
Plans that land, out of six, with touch: the hybrid lands 6 under every phrase and every one of five planner seeds; the smooth network lands 3 to 6 for plain "to the mark", 4 to 6 for "briefly" and 0 to 3 for "gently" over eight seeds, and the widget's, seed 3, lands 3, 6 and 0. A price on force sends the planner to the smallest push that still moves the block, which is the threshold, which is where a blurred switch is most wrong. For the 25 cm mark the hybrid's gentle plans sit 0.84 N (block A) and 0.68 N (block B) above μsN; for the 15 cm mark the widget's smooth model plans 4.46 N for block A, below its threshold, imagines the block reaching the mark, and the real block does not move. Without touch no model lands more than 1 of six. This is lesson 8's optimiser's curse with a sentence in it: a cost that says "gently" makes the planner look for the place where the model is most wrong.
Real systems put words in at two seats. A policy maps (observation, words) to an action: RT-2 (2023), which coined "vision-language-action", OpenVLA (2024) and π0 (2024), a flow-matching action model on a pretrained VLM; none of the three abstracts calls its model a world model. A world model maps (state, action) to a consequence, and words reach it as a cost over its predictions, as above; as goal images, as in V-JEPA 2-AC (2025), which plans with CEM to minimise the L1 distance between imagined and goal-image representations, 16 s per action where a Cosmos baseline with a tenth of the samples took 4 minutes; or as an input that conditions it, as in 1X (2026), which conditions a 14 B-parameter video model on text and uses it as the policy, an inverse dynamics model (lesson 14) turning the generated video into actions. A world-action model (WAM) does both: DreamZero (2026), "Unlike VLAs, WAMs learn physical dynamics by predicting future world states and actions", and Cosmos Policy (2026), which makes actions, future states and values latent frames of one diffusion process, so planning is possible at test time. Hogan and Rodriguez (2016) controlled a modelled pusher–slider by hybrid MPC over its contact modes, and Hogan, Grau and Rodriguez (2018) learned the mode sequences offline: modes apart from laws, for control, not prediction.
7 · Scored by one step, judged by decisions
The model now has a body, and nothing yet says how to know it is good. Test it twice on lesson 1's nudge task, with the true state given: the ball is launched, one impulse is allowed at step 12, and the ball must be in the goal 80 steps later. The agent imagines 128 candidate nudges, keeps the one whose imagined ending is nearest the goal, and executes it in the true world. Sixty launches, with a model for each axis (the widget's last button):
| model | one-step error, both axes (m/s) | nudge success | miss, imagined and real (m) | picks that bounce |
|---|---|---|---|---|
| hybrid | 0.0152 | 100 % | 0.23, 0.23 | 92 % |
| friction only (no walls) | 0.183 | 41.7 % | 0.57, 0.57 | 0 % |
| smooth networks, six seeds | 0.121 to 0.135 | 0 to 10 % | 0.77, 1.53 (means) | 5 to 53 % |
Rank the models by held-out one-step error and every smooth network beats the model that knows nothing about walls. Rank them by what they decide and every smooth network loses to it, by at least 30 points; on other launch sets the friction-only model scores 35 to 50 % and the widget's smooth network, the first of the six and the best of them on both counts, at most 8 %. The friction-only model is exact in free flight and its picks never touch a wall, so what it imagines is what happens. A smooth network is a little wrong everywhere and very wrong near walls, and a planner that takes the best of 128 imagined nudges selects for those errors: the smooth networks imagine landing 0.77 m from the goal and land 1.53 m away.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
A model can now represent a body: a mode, a law inside it, a switch on a force that a touch channel supplies, and a cost compiled from a sentence. And two exams now disagree. By the usual score, the held-out one-step error, the smooth networks beat a model that has never met a wall, 0.121 to 0.135 m/s against 0.183, and by decisions they lose to it, 0 to 10 % success against 41.7 %. How do we know that a particular model passes the exam that matters, a better decision in the world, and not only the ones that look good?
Interview prompts
- Why does a smooth network trained on bounces predict a soft wall? (§2 — a bounded slope turns a jump of height h into a ramp of width at least h/L, costing h²w/12.)
- What is a hybrid dynamics model, and when does it beat a smooth one? (§3 — a mode, one law per mode and a switch; it wins when the switch is within w/12 of the edge, since a misplaced switch costs h²δ.)
- How far does a block pushed with force F for time T travel? (§4 — nothing if F ≤ μsN; otherwise ½aT² + (aT)²/(2μkg), a = (F − μkN)/m.)
- Why can a camera not supply the switch variable of friction? (§5 — N is not in the picture, a block that has not moved shows nothing, and the acceleration noise is as large as the friction drop.)
- How does "gently" reach a world model, and what goes wrong? (§6 — as a cost over predicted quantities; it pulls the planner onto the discontinuity.)
- Can a model with a lower one-step error decide worse? (§7 — yes: a planner that takes the best of many imagined plans selects for the model's errors.)
Companion reads: Robot Model Training · 10 What cameras cannot see: contact and force (the same missing column, for policies), Robot Model Training · 17 What an hour costs: prices and the ledger (what force data costs), Lesson 22 · How the handle gets in (how an action or a sentence enters a model) and Lesson 17 · Two seats, one tuple (the two seats).