all_lessons/World Models/07 · What actions causelesson 7 / 31

What does my action cause?

Every model so far was fitted to a log of someone acting. Give that someone a reason the model cannot see, a gust of wind they can feel and push against, and the usual regression learns that a nudge of 1 m/s moves the ball by 0.13 m, where the simulator says 1.44 m. The log is not wrong; it answers a different question. This lesson derives the gap exactly, as the share of each nudge that the wind did not choose, cuts the link between wind and nudge three ways (randomise, measure, infer), tests a model against the question it will actually be asked, and then finds that even an unconfounded model speaks only where the nudges went.

The thesis, here
A log records what happened after an action was chosen, and the choice carries information about the world. A regression on it learns what one sees: the outcome among the episodes in which the nudge happened to be a. A model of consequences must answer what one gets when the nudge is set by fiat, whatever the wind is doing. The two agree only for the part of the action that nothing else decided, which is 9 % of the careful operator's nudge variance and none of a perfect operator's.
Linear position
Forced by: A rollout has a trustworthy horizon that we can now estimate, and re-observing resets it. But every model so far was fitted to logs that someone generated by acting. If whoever acted reacted to something the model cannot see, such as a gust of wind, the logs confound what an action did with why it was taken. Can a model fitted to logs tell what an action causes?
New idea: a regression on a log learns what an action causes only from the part of the action that nothing else chose. To learn the cause, make that part larger by randomising the action, or by measuring or inferring what drove it, and trust the answer only where the actions went.
Forces next: A model fitted to varied, randomized actions can say what an action causes, at least where the data went. The cheapest use of it is practice: let a policy improve on the model's imagined rollouts instead of in the world. But a policy trained against a model finds the model's flattering mistakes and climbs them. How do we learn from imagined experience without being fooled by it?
The plan
Six moves. (1) Build the log of an operator who feels a hidden wind, and fit the model every earlier lesson would fit. (2) Derive what that model learns, exactly. (3) Cut the arrow from wind to nudge three ways. (4) Test a model on the question it will be asked. (5) Ask where even a clean log is silent. (6) Spend the next nudges where an ensemble disagrees.

1 · The log, and the model that learns from it

The world is the Courtyard, seen along its x axis. A launcher fires the ball, the agent watches it for k = 4 steps (0.4 s), adds one nudge to its velocity along x, and the log records how far the ball moves in the next 2 s. A hidden wind blows along x, constant within an episode and different between episodes: an acceleration drawn from N(0, 0.3²) m/s², the prior that lesson 2's wind state started from. Every step also adds a random kick of 3 cm/s to the velocity. The log keeps the ball's velocity exactly (lessons 2 to 4 showed how an agent obtains it) and does not keep the wind.

symbolmeaningunit
vthe ball's velocity along x when the nudge arrives: the state the model seesm/s
athe nudge, an impulse added to vm/s
wthe hidden wind, a constant accelerationm/s²
yhow far the ball moves along x in the 2 s after the nudge: what the model predictsm

Friction removes velocity at rate γ = 0.35 s−1 and the wind adds to it, so dv/dt = −γv + w. Integrating over τ = 2 s, with the nudge added to v at the start, gives

y = B·(v + a) + C·w + kicks

where B = (1 − e−γτ)/γ is how far a unit nudge carries the ball and C how far a unit wind pushes it. Reading both off the simulator (rerun one episode with a unit nudge, then with a unit wind) gives B = 1.44 m per m/s and C = 1.62 m per m/s². The friction integral gives the same B and a C of 1.60; the difference is the engine's five sub-steps per step. Where the ball is does not enter until it meets a wall (§5), so the state the model needs is its velocity.

The log is made by an operator who feels the wind and pushes against it. The push that cancels the wind's effect on average is a = −κw with κ = C/B = 1.13 s, and the operator adds a habitual noise ε of standard deviation σε = 0.10 m/s: a = −κw + ε. The log holds 1000 such episodes.

Now fit what lessons 4 to 6 fitted, least squares of the outcome on what the model sees, y ≈ θ0 + θv·v + β̂·a, and ask it what a nudge does. It answers β̂ = 0.13 m per m/s, 9 % of the truth. By the usual exam it is a good model: on fresh episodes from the same operator its error is 0.18 m, under a fifth of a metre. The model has learned that pushing the ball barely moves it.

2 · Seeing is not doing

wind w ↙ ↘ nudge a ──→ outcome y

The wind moves the outcome directly (C) and moves the nudge, because the operator reacts to it. The nudge moves the outcome (B). A regression of y on a cannot separate the arrow from a to y from the detour through w. It reports what one sees, E[y | a], the average outcome among logged episodes whose nudge was a. A model of consequences needs what one gets, E[y | do(a)], the average when the nudge is set by fiat and the arrow from w to a is cut. Here that is B(v + a), because the wind averages to zero. Among the logged episodes with nudge a it does not.

Let a⊥ and y⊥ be the nudge and the outcome after the model's other inputs have been regressed out of them. By the Frisch–Waugh theorem the coefficient of a in the full regression is the slope of y⊥ on a⊥. The outcome obeys y⊥ = B·a⊥ + C·w⊥ + noise, where w⊥ is the part of the wind the other inputs leave unexplained, of variance V. So

β̂ = Cov(a⊥, y⊥) / Var(a⊥) = B + C·Cov(a⊥, w⊥) / Var(a⊥)

With a = −κw + ε and ε independent of everything else, Cov(a⊥, w⊥) = −κV and Var(a⊥) = κ²V + σε², so, exactly (walls aside),

β̂ = B − C·κ·V / (κ²V + σε²), and for the operator who cancels the wind (κ = C/B): β̂ / B = σε² / (σε² + κ²V)

Read β̂/B as a share: the fraction of the nudge's variance that the wind did not choose. The ball's velocity at 0.4 s says little about the wind, so V is 86 % of σw² and the share at σε = 0.10 m/s is 9 %. For a perfect operator (σε = 0) the share is 0: the log says pushing does nothing, because every nudge was chosen to cancel exactly the wind that came with it. A nudge of +0.34 m/s, one standard deviation of the operator's spread, came with a wind of −0.30 m/s². It pushed the ball +0.49 m, the wind pushed it −0.49 m, and the log recorded 0.

σε (m/s)closed formfit, 1000 episodesfit, 200 000 episodes
00.000−0.0180.000
0.100.1360.1330.136
0.350.8070.8120.807
1.01.3131.2771.275

The table is the slope in m per m/s the state-only model learns, against the closed form with the measured V. They agree to the sampling error. At σε = 1 the fit falls a little below the formula because the largest nudges send the ball into a wall, in 1.8 % of the episodes, and where it bounces the response stops being a straight line; with the walls taken away the same log gives 1.31.

Road not taken · more data, a bigger model
When a model is wrong, the instinct of earlier lessons is to feed it more or give it more capacity. At 200 000 episodes, 200 times more, the slope at σε = 0.10 is 0.136 with a standard error of 0.001: the error bar shrinks around the wrong value, so the gap is bias, not noise. A tanh network with two layers of 16 units, fitted to the same log, has a see-error of 0.21 m and a do-error of 0.94 m, and says that +1 m/s moves the ball 0.11 m. A flexible model fits the relation in the log, and the relation in the log is the confounded one.

3 · Three ways to cut the arrow

The slope is learned from the variation of the nudge that the wind did not choose, and that variation is the σε² in the share. Three ways to have more of it.

Randomise. Let a coin set the nudge, κ = 0: then the bias term vanishes and β̂ = B (1.42 at a coin of spread 0.5 m/s). Or keep the operator and add noise: the share reaches 95 % only when σε² = 19κ²V, that is σε = 1.37 m/s, 4.0 times the operator's own spread. This has a price, and it is paid in the log: the logged episodes of the careful operator end 0.15 m from where a calm, un-nudged coast would have put the ball; at σε = 1.0 they end 1.45 m away, and with a coin of spread 0.5 m/s, 0.88 m. Exploring costs performance. A coin is also the cleanest instrument: it moves the nudge and, by construction, nothing else.

Measure what drove the choice. If the model's other inputs include something that explains part of the wind, the derivation holds with σw² replaced by the variance V that remains. With the wind itself as an input, V = 0 and β̂ = B for any σε > 0: the fit gives 1.47 at σε = 0.10. At σε = 0 no sensor repairs the log: the nudge is a function of the wind, the two cannot be separated, and the normal equations are singular. A nudge must have been tried at more than one value in the same wind. This condition is called positivity.

Infer it. The wind is hidden but constant, so the watch phase is evidence. With no nudge yet, the velocity increment beyond friction, u = v′ − e−γΔtv, is c·w plus a kick, where c = 0.098 m/s per m/s² is the velocity a unit wind adds in one step. Each watched step is a measurement of w with noise 0.31 m/s², about as informative as the prior itself (σw = 0.3). Precisions add, as in lesson 2:

1/Pk = 1/σw² + k·c²/σkick², so Pk/σw² = 1/(1 + 0.96·k)

This is lesson 2's filter with the wind as its state. After 1, 4 and 12 steps the wind left unexplained is 51 %, 21 % and 8 % of σw². Condition the model on the belief's mean and the same formula holds with V = Pk: the confounding that remains is the belief's variance. The raw velocity extracts far less (86 % unexplained at k = 4, 34 % at 12), because a snapshot cannot subtract the launch speed and a filter can.

what the model is givenwhat that needsslope learned (true 1.44)
the operator's log as it isnothing0.13
a coin picks the nudge (spread 0.5)the right to choose the nudges1.42
the wind itself as an inputa wind sensor, and a nudge that varies in the same wind1.47
the belief about the wind, k = 44 watched steps and the free-flight law0.43
the belief about the wind, k = 1212 watched steps (a later nudge)0.76

4 · The exam that sees through it

Cutting the arrow is a claim about the log. How would we know, without knowing the wind, that a model has learned the effect? Hold out episodes the model could not have been confounded by: fresh ones in which the nudge is set by fiat, independent of the wind. Compare its error there (the do-error) with its error on fresh episodes from the same process as the log (the see-error). The confounded model has a see-error of 0.18 m and a do-error of 0.92 m, 5 times larger. A model fitted to a coin's log has 0.49 m and 0.48 m: equal. (The 0.5 m they share is the wind itself, which neither can see; a model given the wind has a do-error of 0.12 m.) The rule: a model has learned an action's effect only if its error does not rise when the actions are set by fiat.

A simulator can ask the question directly, which a real log never can: rerun the same episode, same wind and same kicks, with a nudge of +1 m/s and with 0. The ball moves 1.44 m farther; the confounded model says 0.13 m. That counterfactual error of 1.30 m is what a planner would inherit.

The same trap exists on the policy side. Causal confusion in imitation (de Haan, Jayaraman and Levine, 2019) shows that ignoring causal structure under distribution shift causes "causal misidentification": more information can give a worse policy. Dreamer 4 (Hafner et al., 2025) builds a guard into its architecture: its agent tokens attend to every modality and no modality attends back, which the authors call crucial for avoiding causal confusion, since the world model's future predictions "can only be directly influenced by actions, not by the current task".

5 · Where the log is silent

Suppose the arrow is cut and the log is randomised. Is the model now right about every nudge? Only about the nudges that were tried. Isolate the question: a calm world with no kicks, the ball at x = 3.6 m moving at 1.6 m/s, and the response of the 2 s displacement to a nudge a between −3 and +3 m/s, the actuator's range. The world is the straight line B(1.6 + a) until the nudge carries the ball to the right wall, 4.3 m away, at a = 1.39 m/s. Beyond that the ball bounces back: at a = 3 it ends 2.21 m from where it started, where the line says 6.62 m.

The log is a dozen nudges drawn uniformly from ±0.8 m/s, which stops 0.59 m/s short of the wall. Fit five networks (tanh, two layers of 12 units) to bootstrap resamples of it. Inside the log's range the ensemble mean is wrong by 0.016 m (RMS); beyond it by 1.25 m, 79 times more. At a = 3 it predicts 5.67 m, which puts the ball at x = 9.27 m, 1.37 m past the wall, and the world delivers 2.21 m. The members know something is wrong, but not how much: the spread of their five predictions (the standard deviation across the members, averaged over a region) is 0.012 m inside the range and 0.086 m beyond it, and 0.17 m at a = 3, a 20th of the error.

Positivity applies to the action range as well: every nudge we will ask about must have been tried, with positive probability, in situations like the one we ask about. The operator violated it twice over: the nudges stayed near zero, and given the wind each was a single number.

6 · Spending the next nudges

Eight more nudges can be spent. Where? Uniformly over ±3 m/s, or where the members of the ensemble disagree most. PETS (Chua et al., 2018) reads epistemic uncertainty, the part more data would remove, from the disagreement of a bootstrap ensemble, and aleatoric uncertainty from a predicted variance. Plan2Explore (Sekar et al., 2020) turns the same disagreement into an intrinsic reward, the variance across an ensemble of one-step predictors of the next image embedding, and trains its exploration policy purely in imagination of a world model. Here: two rounds of four, retraining the ensemble after each, choosing the largest spreads at least 0.4 m/s apart. Averaged over 20 seeds, at equal data, the worst error over ±3 m/s is 1.97 m after random nudges and 0.70 m after nudges chosen by disagreement, which wins in 18 of the 20 seeds. On the widget's seed it is 2.44 m against 0.81 m.

Three limits. Disagreement ranks regions, it does not measure error: at a = 3 the spread was a twentieth of the error, and averaged over 20 seeds the edge error is 3.38 m against a spread of 0.19 m. It is epistemic, so it cannot reveal a variable the model does not contain: 200 bootstrap refits of the confounded slope agree to 0.019 and are all wrong by 1.31. And members can agree across a discontinuity none of them has seen: the wall is found only when a nudge reaches it.

The widget

Seeing, doing, and where the log is silent
Top left: the 1000 logged episodes, each a dot at its nudge and the distance the ball moved once the velocity's share is subtracted. The purple line is the slope the log teaches (state only), the green dashed line what a nudge really does. Top right: the slope each model learns as the noise in the nudges grows; curves are the closed form, dots the fits (purple sees the velocity, cyan also the belief about the wind, green the wind itself). Bottom: the calm world of §5; black is what a nudge really does, the shaded band where the log's nudges went, purple the five networks, and after a click cyan (random nudges) and amber (where the members disagree) while the first ensemble stays as a dashed purple line. Slopes are in m per m/s, errors in m.
slope, state only
—
closed form
—
slope, + belief
—
slope, + the wind
—
wind unexplained: velocity
—
wind unexplained: belief
—
see-error
—
do-error
—
miss of the logged episodes
—
ensemble error, inside the log
—
ensemble error, beyond it
—
spread inside | beyond
—
worst case, random nudges
—
worst case, by disagreement
—
Show the core JS
CA.nudge = function (ep, who, se) { return (who === 'operator' ? -CA.g.kappa * ep.w : 0) + se * ep.e; };
CA.outcome = function (ep, k, a, wind) {                            // the 2 s after a nudge a at step k: same wind, same kicks for every a
  var rk = CY.rng(ep.ko), q, t;
  CA.W.wind = [wind === undefined ? ep.w : wind, 0];
  q = CY.step(CA.W, ep.tr[k], [a, 0], rk);
  for (t = 1; t < CA.TAU; t++) q = CY.step(CA.W, q, null, rk);
  return q[0] - ep.tr[k][0];
};
CA.belief = function (ep, k) {                                      // N(m, P): the wind after k nudge-free steps; precisions add
  var g = CA.g, prec = 1 / (CA.SW * CA.SW) + k * g.c * g.c / (CA.SV * CA.SV), num = 0, t;
  for (t = 0; t < k; t++) num += g.c * ep.us[t] / (CA.SV * CA.SV);
  return { m: num / prec, P: 1 / prec };
};
...
  if (saa / n < 1e-8) return { ok: false, ap: ap, yp: ypa };             // the nudge is a function of the other inputs: nothing to learn its effect from
  var beta = say / saa;
...
CA.formula = function (V, se, kappa) { var g = CA.g; return g.B - g.C * kappa * V / (kappa * kappa * V + se * se); };
...
  var v = nets.map(function (n) { return n.predict([a / 3])[0] * 4; }), mean = CY.stats.mean(v), s = 0;
  v.forEach(function (x) { s += (x - mean) * (x - mean) / v.length; });
  return { m: mean, s: Math.sqrt(s) };
...
CA.pickDisagree = function (nets, nb, sep) {                       // the nudges where the members disagree most, at least sep apart
  var cand = CA.GRID.map(function (a) { return { a: a, s: CA.band(nets, a).s }; }).sort(function (p, q) { return q.s - p.s; }), picks = [], i;
  for (i = 0; i < cand.length && picks.length < nb; i++) if (picks.every(function (p) { return Math.abs(p - cand[i].a) >= sep; })) picks.push(cand[i].a);
  return picks;
};

What to try. Start as loaded: the operator, σε = 0.10 m/s, 4 steps watched. The purple line through the log has slope 0.13 against the green line's 1.44; the closed form says 0.13. Fresh episodes from the same operator give a see-error of 0.18 m, and with nudges set by fiat the do-error is 0.92 m. Slide σε to 0: the cloud goes flat (slope −0.02) and the slope given the wind reads n/a, because the nudge is a function of the wind. Slide it up: at 0.35 m/s the slope is 0.81 (closed form 0.80), at 1.0 it is 1.28 (1.31; the shortfall is the wall), and the logged episodes end 0.15, 0.52 and 1.45 m from the calm coast: the price of learning. Switch to the coin: every spread recovers the slope (1.42 at 0.5), but the estimate scatters more as the nudges shrink (a standard error of 0.16 at 0.1). Raise the steps watched: the slope with the belief is 0.43 at k = 4 and 0.76 at 12, while the raw velocity reaches 0.30. In the lower panel the log stops at ±0.8 m/s: the ensemble is wrong by 0.016 m inside and 1.25 m beyond, and at a = 3 it predicts 5.67 m against the world's 2.21. Drag the range to 2.0: the log reaches the wall, the worst case falls from 3.45 to 1.90 m, and the error inside rises to 0.22 m, because twelve nudges are too few to pin down a response with a kink in it (eighty bring it to 0.08 m). Click spend 4 more nudges twice: random nudges leave a worst case of 2.44 m; nudges where the members disagree leave 0.81 m.

What this lesson did not do
It treated a scalar nudge along the wind's axis and a wind constant within an episode. A wind that changes, or that answers the nudge, breaks the belief of §3 and needs a filter whose wind state can drift (lesson 2's wind state with process noise). It did not use the model: practising on it is lesson 8, where penalising by ensemble disagreement appears (MOPO and MOReL, both 2020), and planning with it is lesson 9. Where actions are missing rather than confounded, an inverse dynamics model can label video with them (VPT, Baker et al., 2022; lesson 14). Instrumental variables other than the coin, and the front-door criterion, need more of the graph than the Courtyard offers. Pictures (lesson 11) and objects (lesson 13) change the model class, not the argument.

Common mistakes / failure modes

"a model that fits the log well knows what my actions do"
The careful operator's log gives a see-error of 0.18 m and a do-error of 0.92 m (§4).
"more data will fix it"
At 200 000 episodes the slope is 0.136 ± 0.001, still 9 % of the truth (§2).
"a bigger network will fix it"
It learns the confounded relation: +1 m/s moves the ball 0.11 m in its eyes (§2).
"a little exploration noise is enough"
At σε = 0.35 m/s the log teaches 0.81 of 1.44; 95 % needs 1.37 m/s, and the logged miss grows from 0.15 to 1.97 m (§3).
"if the confounder is an input, it is solved"
The nudge must still vary in the same wind: at σε = 0 the slope given the wind cannot be fitted at all (§3).
"the ensemble will tell me where the model is wrong"
At a = 3 the spread is 0.17 m for an error of 3.45 m, and it is blind to an omitted variable (§5, §6).

Checkpoint exercise

Try it
A heavier ball has B = 2 m per m/s and C = 3 m per m/s², so an operator who cancels the wind uses κ = 1.5 s. The wind has σw = 0.4 m/s² and the model's other inputs leave all of it unexplained. (a) The operator's noise is σε = 0.3 m/s. What slope does the log teach? (b) How much total nudge noise would teach 90 % of B? Answer: (a) The share is σε²/(σε² + κ²σw²) = 0.09/(0.09 + 0.36) = 0.2, so the slope is 0.4 m per m/s. (b) A share of 0.9 needs σε² = 9κ²σw² = 3.24, so σε = 1.8 m/s, 3 times the operator's own spread κσw = 0.6 m/s.

Where this points next

A model fitted to varied, randomised nudges does say what a nudge causes, but where the nudges went. From a dozen nudges within ±0.8 m/s the ensemble is wrong by 0.016 m inside that range and 1.25 m beyond it, and eight more nudges chosen where its members disagree leave a worst case of 0.70 m on average, against 1.97 m for eight random ones, which still does not remove the edge. Now use it. The cheapest use is practice: let a policy improve on the model's imagined rollouts instead of in the world. Ask this model which nudge carries the ball farthest and it answers a = 3, the edge, where it predicts a ball 1.37 m past the wall and the world returns it 3.45 m short of the prediction. An optimiser does not wander; it goes to the place where the model is most wrong in its own favour. How do we learn from imagined experience without being fooled by it?

Takeaway
A log records actions that were chosen for reasons, so a regression on it learns what one sees, E[y | a], not what one gets, E[y | do(a)]. With a hidden wind and an operator who cancels it, the slope the log teaches is B·σε²/(σε² + κ²V): 0.13 against a true 1.44 for a careful operator, zero for a perfect one, and unchanged by 200 000 episodes or a bigger network. The cure is variation in the nudge that the wind did not choose: randomise it (and pay in the log's own performance), measure what drove it, or infer it from the episode's history as a belief whose variance Pk is the confounding that remains. The exam is to set actions by fiat: the confounded model's error rises from 0.18 to 0.92 m. A clean log still answers only where the nudges went, 0.016 m inside the range and 1.25 m beyond it, and the ensemble's disagreement can rank the silent regions but understates their error and cannot reveal a variable the model does not contain.

Interview prompts

Companion reads: Reinforcement Learning · Exploration vs exploitation (epistemic uncertainty used to explore), Reinforcement Learning · Imitation learning & inverse RL (the policy-side version of causal confusion), Reinforcement Learning · Offline RL (learning from a log you cannot extend), Lesson 22 · How the handle gets in (action conditioning in practice) and Lesson 23 · The actions do not exist, infer them (when the log has no actions at all).