What does my action cause?
Every model so far was fitted to a log of someone acting. Give that someone a reason the model cannot see, a gust of wind they can feel and push against, and the usual regression learns that a nudge of 1 m/s moves the ball by 0.13 m, where the simulator says 1.44 m. The log is not wrong; it answers a different question. This lesson derives the gap exactly, as the share of each nudge that the wind did not choose, cuts the link between wind and nudge three ways (randomise, measure, infer), tests a model against the question it will actually be asked, and then finds that even an unconfounded model speaks only where the nudges went.
New idea: a regression on a log learns what an action causes only from the part of the action that nothing else chose. To learn the cause, make that part larger by randomising the action, or by measuring or inferring what drove it, and trust the answer only where the actions went.
Forces next: A model fitted to varied, randomized actions can say what an action causes, at least where the data went. The cheapest use of it is practice: let a policy improve on the model's imagined rollouts instead of in the world. But a policy trained against a model finds the model's flattering mistakes and climbs them. How do we learn from imagined experience without being fooled by it?
1 · The log, and the model that learns from it
The world is the Courtyard, seen along its x axis. A launcher fires the ball, the agent watches it for k = 4 steps (0.4 s), adds one nudge to its velocity along x, and the log records how far the ball moves in the next 2 s. A hidden wind blows along x, constant within an episode and different between episodes: an acceleration drawn from N(0, 0.3²) m/s², the prior that lesson 2's wind state started from. Every step also adds a random kick of 3 cm/s to the velocity. The log keeps the ball's velocity exactly (lessons 2 to 4 showed how an agent obtains it) and does not keep the wind.
| symbol | meaning | unit |
|---|---|---|
| v | the ball's velocity along x when the nudge arrives: the state the model sees | m/s |
| a | the nudge, an impulse added to v | m/s |
| w | the hidden wind, a constant acceleration | m/s² |
| y | how far the ball moves along x in the 2 s after the nudge: what the model predicts | m |
Friction removes velocity at rate γ = 0.35 s−1 and the wind adds to it, so dv/dt = −γv + w. Integrating over τ = 2 s, with the nudge added to v at the start, gives
y = B·(v + a) + C·w + kicks
where B = (1 − e−γτ)/γ is how far a unit nudge carries the ball and C how far a unit wind pushes it. Reading both off the simulator (rerun one episode with a unit nudge, then with a unit wind) gives B = 1.44 m per m/s and C = 1.62 m per m/s². The friction integral gives the same B and a C of 1.60; the difference is the engine's five sub-steps per step. Where the ball is does not enter until it meets a wall (§5), so the state the model needs is its velocity.
The log is made by an operator who feels the wind and pushes against it. The push that cancels the wind's effect on average is a = −κw with κ = C/B = 1.13 s, and the operator adds a habitual noise ε of standard deviation σε = 0.10 m/s: a = −κw + ε. The log holds 1000 such episodes.
Now fit what lessons 4 to 6 fitted, least squares of the outcome on what the model sees, y ≈ θ0 + θv·v + β̂·a, and ask it what a nudge does. It answers β̂ = 0.13 m per m/s, 9 % of the truth. By the usual exam it is a good model: on fresh episodes from the same operator its error is 0.18 m, under a fifth of a metre. The model has learned that pushing the ball barely moves it.
2 · Seeing is not doing
The wind moves the outcome directly (C) and moves the nudge, because the operator reacts to it. The nudge moves the outcome (B). A regression of y on a cannot separate the arrow from a to y from the detour through w. It reports what one sees, E[y | a], the average outcome among logged episodes whose nudge was a. A model of consequences needs what one gets, E[y | do(a)], the average when the nudge is set by fiat and the arrow from w to a is cut. Here that is B(v + a), because the wind averages to zero. Among the logged episodes with nudge a it does not.
Let a⊥ and y⊥ be the nudge and the outcome after the model's other inputs have been regressed out of them. By the Frisch–Waugh theorem the coefficient of a in the full regression is the slope of y⊥ on a⊥. The outcome obeys y⊥ = B·a⊥ + C·w⊥ + noise, where w⊥ is the part of the wind the other inputs leave unexplained, of variance V. So
β̂ = Cov(a⊥, y⊥) / Var(a⊥) = B + C·Cov(a⊥, w⊥) / Var(a⊥)
With a = −κw + ε and ε independent of everything else, Cov(a⊥, w⊥) = −κV and Var(a⊥) = κ²V + σε², so, exactly (walls aside),
β̂ = B − C·κ·V / (κ²V + σε²), and for the operator who cancels the wind (κ = C/B): β̂ / B = σε² / (σε² + κ²V)
Read β̂/B as a share: the fraction of the nudge's variance that the wind did not choose. The ball's velocity at 0.4 s says little about the wind, so V is 86 % of σw² and the share at σε = 0.10 m/s is 9 %. For a perfect operator (σε = 0) the share is 0: the log says pushing does nothing, because every nudge was chosen to cancel exactly the wind that came with it. A nudge of +0.34 m/s, one standard deviation of the operator's spread, came with a wind of −0.30 m/s². It pushed the ball +0.49 m, the wind pushed it −0.49 m, and the log recorded 0.
| σε (m/s) | closed form | fit, 1000 episodes | fit, 200 000 episodes |
|---|---|---|---|
| 0 | 0.000 | −0.018 | 0.000 |
| 0.10 | 0.136 | 0.133 | 0.136 |
| 0.35 | 0.807 | 0.812 | 0.807 |
| 1.0 | 1.313 | 1.277 | 1.275 |
The table is the slope in m per m/s the state-only model learns, against the closed form with the measured V. They agree to the sampling error. At σε = 1 the fit falls a little below the formula because the largest nudges send the ball into a wall, in 1.8 % of the episodes, and where it bounces the response stops being a straight line; with the walls taken away the same log gives 1.31.
3 · Three ways to cut the arrow
The slope is learned from the variation of the nudge that the wind did not choose, and that variation is the σε² in the share. Three ways to have more of it.
Randomise. Let a coin set the nudge, κ = 0: then the bias term vanishes and β̂ = B (1.42 at a coin of spread 0.5 m/s). Or keep the operator and add noise: the share reaches 95 % only when σε² = 19κ²V, that is σε = 1.37 m/s, 4.0 times the operator's own spread. This has a price, and it is paid in the log: the logged episodes of the careful operator end 0.15 m from where a calm, un-nudged coast would have put the ball; at σε = 1.0 they end 1.45 m away, and with a coin of spread 0.5 m/s, 0.88 m. Exploring costs performance. A coin is also the cleanest instrument: it moves the nudge and, by construction, nothing else.
Measure what drove the choice. If the model's other inputs include something that explains part of the wind, the derivation holds with σw² replaced by the variance V that remains. With the wind itself as an input, V = 0 and β̂ = B for any σε > 0: the fit gives 1.47 at σε = 0.10. At σε = 0 no sensor repairs the log: the nudge is a function of the wind, the two cannot be separated, and the normal equations are singular. A nudge must have been tried at more than one value in the same wind. This condition is called positivity.
Infer it. The wind is hidden but constant, so the watch phase is evidence. With no nudge yet, the velocity increment beyond friction, u = v′ − e−γΔtv, is c·w plus a kick, where c = 0.098 m/s per m/s² is the velocity a unit wind adds in one step. Each watched step is a measurement of w with noise 0.31 m/s², about as informative as the prior itself (σw = 0.3). Precisions add, as in lesson 2:
1/Pk = 1/σw² + k·c²/σkick², so Pk/σw² = 1/(1 + 0.96·k)
This is lesson 2's filter with the wind as its state. After 1, 4 and 12 steps the wind left unexplained is 51 %, 21 % and 8 % of σw². Condition the model on the belief's mean and the same formula holds with V = Pk: the confounding that remains is the belief's variance. The raw velocity extracts far less (86 % unexplained at k = 4, 34 % at 12), because a snapshot cannot subtract the launch speed and a filter can.
| what the model is given | what that needs | slope learned (true 1.44) |
|---|---|---|
| the operator's log as it is | nothing | 0.13 |
| a coin picks the nudge (spread 0.5) | the right to choose the nudges | 1.42 |
| the wind itself as an input | a wind sensor, and a nudge that varies in the same wind | 1.47 |
| the belief about the wind, k = 4 | 4 watched steps and the free-flight law | 0.43 |
| the belief about the wind, k = 12 | 12 watched steps (a later nudge) | 0.76 |
4 · The exam that sees through it
Cutting the arrow is a claim about the log. How would we know, without knowing the wind, that a model has learned the effect? Hold out episodes the model could not have been confounded by: fresh ones in which the nudge is set by fiat, independent of the wind. Compare its error there (the do-error) with its error on fresh episodes from the same process as the log (the see-error). The confounded model has a see-error of 0.18 m and a do-error of 0.92 m, 5 times larger. A model fitted to a coin's log has 0.49 m and 0.48 m: equal. (The 0.5 m they share is the wind itself, which neither can see; a model given the wind has a do-error of 0.12 m.) The rule: a model has learned an action's effect only if its error does not rise when the actions are set by fiat.
A simulator can ask the question directly, which a real log never can: rerun the same episode, same wind and same kicks, with a nudge of +1 m/s and with 0. The ball moves 1.44 m farther; the confounded model says 0.13 m. That counterfactual error of 1.30 m is what a planner would inherit.
The same trap exists on the policy side. Causal confusion in imitation (de Haan, Jayaraman and Levine, 2019) shows that ignoring causal structure under distribution shift causes "causal misidentification": more information can give a worse policy. Dreamer 4 (Hafner et al., 2025) builds a guard into its architecture: its agent tokens attend to every modality and no modality attends back, which the authors call crucial for avoiding causal confusion, since the world model's future predictions "can only be directly influenced by actions, not by the current task".
5 · Where the log is silent
Suppose the arrow is cut and the log is randomised. Is the model now right about every nudge? Only about the nudges that were tried. Isolate the question: a calm world with no kicks, the ball at x = 3.6 m moving at 1.6 m/s, and the response of the 2 s displacement to a nudge a between −3 and +3 m/s, the actuator's range. The world is the straight line B(1.6 + a) until the nudge carries the ball to the right wall, 4.3 m away, at a = 1.39 m/s. Beyond that the ball bounces back: at a = 3 it ends 2.21 m from where it started, where the line says 6.62 m.
The log is a dozen nudges drawn uniformly from ±0.8 m/s, which stops 0.59 m/s short of the wall. Fit five networks (tanh, two layers of 12 units) to bootstrap resamples of it. Inside the log's range the ensemble mean is wrong by 0.016 m (RMS); beyond it by 1.25 m, 79 times more. At a = 3 it predicts 5.67 m, which puts the ball at x = 9.27 m, 1.37 m past the wall, and the world delivers 2.21 m. The members know something is wrong, but not how much: the spread of their five predictions (the standard deviation across the members, averaged over a region) is 0.012 m inside the range and 0.086 m beyond it, and 0.17 m at a = 3, a 20th of the error.
Positivity applies to the action range as well: every nudge we will ask about must have been tried, with positive probability, in situations like the one we ask about. The operator violated it twice over: the nudges stayed near zero, and given the wind each was a single number.
6 · Spending the next nudges
Eight more nudges can be spent. Where? Uniformly over ±3 m/s, or where the members of the ensemble disagree most. PETS (Chua et al., 2018) reads epistemic uncertainty, the part more data would remove, from the disagreement of a bootstrap ensemble, and aleatoric uncertainty from a predicted variance. Plan2Explore (Sekar et al., 2020) turns the same disagreement into an intrinsic reward, the variance across an ensemble of one-step predictors of the next image embedding, and trains its exploration policy purely in imagination of a world model. Here: two rounds of four, retraining the ensemble after each, choosing the largest spreads at least 0.4 m/s apart. Averaged over 20 seeds, at equal data, the worst error over ±3 m/s is 1.97 m after random nudges and 0.70 m after nudges chosen by disagreement, which wins in 18 of the 20 seeds. On the widget's seed it is 2.44 m against 0.81 m.
Three limits. Disagreement ranks regions, it does not measure error: at a = 3 the spread was a twentieth of the error, and averaged over 20 seeds the edge error is 3.38 m against a spread of 0.19 m. It is epistemic, so it cannot reveal a variable the model does not contain: 200 bootstrap refits of the confounded slope agree to 0.019 and are all wrong by 1.31. And members can agree across a discontinuity none of them has seen: the wall is found only when a nudge reaches it.
The widget
What to try. Start as loaded: the operator, σε = 0.10 m/s, 4 steps watched. The purple line through the log has slope 0.13 against the green line's 1.44; the closed form says 0.13. Fresh episodes from the same operator give a see-error of 0.18 m, and with nudges set by fiat the do-error is 0.92 m. Slide σε to 0: the cloud goes flat (slope −0.02) and the slope given the wind reads n/a, because the nudge is a function of the wind. Slide it up: at 0.35 m/s the slope is 0.81 (closed form 0.80), at 1.0 it is 1.28 (1.31; the shortfall is the wall), and the logged episodes end 0.15, 0.52 and 1.45 m from the calm coast: the price of learning. Switch to the coin: every spread recovers the slope (1.42 at 0.5), but the estimate scatters more as the nudges shrink (a standard error of 0.16 at 0.1). Raise the steps watched: the slope with the belief is 0.43 at k = 4 and 0.76 at 12, while the raw velocity reaches 0.30. In the lower panel the log stops at ±0.8 m/s: the ensemble is wrong by 0.016 m inside and 1.25 m beyond, and at a = 3 it predicts 5.67 m against the world's 2.21. Drag the range to 2.0: the log reaches the wall, the worst case falls from 3.45 to 1.90 m, and the error inside rises to 0.22 m, because twelve nudges are too few to pin down a response with a kink in it (eighty bring it to 0.08 m). Click spend 4 more nudges twice: random nudges leave a worst case of 2.44 m; nudges where the members disagree leave 0.81 m.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
A model fitted to varied, randomised nudges does say what a nudge causes, but where the nudges went. From a dozen nudges within ±0.8 m/s the ensemble is wrong by 0.016 m inside that range and 1.25 m beyond it, and eight more nudges chosen where its members disagree leave a worst case of 0.70 m on average, against 1.97 m for eight random ones, which still does not remove the edge. Now use it. The cheapest use is practice: let a policy improve on the model's imagined rollouts instead of in the world. Ask this model which nudge carries the ball farthest and it answers a = 3, the edge, where it predicts a ball 1.37 m past the wall and the world returns it 3.45 m short of the prediction. An optimiser does not wander; it goes to the place where the model is most wrong in its own favour. How do we learn from imagined experience without being fooled by it?
Interview prompts
- What is the difference between E[y | a] and E[y | do(a)], and which does a regression on a log estimate? (§2 — the first conditions on the logged nudge, the second sets it by fiat and cuts the arrow from the confounder; the regression estimates the first.)
- Why does the log of a perfect operator show that the control does nothing? (§2 — every control was chosen to cancel the disturbance that came with it, so the share of control variance the disturbance did not choose is zero.)
- Derive the least-squares slope when a confounder is left out. (§2 — β̂ = B + C·Cov(a, w)/Var(a), which for a = −κw + ε is B − CκV/(κ²V + σ_ε²).)
- Name three ways to cut the arrow from a confounder to an action and what each costs. (§3 — randomise: the log's own performance; measure: a sensor, and a nudge that still varies; infer: watched steps, and the belief's variance remains.)
- What is positivity and where did the operator violate it? (§3, §5 — every action must have positive probability in situations like the one asked about; given the wind the operator's nudge was one number.)
- How do you test a model for causal validity without re-running episodes? (§4 — hold out episodes with actions set by fiat and compare its error with the error on held-out log data; a factor of five means confounded.)
- Why is an ensemble's disagreement not a measure of what a model gets wrong? (§5, §6 — it is epistemic, it understates the error at the edge (0.17 m against 3.45 m), and it cannot see an omitted variable.)
Companion reads: Reinforcement Learning · Exploration vs exploitation (epistemic uncertainty used to explore), Reinforcement Learning · Imitation learning & inverse RL (the policy-side version of causal confusion), Reinforcement Learning · Offline RL (learning from a log you cannot extend), Lesson 22 · How the handle gets in (action conditioning in practice) and Lesson 23 · The actions do not exist, infer them (when the log has no actions at all).