all_lessons/World Models/02 · Belieflesson 2 / 31

The state you cannot see: belief

Lesson 1's planner had an exact model and still failed, because it was handed what a sensor gives, not the state: a position off by 10 cm, no velocity, and nothing while the ball is behind the curtain. This lesson asks what the agent should carry in its head instead. The reading is not the state; the whole history is one, but it grows; a window of readings forgets. What loses nothing is a belief, a probability distribution over the state, advanced by predicting with the model and correcting with each reading. For a ball with friction and a Gaussian sensor it is a mean and a covariance, and the update is the Kalman filter. We derive it, watch its doubt grow behind the curtain, test that doubt, and plan with it. It is exactly as good as the model written into it, and we wrote that by hand.

The thesis, here
What an agent should carry in place of the state is a distribution over it: a best guess alone cannot say how much weight the next reading deserves, or how likely a plan is to work. Predict with the model, correct with the reading, repeat, and a history that grows without bound shrinks to a belief of fixed size, 10 numbers in the Courtyard where a grid would need 230,400,000. The belief is only as honest as the model behind it.
Linear position
Forced by: A model that answers "if I push this way, where does the ball end up?" makes thinking cheaper than trying: one push in the real world instead of dozens, provided the model is handed the true state of the ball. An agent is never handed the state. It sees a noisy position, and nothing at all while the ball is behind the curtain. What should the agent carry in its head instead of the state it cannot see?
New idea: an agent that cannot see the state carries a belief, a probability distribution over it, and advances it in two steps: predict with the model, correct with the reading. In a world that is linear with Gaussian noise the belief is a mean and a covariance and the update is the Kalman filter; behind the curtain the correction is skipped and the doubt grows.
Forces next: The belief, a distribution over the hidden state updated by predict-then-correct, solves the curtain problem. But we built it from a model of the ball and a model of the sensor that we wrote down by hand, using state variables we chose. A real agent receives pixels, and nobody tells it which of a million numbers matter. What should an agent keep from what it sees, when the only teacher is the stream itself?
The plan
Seven moves. (1) Show that the reading is not the state. (2) Try summaries of the history by hand and see what each gets wrong. (3) Derive the belief and its two-step update. (4) Close the update for a linear world with Gaussian noise: the Kalman filter. (5) Follow it behind the curtain and test whether its doubt is honest. (6) Plan the nudge from the belief. (7) Find where the filter stops.

1 · The reading is not the state

Two things must be told apart. The state s = (x, y, vx, vy) is what the world needs in order to go on. The observation ot is what the agent receives at step t: the position plus an error of σO = 0.1 m on each axis, or nothing while the ball is inside the curtain. In 83 % of the launches below the ball is behind the curtain at the moment of the nudge.

Lesson 1 showed what this costs; here it is measured again on 120 launches of the nudge task, with the planner trying 128 candidate nudges. Handed the true state at the nudge, it succeeds on 96.7 % of them. Handed the last two readings, taken at face value, it succeeds on 1.7 %. With the step Δt = 0.1 s, the velocity it believes is wrong by √2·σO/Δt = 1.41 m/s on each axis (2.04 m/s as a vector, measured), and a velocity error δv moves the stopping point by δv/γ = 2.86 m per m/s (γ = 0.35 s−1 is the friction rate), so by about 4.0 m against a goal 1 m across.

The noise is not the deepest problem. Put two balls on the same spot, one moving at 1.0 m/s and one at 1.8 m/s. A reading, however sharp, is the same for both; one step later they are 0.079 m apart, less than this sensor's own error, and after the 8 s of the nudge task they are 2.15 m apart. A state is what makes the future independent of the past, given the action, the Markov property: p(st+1 | s0:t, a0:t) = p(st+1 | st, at). The four numbers have that property. A reading does not, since the readings before it say something about the velocity that the latest one lacks. What the future depends on is something the agent must reconstruct from every reading it has had.

2 · A summary of the history, by hand

The history of readings is a state in the trivial sense that nothing else exists to know. But it grows by two numbers a step, and every decision would re-read all of it. So try fixed-size summaries, in order of effort, and score each by what it does to the planner. Along each axis friction gives the ball the exact one-step motion p′ = p + g·v, v′ = d·v with d = e−γΔt = 0.9656 and g = (1 − d)/γ = 0.0983 s, so k steps on, pk = p0 + Ak·v0 with Ak = (1 − dk)/γ.

summary of the historystoresvelocity error at the nudgeplanner success
the last two readings, differenced4 numbers2.04 m/s1.7 %
a straight line through the last ten readings200.49 m/s6.7 %
the same fit, with friction in it200.12 m/s71.7 %
a fit through all readings so fargrows0.11 m/s77.5 %
the belief (§3 and §4)100.09 m/s79.2 %

Differencing amplifies the noise. A least-squares line through m equally spaced readings has slope variance σO²/Σ(ti − t̄)², and Σ(ti − t̄)² = Δt²·m(m² − 1)/12, so

σv = √12 · σO / (Δt · √(m(m² − 1))) = 1.41 m/s for m = 2, 0.11 m/s for m = 10

per axis. With ten readings the noise is nearly gone and a large error remains, because friction slows the ball by 3.4 % each step: the slope of a line is the velocity in the middle of the window, 0.45 s before its last reading, when the ball was 17 % faster, and the line then carries that velocity forward without slowing it. Averaging cannot remove a bias; the model has to be in the fit. Ask instead for the (p0, v0) that make p0 + Akv0 follow the readings, and carry it to the present with dk: the error falls to 0.12 m/s and the planner succeeds on 71.7 %. Using every reading raises it again (a window of ten throws information away), and adding what the agent knew about the launcher before the first reading gives the last row. A fit through everything is the best use of the readings under this model, and it has two defects: its size grows with the history, and it is recomputed from scratch at every step. Can it be advanced one reading at a time, in constant memory, and can the object also say how sure it is?

3 · The belief: predict, then correct

Let the belief be the distribution of the state given everything seen and done, bt(s) = p(st = s | o0:t, a0:t−1). Two assumptions hold in the Courtyard by construction (until §7 hides a wind in it): the state is Markov (§1), and a reading depends on the state now and on nothing else, p(ot | s0:t, o0:t−1) = p(ot | st). Then the belief can be advanced in two lines:

predict: b⁻(s′) = ∫ p(s′ | s, a) · b(s) ds correct: b₊(s) = p(o | s) · b⁻(s) / ∫ p(o | s″) · b⁻(s″) ds″

The first is the law of total probability applied to the dynamics: the state may have been anywhere the belief allows, and the dynamics say where each place leads. The second is Bayes' rule: weight every state by how well it explains the reading, and renormalise. With the ball behind the curtain there is no reading, the second line is skipped, and the belief only spreads. Everything the history says about the future now passes through the belief: for any question about what happens next, given the actions to come, p(future | history) = ∫ p(future | st) · bt(st) dst. The belief is a sufficient statistic of the history, when the model is right, and the history can be thrown away. Treating the belief as the state of a new, fully observed problem goes back to Åström (1965) and to Smallwood and Sondik (1973); Kaelbling, Littman and Cassandra (1998) used it to plan and act in partially observable stochastic domains. The price is the integral: in general b is a function of four variables and each step integrates over all of them.

4 · When the belief fits in ten numbers

Make the dynamics linear with Gaussian noise, s′ = F s + B a + ξ with ξ ~ N(0, Q), and the reading linear with Gaussian noise, o = H s + ε with ε ~ N(0, R). If the belief is N(x, P), then predicting pushes a Gaussian through a linear map and adds independent Gaussian noise, which gives a Gaussian with mean Fx + Ba and covariance FPFᵀ + Q. Correcting multiplies two Gaussians. The scalar case shows the result: a belief N(μ, P) about a number and a reading z of it with noise R. The exponent of the product is −(s − μ)²/2P − (z − s)²/2R, a quadratic in s. Its s² coefficient, −(1/P + 1/R)/2, fixes the new variance; its s coefficient, μ/P + z/R, puts the peak at P₊(μ/P + z/R). Hence

1/P₊ = 1/P + 1/R, μ₊ = P₊ (μ/P + z/R) = μ + K (z − μ), K = P / (P + R)

Precisions (1/variance) add; the new mean is the precision-weighted average; the reading moves the mean by the fraction K of the gap, large when the reading is sharper than the belief. With a matrix H that selects what the sensor sees, the same algebra gives the Kalman filter (Kalman, 1960):

predict: x⁻ = F x + B a, P⁻ = F P Fᵀ + Q
correct: ν = z − H x⁻, S = H P⁻ Hᵀ + R, K = P⁻ Hᵀ S⁻¹, x = x⁻ + K ν, P = (I − K H) P⁻

Here ν is the innovation, what the reading says that the belief did not expect, and S is the variance the belief predicted for it. A ball that touches nothing moves independently along x and y, so each axis gets its own filter, with mean x = (p, v) (position and velocity along that axis), F = [[1, g], [0, d]] and H = [1, 0]. Then S = P⁻pp + R and K = (P⁻pp, P⁻pv)ᵀ/S. A position sensor teaches the velocity through P⁻pv. Predict once from an uncorrelated belief with σp = 0.05 m and σv = 0.3 m/s and the covariance is P⁻pv = g·d·σv² = 0.0085 m²/s: where the ball is next depends on how fast it goes, so a reading above the prediction raises the velocity estimate too.

The agent starts from what it knows about the launcher, the first two moments of its launch distribution: x0 = 0.5 ± 0.05 m (the launcher stands at 0.5 m; the 5 cm is slack), y0 = 2.5 ± 0.87 m, vx = 2.46 ± 0.29 m/s, vy = 0 ± 0.43 m/s, the two axes taken as independent. Two axis filters are the whole belief: four means and six covariance entries, 10 numbers (correlations between all four variables would make it 14). The recursion is exact for free flight and a Gaussian start; the launch distribution is not Gaussian, which §5 tests. A bounce is not linear, so the statistics below use the launches that stay clear of the walls for 4 s (226 of 240).

The Courtyard adds no noise to the ball, Q = 0, and then there is something to check. The mean of this recursion should be the fit of §2 through the prior and all readings so far. We solve that fit by the normal equations at every step of every launch and compare: the means agree to better than 10−9 and the covariances to better than one part in 108. The recursion is the full-history fit, advanced one reading at a time in constant memory. If the filter is allowed a velocity kick q each step (a doubt), the covariance stops shrinking, old readings fade, and the gain settles at K∞ = (0.247, 0.347 s⁻¹) for q = 0.05 m/s, the fixed point of the same recursion. With q = 0 the gain decays towards zero (0.0018 after 600 readings): the filter stops listening, which is right if the model is exact and costly if it is not.

Road not taken · keep the whole belief on a grid
The recursion of §3 can be run as it stands, with b stored as a table of cells, and then no Gaussian is needed: any dynamics, any sensor, any shape. The price is the cells. At 5 cm in position and 5 cm/s in velocity, over the arena and ±3 m/s, a four-dimensional grid has 160 × 100 × 120 × 120 = 230,400,000 cells, 1.8 GB at 8 bytes a cell, against 10 numbers for the Gaussian. The two axes are independent in free flight, so two small grids would do (31,200 cells); each hidden variable we add multiplies the cost again, and a bounce couples the axes. A particle filter keeps samples instead of cells and escapes the table, but a sharp reading leaves most of them with almost no weight: after the very first reading the effective sample size, (Σw)²/Σw², is only about 10 % of the samples drawn from the prior. The Gaussian is the case where the integral costs nothing.

5 · Behind the curtain, and is the doubt honest?

No reading, no correction, so predict repeatedly. With q = 0, P⁻ = Fk P (Fk)ᵀ after k blind steps, and by induction Fk = [[1, Ak], [0, dk]]. Reading off the top-left entry:

Ppp(k) = Ppp + 2 Ak Ppv + Ak² Pvv, Ak ≤ 1/γ = 2.86 m per m/s

The doubt about position grows with the doubt about velocity and with their correlation, and friction caps it: a velocity doubt of δv can move the ball by at most δv/γ. After the ten readings before the curtain the x-filter has σp = 0.052 m, σv = 0.067 m/s and correlation 0.86. Blind for 0.5 s its position std is 0.080 m, after 1 s 0.104 m, 2 s 0.143 m, 4 s 0.190 m, and it can never pass 0.237 m, the std of the stopping point p + v/γ. When the ball reappears, the first reading is a scalar fusion, and its weight is set by the doubt it meets: the gain on the position is Kp = 0.39 after half a blind second and 0.67 after two, which a mean without a covariance could not know. After two blind seconds Kv = 0.153 s⁻¹ and the position std falls from 0.143 m to 0.082 m, since P₊ = PR/(P + R) is below both P and R. What the agent knows decays behind the curtain and snaps back.

A belief is a claim, my error has covariance P, and a claim can be tested in two ways. If the filter is right, the innovation satisfies E[νᵀS⁻¹ν] = 2, the dimension of the reading (x and y); averaged over the two axes, the normalised innovation squared (NIS) has mean 1. That test needs only the readings. The other needs the truth, which a simulator has and a real agent does not: the truth should fall inside the belief's 95 % region 95 % of the time, its coverage. For a 2-D Gaussian error u with the claimed covariance P, uᵀP⁻¹u (the squared error in units of the belief's own spread) is below c with probability 1 − e−c/2, so the 95 % ellipse is c = −2 ln 0.05 = 5.99, which is 2.45 standard deviations rather than 1.96. On 3000 launches drawn from the filter's own prior the mean NIS is 1.00 and the coverage 95.0 %, as the algebra demands. On the launcher's own (not Gaussian) launches it is 0.98 and 95.5 % over 1,423 wall-free launches. Now tell the filter the wrong sensor noise, and watch which way each test fails, on the 226 launches of the widget:

the filter believes the sensor error ismean NIS (1 = honest)truth inside the 95 % ellipseerror behind the curtain / the std it claims
half the real one3.7654.8 %0.074 / 0.037 m
the real one, 0.1 m0.9794.9 %0.071 / 0.070 m
twice the real one0.27100 %0.080 / 0.127 m

A filter that thinks its sensor better than it is trusts the readings too much: its covariance is too small, the NIS is above 1 and the truth escapes the ellipse. A filter that thinks the sensor worse is too timid: NIS below 1 and a region that is too big. The error itself barely moves between the rows; what changes is whether the filter's claim about it is true. A failed test does not say what is wrong, only that something is.

6 · Use it: plan from the belief

The planner of lesson 1 needs a state to imagine from. Give it the belief mean at the nudge. On the 120 launches it succeeds on 79.2 % of them, up from 1.7 % with the last two readings and against 96.7 % with the true state (600 other launches give 79.5 %). The belief also says what to expect: draw 48 states from it, fly each out with the chosen nudge, and count how many end in the goal; the belief claims 74.9 % (76.5 % over 600). The claim and the result agree because the belief is honest. Its spread also predicts what is attainable. At the nudge the stopping point of the ball, p + A81v (81 steps: the nudge and the 80 that follow), has a standard deviation of about 0.25 m on each axis, and a cloud that size centred exactly on the goal lands inside its 0.5 m radius with probability 1 − e−r²/2σ² = 87.0 %. The 128-candidate planner reaches 79.2 %, because its candidates cannot put the mean on the centre (the best one misses it by 0.24 m on average); with 512 candidates it reaches 88.3 %, at that limit to within the ±3 points of a 120-launch sample. The remaining failures are what the readings leave unknown. No planner can recover them, and only the belief can say how many there will be.

The whole belief can drive the choice too: score each candidate by the share of 16 sampled states that it sends into the goal. That gives 73.3 %, no better than aiming the mean. Here the ending is nearly the believed state translated (for candidates that end near the goal the size of the cloud of endings differs by 8 % on average), so centring the cloud and aiming the mean are the same act. The spread is for judging risk, not for aiming.

A belief behind the curtain
Arena: the true path (black), the readings (cyan; none inside the grey curtain), the belief mean (purple) and its 95 % ellipse every fourth step; amber: the belief at the nudge, 1.2 s. Error plot: the belief's error along x (cyan) and y (amber) against its own ±2σ band, curtain time shaded. Nudge panel: the nudge chosen from the belief mean, 48 states drawn from the belief and flown out with it (cyan dots), the real ending, how often the nudge works (bars), and the camera's last picture. Statistics use the launches that never touch a wall in 4 s.
std at curtain exit
—
std after first reading
—
error behind curtain
—
std it claims there
—
mean NIS (1 = honest)
—
truth in 95 % ellipse
—
launches in statistics
—
success: true state
—
success: last two readings
—
success: belief mean
—
success the belief claimed
—
this launch
—
Show the core JS
BL.predict = function (f) {                                      // x <- F x,  P <- F P F' + Q
  var n = f.n, F = f.F, x = f.x, P = f.P, T = f.T, i, j, k, s;
  for (i = 0; i < n; i++) { s = 0; for (j = 0; j < n; j++) s += F[i * n + j] * x[j]; f.xt[i] = s; }
  for (i = 0; i < n; i++) x[i] = f.xt[i];
  for (i = 0; i < n; i++) for (j = 0; j < n; j++) { s = 0; for (k = 0; k < n; k++) s += F[i * n + k] * P[k * n + j]; T[i * n + j] = s; }
  for (i = 0; i < n; i++) for (j = 0; j < n; j++) { s = f.Q[i * n + j]; for (k = 0; k < n; k++) s += T[i * n + k] * F[j * n + k]; P[i * n + j] = s; }
};
BL.update = function (f, z) {                                    // nu = z - x0,  S = P00 + R,  K = P[:,0] / S,  x <- x + K nu,  P <- (I - K H) P (I - K H)' + K R K'  (equal to (I - K H) P, kept symmetric)
  var n = f.n, P = f.P, x = f.x, T = f.T, K = f.K, i, j, S = P[0] + f.R, nu = z - x[0];
  for (i = 0; i < n; i++) K[i] = P[i * n] / S;
  for (i = 0; i < n; i++) x[i] += K[i] * nu;
  for (i = 0; i < n; i++) for (j = 0; j < n; j++) T[i * n + j] = P[i * n + j] - K[i] * P[j];
  for (i = 0; i < n; i++) for (j = 0; j < n; j++) P[i * n + j] = T[i * n + j] - T[i * n] * K[j] + f.R * K[i] * K[j];
  f.nu = nu; f.S = S;
};

What to try. Move curtain width first. At the default 0.8 m (launch #16) the purple ellipses shrink along the visible track, widen inside the grey band and shrink again at the first reading; averaged over the launches the position std is 0.082 m at the curtain's exit and 0.066 m after the first reading. At 3 m it becomes 0.176 and 0.087 m (only 135 of the 226 launches get out within 4 s), and at 0 the planner rises from 79.2 % to 87.5 %, since the ball is never lost. Then, with the curtain back at 0.8 m and one other control changed at a time: the sensor menu reproduces the table of §5 and adds what the planner makes of it, since the nudge that half the real one claims will work 89.4 % of the time works 76.7 %, and twice the real one claims only 49.3 % and achieves 71.7 %; doubt 0.05 widens the band, drops the claim to 48.2 % and leaves the result at 79.2 %; and with the wind, the filter unaware, the cyan error leaves its band after about two seconds and the planner claims 76.8 % while achieving 4.2 %, where a wind state makes it claim 5.1 %.

7 · Where the filter stops

The model. The filter is exactly as good as the world we wrote into it. Give the Courtyard a hidden wind, a constant acceleration of 0.2 m/s² toward the launcher, and keep the filter as it is. During the first 1.2 s, the time of the nudge, the wind is invisible: the NIS over those steps is 0.98 (1.04 with the velocity known exactly), because by the last reading before the curtain, at about 0.9 s, the wind has moved the ball by only ½·0.2·0.9² = 0.08 m, less than the sensor's own error. It shows over time (the NIS averaged over the first 3 s is 1.29, over 4 s 1.89), by which point the nudge has been chosen. The planner then imagines a windless future, though the wind adds up to 3.1 m by the end: it claims 76.8 % and achieves 4.2 %, where a planner handed the true state in the same wind succeeds on 87.5 %. Two repairs, tested in the widget:

the filterNIS over 4 struth in 95 % ellipsethe planner claimsand achieves
no wind in the world, plain filter0.9794.9 %74.9 %79.2 %
wind, plain filter1.8952.3 %76.8 %4.2 %
wind, the filter doubts itself (q = 0.05 m/s)0.9895.5 %49.1 %5.8 %
wind, a wind state in the filter0.9596.2 %5.1 %6.7 %

Doubt restores the NIS and the coverage because it widens every band; it hides the wind in the noise, and the planner still claims far more than it achieves. The wind state does the right thing. The state grows to (p, v, w), F gains the column that one step of wind adds to position and velocity (0.0059 m and 0.0979 m/s per m/s² of wind), and the wind starts as 0 ± 0.3 m/s². After 4 s the filter has found it, −0.20 ± 0.025 m/s². But at the nudge, 1.2 s in, the readings have narrowed it only from ±0.30 to ±0.24 m/s², so the belief about where the ball will end is wide and the planner now claims only 5.1 %, close to what it achieves. It does not know the wind, and it knows that it does not know. Someone had to think of adding the wind. An agent has no list of the variables that matter.

The sensor. The filter also assumed a sensor that reports the position, H = [1, 0]. A camera reports a picture. The small picture of the Courtyard drawn in the widget has 24 × 15 = 360 numbers; the ball in it is a Gaussian blob of width σ = 0.8 pixel that changes about 11 of the numbers by more than 0.05, the rest being floor, goal and curtain. The map from the ball's position to the 360 numbers is not linear: the picture with the ball halfway between two positions is not the average of the pictures at the two. In the Courtyard we own the renderer and could write it down; for a real camera nobody can. The belief of this lesson has no way to start from pixels.

What this lesson did not do
It filtered a ball in free flight; walls and posts make the dynamics nonlinear and need the grid or particle filters of §4, or a linearised filter. It ignored what a missing reading says (the ball is inside the curtain), which would cut the belief at the curtain's edges and make it non-Gaussian. It took the model as given: lesson 3 asks what to keep when the observation is a picture, lesson 4 learns the filter and the dynamics from data, and lesson 5 turns to a future that is not one point. It used the belief to aim once; lesson 9 plans over many steps, re-planning after each action, and lesson 7 returns to the wind as a confounder in logged data.

Common mistakes / failure modes

"the observation is the state"
A reading has no velocity, 10 cm of error and gaps: handed the last two readings, the exact planner succeeds on 1.7 % of launches, not 96.7 % (§1).
"average enough readings and the state is known"
A line through ten readings has little noise and a 17 % friction bias: planner success 6.7 %. The model has to be in the fit (§2).
"behind the curtain the agent knows nothing"
It knows what the model implies: the position std grows from 0.052 m to 0.143 m in two seconds, never passes 0.237 m, and falls at the first reading (§5).
"a filter's covariance is a fact"
It is a claim. Told the sensor is twice as good as it is, the filter's NIS is 3.76 and its 95 % ellipse holds the truth 54.8 % of the time (§5).
"the filter passes its test, so the model is right"
With a hidden wind the NIS over the first 1.2 s is 0.98, yet the planner claims 76.8 % and achieves 4.2 %: a test sees only what the readings have shown (§7).
"add process noise to be safe"
Doubt restores the NIS (0.98), not the future: the planner still claims 49.1 % and achieves 5.8 % (§7).

Checkpoint exercise

Try it
The belief about the ball's x-position is 2.00 m with standard deviation 0.20 m. The sensor (error 0.1 m) reads 2.15 m. (a) What are the gain, the new mean and the new standard deviation? (b) Repeat with a filter that wrongly believes the sensor error is 0.3 m. (c) If the belief's 0.20 m is right, what is the mean NIS of the filter in (b)? Answer: (a) With P = 0.04 and R = 0.01, K = P/(P + R) = 0.04/0.05 = 0.8; mean 2.00 + 0.8·0.15 = 2.12 m; std √(P(1 − K)) = 0.089 m. (b) R = 0.09 gives K = 0.31, mean 2.05 m, std 0.17 m: it listens less and doubts more. (c) The innovation really has variance P + Rtrue = 0.05, the filter divides by S = 0.13, so the mean NIS is 0.05/0.13 = 0.38: underconfident.

Where this points next

The belief solves the curtain: a distribution over the hidden state, advanced by predicting with the model and correcting with each reading, in ten numbers in the Courtyard, and the nudge planned from it works on 79.2 % of launches where the last two readings gave 1.7 %. It is also exactly as good as what we wrote into it: with a hidden wind the same filter claims 76.8 % and achieves 4.2 %, and with the wind added to the state it knows that it does not know. Every ingredient was written by hand: the state variables, F, H, Q and R. A camera hands over 360 numbers per frame, of which about 11 carry the ball, through a map nobody wrote down, and the belief of this lesson cannot start from that. What should an agent keep from what it sees, when the only teacher is the stream itself?

Takeaway
The observation is not the state: it has no velocity, it is 10 cm wrong, and it vanishes behind the curtain. The history is a state, but it grows, and summaries by hand fail one defect at a time (two readings amplify the noise, a straight line is biased by friction, a window throws readings away). The belief is the distribution of the state given the whole history, advanced by predict (push it through the dynamics) and correct (weight it by the reading); in a linear world with Gaussian noise that is the Kalman filter, a mean and a covariance, equal to the full-history fit in constant memory. Behind the curtain its doubt grows, capped by friction, and snaps back at the first reading. The doubt is a claim to be tested, by NIS near 1 and 95 % coverage, and the wrong sensor noise fails both in a known direction. Planned from the belief mean, the nudge works on 79.2 % of launches and the belief predicts 74.9 %. The filter is exactly as good as its model: a hidden wind makes it claim 76.8 % where it achieves 4.2 %, and a camera would need a model nobody has.

Interview prompts

Companion reads: Reinforcement Learning · What the agent sees (observation, the Markov line, POMDPs), 3D Vision · 05 Pose by agreement II (the batch version of the same fit: all poses against all observations at once) and Synthetic Vision · 03 The camera is a measurement (where a noise model like R comes from).