all_lessons/Robot Model Training/09 · Videolesson 9 / 24

Watching: learning from video without actions

Lesson 8 pooled demonstrations from other arms by writing every action as the place the hand should go. That pool is small beside the video people have shot of themselves at work, and video shows where things went but has no action column: a regressor has nothing to fit. This lesson recovers the column. The cup's next position is the consequence of the command, so one number measured on a few labelled seconds, what the plant does per unit command, turns each pair of frames into a pseudo-command. Two labelled demonstrations and 2.5 minutes of footage take a clone from 54.5 % to 84.0 %. What decides this is which label is used, and it is not the one with the smallest error. What no label recovers is a push: a push of 15 cm/s and one of 30 cm/s on a table leave footage 1.2 mm apart.

The thesis, here
Footage holds where the cup went and none of the commands that sent it. A command is the cause of the step that follows it, so it can be read back from that step: divide the step by what the plant does per unit command, a number a few labelled seconds measure. Each label has to be its own frame's command; a label that mixes in the step that arrived at the frame, or the frames after it, spoils the policy even when its error is small. What the plant does not turn into motion is not in the footage at all: force, effort, intent.
Linear position
Forced by: Demonstrations from other bodies help when the action is written in a space every body shares, the place the hand should go, and each arm turns that into its own joint motions; written in joint angles they are worse than no data. Pooled this way, all the robot demonstrations ever recorded are a small fraction of what people have filmed themselves doing, and the video has no actions in it. What can be learned from watching, and what is missing from the video that no amount of it can supply?
New idea: an action can be read back from the frame after it by inverting the plant, so a few labelled seconds label any amount of footage, if each label is the command of its own frame. A label is priced by the pull-back it leaves the policy (how hard it pulls the cup back toward the path), not by its error.
Forces next: A model that infers the missing actions from a few hours of labelled motion can turn a large pile of video into demonstrations, and what it learns is what to do next, not how hard to push. Video records where things moved, never the force that moved them: an insertion with a millimetre of clearance, watched through a camera that cannot resolve a millimetre, looks the same when it succeeds and when it jams. What must a policy sense, and how must it act, when the task is decided by something the camera cannot see?
The plan
Six moves. (1) Count what each kind of hour holds and put footage on the Bench. (2) Invert the plant: the next frame shows the command, and a few labelled seconds fix the one number. (3) Rule out three ways to skip the effect. (4) Run the pipeline and see what footage buys. (5) Score a label by its pull-back, not its error. (6) Find the column that is not there.

1 · Two kinds of hour

Ego4D (Grauman et al., 2022) holds 3,670 hours of daily-life video; besides pixels it carries narrations and benchmark labels, and no joint command, end-effector command or force. The real-robot hours GR00T N1 (NVIDIA, 2025) pooled from ten sources come to 3,288.8, and video models are trained on far more: Cosmos (NVIDIA, 2025) collected about 20 million hours of raw video, more than a thousand times the robot hours added up in lesson 8. A robot hour has a state and a command in every frame and is scarce; a video hour has the state and is plentiful.

On the Bench, footage is the expert's runs of the five-post course with the commands thrown away: the cup's position at 20 Hz and nothing else. A labelled run keeps them. The robot owns two labelled demonstrations (365 frames, 18 s). The footage is shot by a hand that wobbles twice as much as the arm: its gust is 0.10 rad/s against the arm's 0.05, so it wanders into states the two demonstrations never visit. One clip is one run, 9.2 s. As in lesson 8 the policy reads the cup's position and answers with the velocity it wants for the cup; the arm's own controller turns that into joint velocities.

A policy is a table of (position, command) pairs (lesson 1), and the footage has no commands: fed 16 clips it has nothing to store. The two labelled demonstrations alone make a clone that reaches the mat in 54.5 % of 200 runs (interval 47.6 to 61.3 %), and more of the robot's own demonstrations help slowly: 63.0 % with 8, 69.0 % with 200. The footage has the coverage the demonstrations lack, and it lacks the label. How does the label come back?

2 · The effect shows the cause

Write zt for the cup's position at frame t as the footage shows it (m), vt for the cup velocity commanded at that frame (m/s, the missing column), Δt = 0.05 s for the frame time, nt for the gust's effect on the cup during the step (m) and g for what the plant does per unit command (1 on the Bench, unknown to the learner). The arm turns a wanted cup velocity into joint velocities, the plant adds the gust, and the cup moves:

zt+1 − zt = g · Δt · vt + nt

The command is the cause of the step, so inverting the plant reads it back:

v̂t = (zt+1 − zt) / (ĝ · Δt)

That is the whole inverse-dynamics model. It has one number, ĝ, and the labelled runs measure it by regressing the realised velocity on the logged command: ĝ = Σ (Δz/Δt) · v / Σ v · v. The two labelled demonstrations give ĝ = 0.997. The label's error is the gust, nt / (g Δt): 5.9 cm/s per axis on the footage, against a speed of 15 cm/s. Relative to the size of the commands the error is 27.8 % on labelled frames and 56.5 % on the footage, twice as much because the hand wobbles twice as much.

What the labelled hours buy is ĝ and the alignment. Resample the labelled data: from one demonstration ĝ is good to ±1.4 %, from sixteen to ±0.4 %, falling as n−1/2 (fitted exponent −0.47), as a sum of squares predicts. How good must it be? Scale every label by a factor and train: 0.8 reaches 84.0 %, and 0.7 reaches 0.0, because a policy that moves too slowly never finishes in the steps it is given. So ĝ has to be right to better than a fifth, and one demonstration gives it to about a hundredth. The labelled frames also say which pair of frames carries the command: of the alignments from two frames early to three late, the best fit is the pair (t, t+1) in 60 of 60 resamples of one demonstration.

Real footage is harder: the labelled hours scale with what the model has to learn, not with the footage. VPT (Baker et al., 2022) trained a model of about half a billion weights on 1,962 hours of labelled Minecraft play, applied it to about 70,000 hours of web video, and found that at least 10 hours of labelled data were needed for any crafting, with gains levelling off after 100 hours.

3 · Three ways to skip the effect

Three tempting shortcuts need no inverse of the plant. Each fails by a computation.

Label the footage with the policy. Let the clone built from the two labelled demonstrations answer for every footage frame and retrain on its answers. The labels hold only what it already says: 0.0 % of the footage frames lie beyond its reach of three bandwidths, so it answers all from what it stores. The retrained policy reaches 48.5 % against the clone's 54.5 %, and its labels err by only 9.6 % of the command. A small label error is no sign that a label carries anything new.

Road not taken · no labels at all
A latent-action model learns a code for what changed between frames and never sees a command. The footage cannot say which code is which: if the code is M v for any invertible M, a decoder that undoes M explains the next frame exactly as well, so every M is as good as every other. A policy trained on the codes outputs codes, and the arm needs metres per second. Labelled frames fix M: fitted by least squares on n of them, gust included, the median error of the recovered map is 46 % with 2 frames, 19 % with 4, 14 % with 8 and 9.3 % with 32. Genie (Bruce et al., 2024) learns 8 latent actions from unlabelled video and needed 200 labelled expert samples on CoinRun to match a policy trained on true actions.
Road not taken · predict the next frame, then decode it
A video model predicts the change of position from the position, and a policy decodes it into a command. Here the video model is the kernel average of the observed steps and decoding divides by ĝ Δt. The kernel average is linear, so averaging the labels and labelling the average are the same operation: the two policies differ by 0.0 cm/s at every footage position, to a billionth. It is the inverse-dynamics model with its steps reordered, and it needs the same ĝ.

The widget

Label the footage, copy it, press the plate
The table from above: footage (grey, positions only), labelled demonstrations (violet) and 12 of 200 runs of the policy trained on them (cyan reached the mat, red touched a post). Top right: each footage frame's label against the command its hand really gave (sideways is up and down in the table view). Bottom left: success against footage for the chosen label (amber), the true commands (green) and no footage (grey). Bottom right: label error, success and the policy's pull-back (section 5).
policy reaches the mat / presses the plate
—
95 % interval
—
labelled demonstrations alone
—
with the true commands
—
label error on the footage
—
pull-back of the policy
—
gain ĝ from the labelled runs
—
push inferred of push given (table frames)
—
Show the core JS
IL.fitGain = function (runs) {
  var sxy = 0, sxx = 0, n = 0;
  runs.forEach(function (r) {
    for (var t = 0; t < r.V.length; t++) for (var a = 0; a < 2; a++) { sxy += (r.Z[t + 1][a] - r.Z[t][a]) / IL.DT * r.V[t][a]; sxx += r.V[t][a] * r.V[t][a]; n++; }
  });
  return { g: sxy / sxx, n: n };
};
...
IL.say = function (M, Z, t) {
  var o = [0, 0];
  for (var k = 0; k < M.js.length; k++) { var zz = Z[t + M.js[k]], z0 = Z[t]; o[0] += M.c[k] * (zz[0] - z0[0]); o[1] += M.c[k] * (zz[1] - z0[1]); }
  return o;
};
IL.label = function (M, runs) {
  var out = [];
  runs.forEach(function (r) { for (var t = M.B; t + M.F < r.Z.length; t++) out.push([r.Z[t], IL.say(M, r.Z, t)]); });
  return out;
};

What to try. At the defaults (16 clips, two labelled demonstrations, label = the next step) the policy reaches the mat in 84.0 % of its runs (interval 78.3 to 88.4 %); the demonstrations alone reach 54.5 % and the true commands 93.5 %; the label error is 56.5 % and the pull-back +0.16 per second. One clip (9.2 s) already gives 79.5 %. The step that arrived has an error of 59.6 %, as large, and reaches 17.5 % with a pull-back of −0.22. The filter fitted over 13 frames has an error of 19.1 % and reaches 50.0 %; the policy's own answer, 9.6 % error, reaches 48.5 %. Footage as steady as the arm: the next step reaches 70.0 %, the true commands 71.5 %. Press the plate: the model-labelled policy presses it in 0.0 % of runs at every footage setting, the true commands in 100.0 %, and the model infers a push of 0.0 where 15.0 cm/s was given.

4 · What the footage buys

Two labelled demonstrations and 16 clips with inferred commands reach 84.0 % where the demonstrations alone reach 54.5 %: a lift of 29.5 points. With the commands the footage's hand really gave, 93.5 %, so the inverse model costs 9.5 points. Almost all of the lift is there after one clip: 0 clips 54.5 %, 1 clip 79.5 %, 16 clips 84.0 %, and 80 clips (12.2 minutes) 85.0 %, against 92.0 % for the true commands. The Bench has one course, which one clip already covers. Further clips close the gap to the true commands slowly, as the regression averages the label's noise: 9.5 points at 16 clips, 7.0 at 80.

What the footage supplies is coverage: states the robot's tidy demonstrations never visit, with the recovery commands the demonstrator gave there (lesson 2). That is why 200 labelled demonstrations of the same kind reach only 69.0 %, and why footage as steady as the arm gives only 71.5 % even with the true commands. In practice what matters is the number of situations the hours cover (lesson 7's coverage law), not their length.

5 · Score a label by its pull-back

The next step has a label error of 56.5 %. The filter fitted over 13 frames, the label a practitioner would fit by regression on labelled data, has 19.1 % on the footage (9.7 % on held-out labelled frames). Each policy was trained on 16 clips labelled one way:

label of frame tlabel errorpull-back (per s)reaches the mat
the true commands0 %+0.1593.5 %
the next step, (zt+1 − zt) / ĝΔt56.5 %+0.1684.0 %
no footage (two labelled demonstrations)—−0.0154.5 %
the average of the next step and the one that arrived41.2 %−0.0355.0 %
a filter fitted over 13 frames19.1 %−0.0150.0 %
the policy's own answer9.6 %−0.1248.5 %
the step that arrived, (zt − zt−1) / ĝΔt59.6 %−0.2217.5 %

The rows are not ordered by error: the next step, whose error is the second largest, beats the policy's own answer, whose error is the smallest of the estimated labels, and the step that arrived, with about the next step's error, loses 66.5 points to it. Something else decides, and it shows near the path. Let δ be the cup's sideways offset from the path and m(δ) the policy's sideways command there. The loop is δ̇ = m(δ) + gust. If m(δ) = −pδ the offset is pulled back at rate p and wanders about sv√(Δt/2p), with sv the gust's spread as a velocity at the cup; if p < 0 it grows as e|p| t. Call p the pull-back. The expert's is 1.13 per second; a kernel of 1.5 cm smooths across the offsets that carry the feedback and keeps 0.15, which is all that holds the cup to the path.

A label error e that correlates with the offset, E[e | δ] = cδ, takes c away from that: ppolicy = p − c. The next step's error is the gust of a step that has not happened, so it cannot know where the cup is: c = 0, and the pull-back stays at 0.16. The step that arrived contains the gust that put the cup at δ. The offset it arrived at is the old offset plus Δt times that gust, so with sv = 5.9 cm/s per axis and an offset spread of sδ = 2.15 cm the best linear guess of the gust from the offset is E[gust | δ] = (sv² Δt / sδ²) δ, and c = 0.38 per second. The loss measured in the table is 0.15 − (−0.22) = 0.37, and the average of the two steps loses half of it, 0.18. A pull-back of −0.22 per second turns an offset of 1 cm into more than 10 cm over a 12-second course. The policy is a function of position, so its labels may not know where the cup came from.

Two controls separate the effects. Take the true command of the frame s later, with no gust in it: s = +1 gives 66.0 %, +2 gives 25.0 %, +3 gives 0.0 %, because a command from the future belongs to a place further along. The command of the frame before gives 96.5 % (two before, 84.5 %): a late label is harmless, but the real step that arrived, with its gust, gives 17.5 %. A label has to be the command of its own frame, which is the step that follows it, and nothing else.

The inverse models of GR00T N1 (NVIDIA, 2025) and DreamGen (Jang et al., 2025) take two images, the current frame and a later one. VPT's model sees 128 frames, past and future, but the policy it labels for is a causal transformer with memory. A policy that reads one position may not be given labels that know the past; the 13-frame filter is the right label only for a policy that reads the past too.

6 · The column that is not there

A plate sits on the table, 15 cm below where the cup starts. The arm must descend and push; the table stops the cup, and the plate clicks when the arm has pushed down at a commanded 8 cm/s or more for half a second. The demonstrator descends at 15 cm/s and keeps pushing at 15 cm/s for the rest of a 3.5 s clip, and clicks the plate in 100 of 100 runs. In the footage 70.3 % of the frames are the cup sitting on the table.

Label the footage with the inverse model and train: the policy presses the plate in 0.0 % of 200 runs at 1, 2, 5 and 16 clips. The model does what it was fitted to do. On the table the step is zero, so the command that produced it is, for the plant it was fitted on, zero: it infers a push of 0.0 cm/s where the demonstrator gave 15.0. The same footage with the commands the demonstrator really gave trains a policy that presses in 100.0 % of runs. The column that separates the two is not in the footage, and no better model recovers it. A demonstrator who pushes at 30 cm/s and one who pushes at 15 leave footage that differs by at most 1.2 mm anywhere. Take a plate that needs 20 cm/s: the first clicks it in 100 of 100 runs, the second in 0. Two outcomes, one footage.

Effort, how hard the arm pushed, is one thing a change of position cannot carry; force, what the table pushed back, is another; intent is a third, since a hand that stops at the plate may be pressing, waiting or done. The same holds for a peg in a hole with a millimetre of clearance: the approach that goes in and the one that jams can differ by a fraction of a millimetre, and UMI (Chi et al., 2024) recovers the gripper's position from a wrist camera to 6.1 mm against motion capture, six times that clearance.

What this lesson did not do
It used one body, one camera and exact positions: a camera that reports each position to 2 mm per axis adds 5.7 cm/s of noise to a two-frame label, and what a camera cannot resolve is lesson 10's subject. The footage's hand is the same expert with a bigger wobble; other people, styles and bodies bring lesson 4's several answers and lesson 8's units. It has one course, so it cannot say how many hours cover how many situations (lesson 7). Its inverse model is one number; a real one reads pixels, and published results keep labelled robot data in the mixture (LAPA, Ye et al., 2025, still fine-tunes on 450 labelled robot trajectories). It did not weigh pseudo-labels against real ones in one mixture, and it gave the policy no force to feel and no way to yield (lesson 10).

Common mistakes / failure modes

"the better the inverse model scores, the better the policy"
The 13-frame filter errs by 19.1 % and reaches 50.0 %; the next step errs by 56.5 % and reaches 84.0 % (§5).
"noise in the pseudo-labels is the enemy"
The next step is among the noisiest labels and the best of them: regression averages noise away. Error that correlates with position is what spoils a label (§5).
"labelled hours must grow with the footage"
One demonstration fixes ĝ to ±1.4 % and 80 clips give 85.0 % against 16's 84.0 % (§2, §4); a pixel model needs hours.
"if the inverse model is accurate on frames, nothing is missing"
It infers a push of 0.0 where 15.0 cm/s was given, and the plate is never pressed (§6).

Checkpoint exercise

Try it
The next-step label has a sideways error of 5.9 cm/s per frame, independent from frame to frame. (a) The kernel averages 100 frames: how much label noise is left? (b) The step that arrived adds a push away from the path, c = sv² Δt / sδ², with sδ = 2.15 cm and Δt = 0.05 s: compute c. (c) With a true-label pull-back of 0.15 per second, what is the pull-back of a policy trained on that label, and which way does an offset move? Answer: (a) 5.9 / √100 = 0.59 cm/s, about 4 % of the 15 cm/s speed. (b) c = (0.059 m/s)² × 0.05 s / (0.0215 m)² = 0.38 per second. (c) 0.15 − 0.38 = −0.23 per second: negative, so every offset grows.

Where this points next

Footage can be turned into demonstrations: two labelled demonstrations (18 s) and 16 clips (2.5 min) take the clone from 54.5 % to 84.0 %. What the model learns is what to do next, not how hard to push. The footage records where the cup went, and the table is where it stops: a push of 15 cm/s and one of 30 cm/s leave footage 1.2 mm apart, the model infers a push of 0.0 cm/s where 15.0 were given, and the policy presses the plate in 0.0 % of runs however many clips it is shown. An insertion with a millimetre of clearance is the same case for a camera that cannot resolve a millimetre. What must a policy sense, and how must it act, when the task is decided by something the camera cannot see?

Takeaway
Footage has the state and no commands, but a command causes the next step, so it can be read back: the step divided by ĝΔt, with ĝ measured on a few labelled seconds. A label is priced not by its error, 56.5 % for the best and 19.1 % for a worse one, but by the pull-back it leaves the policy, which a label that knows the step that arrived turns from +0.15 to −0.22 per second. Used right, two labelled demonstrations and 16 clips reach 84.0 % against 54.5 %; labelling the footage with the policy itself gives nothing (48.5 %). What the plant does not turn into motion is not in the footage: from labels that say 0.0 cm/s where 15.0 were given, the plate is pressed in 0.0 % of runs. Judge pseudo-labels by the policy they make.

Interview prompts

Companion reads: World Models · 23 Inferring the actions (the same recovery from the world-model side), World Models · 14 Latent actions (codes learned without labels) and World Models · 11 Video world models (predicting the next frame).