Watching: learning from video without actions
Lesson 8 pooled demonstrations from other arms by writing every action as the place the hand should go. That pool is small beside the video people have shot of themselves at work, and video shows where things went but has no action column: a regressor has nothing to fit. This lesson recovers the column. The cup's next position is the consequence of the command, so one number measured on a few labelled seconds, what the plant does per unit command, turns each pair of frames into a pseudo-command. Two labelled demonstrations and 2.5 minutes of footage take a clone from 54.5 % to 84.0 %. What decides this is which label is used, and it is not the one with the smallest error. What no label recovers is a push: a push of 15 cm/s and one of 30 cm/s on a table leave footage 1.2 mm apart.
New idea: an action can be read back from the frame after it by inverting the plant, so a few labelled seconds label any amount of footage, if each label is the command of its own frame. A label is priced by the pull-back it leaves the policy (how hard it pulls the cup back toward the path), not by its error.
Forces next: A model that infers the missing actions from a few hours of labelled motion can turn a large pile of video into demonstrations, and what it learns is what to do next, not how hard to push. Video records where things moved, never the force that moved them: an insertion with a millimetre of clearance, watched through a camera that cannot resolve a millimetre, looks the same when it succeeds and when it jams. What must a policy sense, and how must it act, when the task is decided by something the camera cannot see?
1 · Two kinds of hour
Ego4D (Grauman et al., 2022) holds 3,670 hours of daily-life video; besides pixels it carries narrations and benchmark labels, and no joint command, end-effector command or force. The real-robot hours GR00T N1 (NVIDIA, 2025) pooled from ten sources come to 3,288.8, and video models are trained on far more: Cosmos (NVIDIA, 2025) collected about 20 million hours of raw video, more than a thousand times the robot hours added up in lesson 8. A robot hour has a state and a command in every frame and is scarce; a video hour has the state and is plentiful.
On the Bench, footage is the expert's runs of the five-post course with the commands thrown away: the cup's position at 20 Hz and nothing else. A labelled run keeps them. The robot owns two labelled demonstrations (365 frames, 18 s). The footage is shot by a hand that wobbles twice as much as the arm: its gust is 0.10 rad/s against the arm's 0.05, so it wanders into states the two demonstrations never visit. One clip is one run, 9.2 s. As in lesson 8 the policy reads the cup's position and answers with the velocity it wants for the cup; the arm's own controller turns that into joint velocities.
A policy is a table of (position, command) pairs (lesson 1), and the footage has no commands: fed 16 clips it has nothing to store. The two labelled demonstrations alone make a clone that reaches the mat in 54.5 % of 200 runs (interval 47.6 to 61.3 %), and more of the robot's own demonstrations help slowly: 63.0 % with 8, 69.0 % with 200. The footage has the coverage the demonstrations lack, and it lacks the label. How does the label come back?
2 · The effect shows the cause
Write zt for the cup's position at frame t as the footage shows it (m), vt for the cup velocity commanded at that frame (m/s, the missing column), Δt = 0.05 s for the frame time, nt for the gust's effect on the cup during the step (m) and g for what the plant does per unit command (1 on the Bench, unknown to the learner). The arm turns a wanted cup velocity into joint velocities, the plant adds the gust, and the cup moves:
zt+1 − zt = g · Δt · vt + nt
The command is the cause of the step, so inverting the plant reads it back:
v̂t = (zt+1 − zt) / (ĝ · Δt)
That is the whole inverse-dynamics model. It has one number, ĝ, and the labelled runs measure it by regressing the realised velocity on the logged command: ĝ = Σ (Δz/Δt) · v / Σ v · v. The two labelled demonstrations give ĝ = 0.997. The label's error is the gust, nt / (g Δt): 5.9 cm/s per axis on the footage, against a speed of 15 cm/s. Relative to the size of the commands the error is 27.8 % on labelled frames and 56.5 % on the footage, twice as much because the hand wobbles twice as much.
What the labelled hours buy is ĝ and the alignment. Resample the labelled data: from one demonstration ĝ is good to ±1.4 %, from sixteen to ±0.4 %, falling as n−1/2 (fitted exponent −0.47), as a sum of squares predicts. How good must it be? Scale every label by a factor and train: 0.8 reaches 84.0 %, and 0.7 reaches 0.0, because a policy that moves too slowly never finishes in the steps it is given. So ĝ has to be right to better than a fifth, and one demonstration gives it to about a hundredth. The labelled frames also say which pair of frames carries the command: of the alignments from two frames early to three late, the best fit is the pair (t, t+1) in 60 of 60 resamples of one demonstration.
Real footage is harder: the labelled hours scale with what the model has to learn, not with the footage. VPT (Baker et al., 2022) trained a model of about half a billion weights on 1,962 hours of labelled Minecraft play, applied it to about 70,000 hours of web video, and found that at least 10 hours of labelled data were needed for any crafting, with gains levelling off after 100 hours.
3 · Three ways to skip the effect
Three tempting shortcuts need no inverse of the plant. Each fails by a computation.
Label the footage with the policy. Let the clone built from the two labelled demonstrations answer for every footage frame and retrain on its answers. The labels hold only what it already says: 0.0 % of the footage frames lie beyond its reach of three bandwidths, so it answers all from what it stores. The retrained policy reaches 48.5 % against the clone's 54.5 %, and its labels err by only 9.6 % of the command. A small label error is no sign that a label carries anything new.
The widget
What to try. At the defaults (16 clips, two labelled demonstrations, label = the next step) the policy reaches the mat in 84.0 % of its runs (interval 78.3 to 88.4 %); the demonstrations alone reach 54.5 % and the true commands 93.5 %; the label error is 56.5 % and the pull-back +0.16 per second. One clip (9.2 s) already gives 79.5 %. The step that arrived has an error of 59.6 %, as large, and reaches 17.5 % with a pull-back of −0.22. The filter fitted over 13 frames has an error of 19.1 % and reaches 50.0 %; the policy's own answer, 9.6 % error, reaches 48.5 %. Footage as steady as the arm: the next step reaches 70.0 %, the true commands 71.5 %. Press the plate: the model-labelled policy presses it in 0.0 % of runs at every footage setting, the true commands in 100.0 %, and the model infers a push of 0.0 where 15.0 cm/s was given.
4 · What the footage buys
Two labelled demonstrations and 16 clips with inferred commands reach 84.0 % where the demonstrations alone reach 54.5 %: a lift of 29.5 points. With the commands the footage's hand really gave, 93.5 %, so the inverse model costs 9.5 points. Almost all of the lift is there after one clip: 0 clips 54.5 %, 1 clip 79.5 %, 16 clips 84.0 %, and 80 clips (12.2 minutes) 85.0 %, against 92.0 % for the true commands. The Bench has one course, which one clip already covers. Further clips close the gap to the true commands slowly, as the regression averages the label's noise: 9.5 points at 16 clips, 7.0 at 80.
What the footage supplies is coverage: states the robot's tidy demonstrations never visit, with the recovery commands the demonstrator gave there (lesson 2). That is why 200 labelled demonstrations of the same kind reach only 69.0 %, and why footage as steady as the arm gives only 71.5 % even with the true commands. In practice what matters is the number of situations the hours cover (lesson 7's coverage law), not their length.
5 · Score a label by its pull-back
The next step has a label error of 56.5 %. The filter fitted over 13 frames, the label a practitioner would fit by regression on labelled data, has 19.1 % on the footage (9.7 % on held-out labelled frames). Each policy was trained on 16 clips labelled one way:
| label of frame t | label error | pull-back (per s) | reaches the mat |
|---|---|---|---|
| the true commands | 0 % | +0.15 | 93.5 % |
| the next step, (zt+1 − zt) / ĝΔt | 56.5 % | +0.16 | 84.0 % |
| no footage (two labelled demonstrations) | — | −0.01 | 54.5 % |
| the average of the next step and the one that arrived | 41.2 % | −0.03 | 55.0 % |
| a filter fitted over 13 frames | 19.1 % | −0.01 | 50.0 % |
| the policy's own answer | 9.6 % | −0.12 | 48.5 % |
| the step that arrived, (zt − zt−1) / ĝΔt | 59.6 % | −0.22 | 17.5 % |
The rows are not ordered by error: the next step, whose error is the second largest, beats the policy's own answer, whose error is the smallest of the estimated labels, and the step that arrived, with about the next step's error, loses 66.5 points to it. Something else decides, and it shows near the path. Let δ be the cup's sideways offset from the path and m(δ) the policy's sideways command there. The loop is δ̇ = m(δ) + gust. If m(δ) = −pδ the offset is pulled back at rate p and wanders about sv√(Δt/2p), with sv the gust's spread as a velocity at the cup; if p < 0 it grows as e|p| t. Call p the pull-back. The expert's is 1.13 per second; a kernel of 1.5 cm smooths across the offsets that carry the feedback and keeps 0.15, which is all that holds the cup to the path.
A label error e that correlates with the offset, E[e | δ] = cδ, takes c away from that: ppolicy = p − c. The next step's error is the gust of a step that has not happened, so it cannot know where the cup is: c = 0, and the pull-back stays at 0.16. The step that arrived contains the gust that put the cup at δ. The offset it arrived at is the old offset plus Δt times that gust, so with sv = 5.9 cm/s per axis and an offset spread of sδ = 2.15 cm the best linear guess of the gust from the offset is E[gust | δ] = (sv² Δt / sδ²) δ, and c = 0.38 per second. The loss measured in the table is 0.15 − (−0.22) = 0.37, and the average of the two steps loses half of it, 0.18. A pull-back of −0.22 per second turns an offset of 1 cm into more than 10 cm over a 12-second course. The policy is a function of position, so its labels may not know where the cup came from.
Two controls separate the effects. Take the true command of the frame s later, with no gust in it: s = +1 gives 66.0 %, +2 gives 25.0 %, +3 gives 0.0 %, because a command from the future belongs to a place further along. The command of the frame before gives 96.5 % (two before, 84.5 %): a late label is harmless, but the real step that arrived, with its gust, gives 17.5 %. A label has to be the command of its own frame, which is the step that follows it, and nothing else.
The inverse models of GR00T N1 (NVIDIA, 2025) and DreamGen (Jang et al., 2025) take two images, the current frame and a later one. VPT's model sees 128 frames, past and future, but the policy it labels for is a causal transformer with memory. A policy that reads one position may not be given labels that know the past; the 13-frame filter is the right label only for a policy that reads the past too.
6 · The column that is not there
A plate sits on the table, 15 cm below where the cup starts. The arm must descend and push; the table stops the cup, and the plate clicks when the arm has pushed down at a commanded 8 cm/s or more for half a second. The demonstrator descends at 15 cm/s and keeps pushing at 15 cm/s for the rest of a 3.5 s clip, and clicks the plate in 100 of 100 runs. In the footage 70.3 % of the frames are the cup sitting on the table.
Label the footage with the inverse model and train: the policy presses the plate in 0.0 % of 200 runs at 1, 2, 5 and 16 clips. The model does what it was fitted to do. On the table the step is zero, so the command that produced it is, for the plant it was fitted on, zero: it infers a push of 0.0 cm/s where the demonstrator gave 15.0. The same footage with the commands the demonstrator really gave trains a policy that presses in 100.0 % of runs. The column that separates the two is not in the footage, and no better model recovers it. A demonstrator who pushes at 30 cm/s and one who pushes at 15 leave footage that differs by at most 1.2 mm anywhere. Take a plate that needs 20 cm/s: the first clicks it in 100 of 100 runs, the second in 0. Two outcomes, one footage.
Effort, how hard the arm pushed, is one thing a change of position cannot carry; force, what the table pushed back, is another; intent is a third, since a hand that stops at the plate may be pressing, waiting or done. The same holds for a peg in a hole with a millimetre of clearance: the approach that goes in and the one that jams can differ by a fraction of a millimetre, and UMI (Chi et al., 2024) recovers the gripper's position from a wrist camera to 6.1 mm against motion capture, six times that clearance.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
Footage can be turned into demonstrations: two labelled demonstrations (18 s) and 16 clips (2.5 min) take the clone from 54.5 % to 84.0 %. What the model learns is what to do next, not how hard to push. The footage records where the cup went, and the table is where it stops: a push of 15 cm/s and one of 30 cm/s leave footage 1.2 mm apart, the model infers a push of 0.0 cm/s where 15.0 were given, and the policy presses the plate in 0.0 % of runs however many clips it is shown. An insertion with a millimetre of clearance is the same case for a camera that cannot resolve a millimetre. What must a policy sense, and how must it act, when the task is decided by something the camera cannot see?
Interview prompts
- How do you get actions out of video that has none? (§2 — the command causes the next step, so divide the step by the plant's gain, which a few labelled seconds measure.)
- An inverse model has 19.1 % error and a cruder one 56.5 %. Which is the better teacher, and how would you find out? (§5 — train a policy on each: the crude one reaches 84.0 %, the accurate one 50.0 %.)
- Why is pseudo-label noise cheap and a label that knows the previous step expensive? (§5 — regression over positions averages noise away, but error correlated with position subtracts from the pull-back, 0.15 to −0.22 per second.)
- Why not label the footage with the policy you already have, or learn latent actions with no labels? (§3 — the first adds no information, 48.5 % against 54.5 %; the second is defined only up to an invertible map that labelled frames must fix.)
- How many labelled hours does an inverse model need? (§2 — it depends on what it must learn: seconds for one gain on the Bench, 100 hours in VPT's Minecraft, where it reads pixels; not on how much footage there is.)
- What can a policy never learn from footage, even with a perfect inverse model? (§6 — whatever the plant does not turn into motion: a push of 15 and one of 30 cm/s leave footage 1.2 mm apart.)
Companion reads: World Models · 23 Inferring the actions (the same recovery from the world-model side), World Models · 14 Latent actions (codes learned without labels) and World Models · 11 Video world models (predicting the next frame).