Moving the bits: the infrastructure bill
Lesson 22 counted the hours that survive, and they are still video. One effective hour from three 1080p cameras at 30 frames per second is 10.8 GB on disk and 1,008 GB decoded, and a model reads tokens, not pixels: an H100 trains on 2,771 camera frames a second where the eight cores beside it decode 960, so it is busy 34.6 % of the time. This lesson counts the pipeline in training tokens: each stage gets a capacity per token, the slowest sets the busy fraction, and the pixels decoded per token (8,100) say which stage that is. The bill, bytes, cores and idle accelerators, is linear in the hours, so it is a fixed share of any budget: 5.3 % of an own hour, 44 % of a twin hour.
New idea: count the pipeline in training tokens: each stage's capacity per token, the slowest sets the accelerator's busy fraction, and the pixels decoded per token say whether that stage is the decoder. Its costs are linear in the hours moved, so the bill is a price per hour like lesson 17's.
Forces next: Moving the data costs a fixed share of the budget and decides whether the processors are ever busy; it is the last bill. We can now price every source, its value, its shelf life, its duplicates and its transport. Put together, which kind of data does embodied training eat the most, by volume and by value, and what should be bought next?
1 · An hour, weighed
Take an hour that lesson 22 left standing. Its format is an assumption: three cameras (two on tripods and one on the wrist, as in DROID, Khazatsky et al., 2024), each 1920 × 1080 at 30 frames per second, stored as H.264 at 8 Mbit/s, the rate YouTube recommends for 1080p uploads at 24 to 30 frames per second. A robot logger sets its own rate, and the byte terms of the bill scale with it.
| per camera | computation | result |
|---|---|---|
| stored stream | 8 Mbit/s ÷ 8 | 1.000 MB/s, 3.60 GB an hour |
| decoded frame | 1920 × 1080 × 1.5 bytes (YUV 4:2:0: a brightness plane and two colour planes of a quarter the pixels) | 3,110,400 bytes |
| decoded stream | 3,110,400 bytes × 30 frames per second | 93.3 MB/s, 335.9 GB an hour |
| decoded ÷ stored | 12 bits per pixel ÷ 0.129 bits of stream per pixel | 93.3 |
Three cameras: 10.8 GB stored and 1,008 GB decoded an hour, in 324,000 frames; a corpus of 1,000 effective hours, the scale of this lesson's illustrations, is 10.8 TB stored and 1.008 PB decoded. DROID's stereo MP4 video is 8.7 TB for 350 hours: 24.9 GB per recorded hour from six 720p streams, 0.67 bits per pixel, at most 5.2 times the rate assumed here (failed episodes sit outside the 350 hours): the stored side is, if anything, light, and the decoded side is fixed by the pixels. Both are bytes, and the accelerator does not train on bytes: one H100 trains a 93 M model on this hour in 116.9 s, and the eight cores beside it need 337.5 s to decode it (§2 and §3 derive both). What does the accelerator eat, and how fast?
2 · What the accelerator eats
A model reads tokens. It resizes each camera frame to the side it was built for, 224 pixels here, and cuts it into patches of 14 × 14: floor(224 / 14)² = 16² = 256 tokens. At 384 pixels the same patch gives floor(384 / 14)² = 27² = 729, since 384 / 14 = 27.43 is floored, not rounded to 28. The decoder pays for the whole 1080p frame; the model keeps 256 tokens of it. Three cameras at 30 frames per second make 324,000 × 256 = 82.9 M tokens an hour.
A token passes through the P weights of the model. Going forward each weight does a multiply and an add, 2 FLOPs; going backward costs twice the forward pass (one pass for the gradient of the input, one for that of the weights). That is 6P FLOPs per token; attention across 256 tokens adds a few per cent and is ignored. An accelerator with dense BF16 peak Φ (A100: 312 TFLOPS, its datasheet; H100 SXM: 989.5, half of the 1,979 NVIDIA's page prints with sparsity) that reaches a share μ = 0.4 of it (assumed; only P / μ enters) trains, at n tokens a frame, on
D = μ Φ / (6 P n) frames per second.
For P = 93 M: 558 MFLOPs per token, 709,319 tokens a second, 2,771 frames a second on an H100.
| model (every image token through every weight) | P | frames per second, one H100 | H100 time for 324,000 frames |
|---|---|---|---|
| a small visuomotor network (assumed) | 25 M | 10,307 | 31 s |
| Octo-Base (Octo Model Team, 2024) | 93 M | 2,771 | 117 s |
| SigLIP so400m (Zhai et al., 2023): 27 layers of 4 × 1152² + 2 × 1152 × 4304 = 15.22 M weights | 411 M | 627 | 517 s |
| the ViT-g encoder of V-JEPA 2 (Assran et al., 2025), "over 1 billion" | 1 B | 258 | 1,257 s |
| a vision-language-action model of OpenVLA's size (Kim et al., 2024) | 7 B | 37 | 2.44 h |
This is demand: the rate at which the machine would train if nothing made it wait. The smaller the model, the more it asks of what feeds it.
3 · What each token costs the decoders
A decoder works in pixels: to show a frame of an inter-coded stream it rebuilds the picture from the ones before it, every pixel, whatever the model keeps. To first order its cost is proportional to the pixels. One core, counting colour conversion and the resize, decodes 120 frames per second of 1080p, 248.8 Mpx/s (assumed: nothing this lesson relies on prints a measured figure). One NVDEC engine, the decoder on the accelerator, decodes 771 frames per second of 1080p H.264 in NVIDIA's application note: "indicative", per engine, in the Turing row that the A100 and H100 follow scaled by their video clocks. An H100 has 7 engines and an A100 5: at most 5,397 and 3,855 frames per second, upper bounds that add up only across simultaneous sessions.
The decoders pay for pixels and the accelerator eats tokens, so the exchange rate of the pipeline is
π = g W H / n pixels decoded per token,
for a stored frame of W × H, n tokens kept and g frames decoded for each frame used. Reading in order g = 1. A random frame of a stream with a keyframe every G frames needs the decoder to start at the last keyframe: (G + 1) / 2 = 15.5 frames on average for G = 30. At 1920 × 1080 and 256 tokens π = 8,100; a frame stored at the model's own 224 × 224 gives 14² = 196; random frames give 125,550.
Every stage can now be put in the accelerator's unit: its rate divided by its units per token is its capacity in tokens per second (a frame is 256 tokens). A token costs the decoder 1.5 × 8,100 = 12,150 bytes of pixels and the disk 130 bytes of stream, a 33,333-byte frame ÷ 256.
| stage, one accelerator | its rate | per token | tokens per second |
|---|---|---|---|
| accelerator (H100 at 40 %) | 395.8 TFLOPS | 6P = 558 MFLOPs | 709,319 |
| decoders, 8 cores | 1,990.7 Mpx/s | π = 8,100 pixels | 245,760 |
| read link, encoded video | 1.5 GB/s | 130 bytes | 11,520,000 |
| read link, decoded frames | 1.5 GB/s | 12,150 bytes | 123,457 |
Reading encoded video the slowest stage is the decoders, 245,760 tokens a second against the accelerator's 709,319: busy 34.6 %. Reading decoded frames it is the link, 123,457: busy 17.4 % (§5).
How many cores. Little's law (Little, 1961): in a stable system the average number of items inside is their arrival rate times the time each stays, L = λτ, whatever the distributions. To keep the accelerator at full speed frames must enter the decoders at the rate it takes them, λ = D = 2,771 a second, and each stays τ = 1 / 120 s = 8.33 ms in a core. So L = 23.1 decoders are working: that is the cores one accelerator needs, D g / (frames per second of one core at this size). With c cores the decoders give c × 120 = 960 frames per second; with S the supply of the slowest stage, here the decoders', the accelerator is busy for
busy = min(1, S / D) = 960 / 2,771 = 34.6 %.
| cores needed per H100, frames in order | 480p | 1080p | 2160p |
|---|---|---|---|
| 25 M | 17.0 | 85.9 | 343.6 |
| 93 M | 4.6 | 23.1 | 92.4 |
| 411 M | 1.0 | 5.2 | 20.9 |
| 1 B | 0.4 | 2.1 | 8.6 |
The baton's claim becomes a condition: a model that trains on video spends its time decoding it when it has fewer cores than D g / (core rate). At 1080p, in order, with 8 cores, that is every model below 268 M parameters: one of Octo-Base's size is below it and a 7 B model is far above. 4K raises the threshold to 1,074 M and random frames to 4,160 M. An H100's NVDEC engines lift the 93 M model's supply to 6,357 frames per second, enough at 1080p; the 25 M model reaches 61.7 % and the 93 M model at 4K 57.4 %.
4 · The pipeline as a queue
The formula is a claim about a queue, so the widget runs one. A read stage streams a frame every 1 / (read rate) seconds. A decoder takes the next frame, decodes it for g W H / (its rate) seconds and puts it in a buffer of 64 frames, waiting if the buffer is full. The accelerator takes a frame from the buffer and trains on it for 1 / D seconds. The simulated busy fraction agrees with the formula's to about a point in every state of this lesson.
It runs on these constants. Only their products enter, so the cores slider stands for cores × core speed, the model selector for P / μ, and the price selector for the byte prices.
| assumption | value | status |
|---|---|---|
| stream; model input | 8 Mbit/s at 1080p30, 0.129 bits per pixel per frame at any size; 224 × 224 in 14 × 14 patches, 256 tokens | assumed; published (SigLIP so400m) |
| dense BF16 peak; share reached | H100 989.5, A100 312 TFLOPS; μ = 0.4 | published; assumed |
| decoders; read link | 8 cores per accelerator, 120 frames per second of 1080p each; an NVDEC engine 771, 7 on an H100; 1.5 GB/s per accelerator (100 Gbit/s shared by eight) | assumed; the engine published, an upper bound |
| prices; epochs | $0.02 per GB-month for 12 months, $0.02 per GB read, $0.04 per core-hour, $2.50 per accelerator-hour; 10 epochs | assumed; the accelerator's is the ledger's |
The widget
What to try. Leave the defaults: three 1080p cameras at 30 fps, a 93 M model, 8 cores, frames in order. The accelerator wants 2,771 frames a second and the cores give 960: busy 34.6 %, decode binds, 23.1 cores are needed, each token costs 8,100 decoded pixels, and an hour costs $6.58 to move, 5.3 % of an own hour. At 2160p: 43.2 GB an hour, 92.4 cores needed, busy 8.7 %; at 480p: 4.6 needed, busy 100 %, 2.13 GB, $1.04. Back at 1080p, 23 cores give 99.6 % and 24 give 100 %, and the bill falls to $5.06. With the 411 M model 5.2 cores are needed, busy 100 %; with 25 M, 85.9 needed, busy 9.3 %, and 61.7 % with the card's NVDEC. Decoded frames on disk: 1,008 GB, busy 17.4 % with the read stage binding, bill $447.9. Random frames: 357.9 cores needed, busy 2.2 %. Byte prices ten times higher: $49.4, idle accelerators 3 % of it; a tenth: $2.31, 66 %.
5 · Three roads, rejected by computation
Three ways out of a stall are tempting. Each moves one number, and each fails by arithmetic already done.
| option | what it costs | what goes wrong |
|---|---|---|
| decode once, keep the pixels | storage × 93.3: $241.9 a year for one hour, 196 % of its $123.61 price | it saves $0.03 of decoding an epoch (break-even 665 epochs a month), and the stall moves: the link gives 1.5 GB/s of the 8.6 wanted (3.11 MB × 2,771 frames a second), busy 17.4 % |
| lowering: halve the frame rate | half the frames of every second of motion; 5.4 GB and $3.29 an hour | busy stays 34.6 %: each frame costs the same decode and FLOPs, so a rate changes the bill and never the busy fraction |
| lowering: store 480p | bytes ÷ 5.06; 4.6 cores needed, busy 100 % | a feature under two pixels is taken as lost: a 1 mm clearance in a view 500 mm wide spans 3.8 pixels at 1080p, 2.6 at 720p, 1.7 at 480p, 0.4 at 224 (lesson 10's peg is decided below the sensor's resolution), so the task sets the size, here at least 1,000 pixels across, and not the decoder |
| lowering: make the model read 112 pixels | a coarser view, tokens 256 → 64 | the card trains on 4 times the frames, 11,083 a second, so 92.4 cores are needed and it is busy 8.7 % |
| more accelerators on the same cores | 4 cores each instead of 8: busy 17.3 % | an epoch takes as long (the decoders set the rate) and the accelerators' bill for it doubles, $0.234 to $0.469 an hour of data; a faster card is the same move: with 8 cores an A100 is busy 100 % (7.3 needed), an H100, which trains 3.17 times the frames, 34.6 % |
Cores to the need are cheap: the 24 above cost $0.01 more over ten epochs, remove the $1.53 of idle accelerator time (the bill falls by $1.52) and bring an epoch over 1,000 hours on 8 accelerators from 11.7 to 4.1 hours.
6 · The bill, per hour and per source
| per effective hour, three 1080p cameras, 10 epochs | computation | dollars |
|---|---|---|
| storage, 12 months | 10.8 GB × $0.02 × 12 | 2.59 |
| network | 10 × 10.8 GB × $0.02 | 2.16 |
| cores | 10 × 8 × 337.5 s × $0.04 / 3600 | 0.30 |
| idle accelerators | 10 × (337.5 − 116.9) s × $2.50 / 3600 | 1.53 |
| the pipeline | 6.58 | |
| the learning itself, not a transport cost | 10 × 116.9 s × $2.50 / 3600 | 0.81 |
At these prices bytes are 72 % of the bill, the stall 23 % and the cores 5 %; the split moves with the prices, which are assumptions (the last selector).
A fixed share. Every term is linear in the hours moved: bytes in the hours, accelerator-seconds in the frames, cores in the accelerator-seconds. So t = $6.58 is a price per effective hour whatever the number of hours. A budget B spent on hours priced p buys B / p hours and moving them costs (B / p) t, a share t / p of B whatever B is: $100 buys 0.81 own hours and $5.33 of moving, $100,000 buys 809 and $5,326, both 5.33 %. Lesson 18's rule, equal marginal value per dollar, holds with ps + ts in the place of ps, and lesson 17's cost per useful hour becomes (ps + ts) / ρs, with ρs the exchange rate of lesson 16.
By kind of hour. The same arithmetic with each kind's format (an assumption for the last three), against lesson 17's price:
| kind of hour, format | stored GB an hour | pixels per token | moving it, $ an hour | lesson 17's price | share |
|---|---|---|---|---|---|
| own robot, 3 × 1080p at 30 fps | 10.8 | 8,100 | 6.58 | $123.61 | 5.3 % |
| twin arm, the same | 10.8 | 8,100 | 6.58 | $15.00 | 44 % |
| older arm, 2 × 480p at 15 fps | 0.71 | 1,601 | 0.35 | $15.00 | 2.3 % |
| simulator, rendered at the model's size, not stored | 0 | none | 0 | $0.1875 | 0 % |
| force-bearing demonstration: own format and a force channel of 6 × 4 bytes at 1 kHz | 10.9 | 8,100 | 6.62 | $214.44 | 3.1 % |
The twin's hour costs as much to move as the own robot's, 44 % of its price against 5 %. The force channel adds 0.086 GB, 0.8 % of a camera hour's bytes, to the dearest hour of the ledger. The older arm's small frames fit 8 cores (4.6 needed) and never stall. A simulated hour moves nothing, but stored as video it would cost 35 times its price to move. By bytes the kinds rank force-bearing hour, own = twin, older arm, simulator; by price, force-bearing hour, own, twin = older arm, simulator; by the share transport takes, twin, own, force-bearing hour, older arm.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
Moving an hour costs bytes, cores and the time an accelerator spends waiting: $6.58 an hour at the defaults, with the accelerator busy 34.6 % of the time, and a share that holds whatever is bought, 5.3 % of an own hour and 44 % of a twin hour. The roads that look free fail by the same arithmetic, and cores to the need cost cents. What no column can do alone is rank the kinds of data: a camera hour is 10.8 GB and its force channel adds 0.086, the hour that carries it is the dearest at $214.44, and transport takes 44 % of a twin hour. We can now price every source, its value, its shelf life, its duplicates and its transport. Put together, which kind of data does embodied training eat the most, by volume and by value, and what should be bought next?
Interview prompts
- Why can an accelerator be a third busy with all its data on disk? (§3 — 8,100 decoded pixels per token: 8 cores give 960 frames a second against the 2,771 it wants, 34.6 %.)
- How many cores must feed one accelerator, and where does the formula come from? (§3 — Little's law: arrival rate D times the time in a core, 2,771 × 8.33 ms = 23.1.)
- Why does a smaller model input make the decoders' work per token larger? (§5 — 64 tokens a frame instead of 256 means 4 times the frames a second: 92.4 cores needed, busy 8.7 %.)
- Why is pre-decoding the corpus a bad trade, and what is its break-even? (§5 — storage × 93.3 to save $0.03 of decoding an epoch: 665 epochs a month.)
- The transport bill is a fixed share of the budget: show it. (§6 — B / p hours cost t each to move, a share t / p whatever B: 5.3 % of an own hour, 44 % of a twin hour.)
- What does the NVDEC figure of 771 frames a second say, and what does it not? (§3 — per engine, indicative, 1080p H.264, the Turing row scaled by an unprinted clock; 5 and 7 engines make upper bounds of 3,855 and 5,397.)
Companion reads: System ML · 23 Profiling, MFU and finding the bottleneck (the share of peak a run reaches), System Design · 02 Latency, throughput and queueing (Little's law) and Lesson 10 · What cameras cannot see (a task decided below the sensor's resolution).