all_lessons/Robot Model Training/23 · Moving the bitslesson 23 / 24

Moving the bits: the infrastructure bill

Lesson 22 counted the hours that survive, and they are still video. One effective hour from three 1080p cameras at 30 frames per second is 10.8 GB on disk and 1,008 GB decoded, and a model reads tokens, not pixels: an H100 trains on 2,771 camera frames a second where the eight cores beside it decode 960, so it is busy 34.6 % of the time. This lesson counts the pipeline in training tokens: each stage gets a capacity per token, the slowest sets the busy fraction, and the pixels decoded per token (8,100) say which stage that is. The bill, bytes, cores and idle accelerators, is linear in the hours, so it is a fixed share of any budget: 5.3 % of an own hour, 44 % of a twin hour.

The thesis, here
An hour of video is moved in pixels and consumed in tokens. Count each stage in tokens per second, its own rate divided by its own units per token: the machine runs at the slowest stage, the accelerator is busy for min(1, supply / demand) of its time, and the pixels the decoders must make for each token say where it waits. Bytes, cores and accelerator-seconds are linear in the hours moved, so moving an hour costs a fixed amount t, a share t / p of whatever is spent on hours priced p.
Linear position
Forced by: Counting effective hours instead of recorded ones shrinks a corpus, sometimes by orders of magnitude, and the hours that survive still have to reach the accelerators: a model that trains on video spends its time decoding it. If the processors wait for the data, the data budget is not the only bill. What does it cost to move the effective hours, and where does the machine stall?
New idea: count the pipeline in training tokens: each stage's capacity per token, the slowest sets the accelerator's busy fraction, and the pixels decoded per token say whether that stage is the decoder. Its costs are linear in the hours moved, so the bill is a price per hour like lesson 17's.
Forces next: Moving the data costs a fixed share of the budget and decides whether the processors are ever busy; it is the last bill. We can now price every source, its value, its shelf life, its duplicates and its transport. Put together, which kind of data does embodied training eat the most, by volume and by value, and what should be bought next?
The plan
Six moves. (1) Weigh an hour. (2) Count what the accelerator eats, in tokens. (3) Count what each token costs the decoders. (4) Run the pipeline as a queue. (5) Reject pre-decoding, blind lowering and more accelerators by computation. (6) Add the bill, per hour and per source.

1 · An hour, weighed

Take an hour that lesson 22 left standing. Its format is an assumption: three cameras (two on tripods and one on the wrist, as in DROID, Khazatsky et al., 2024), each 1920 × 1080 at 30 frames per second, stored as H.264 at 8 Mbit/s, the rate YouTube recommends for 1080p uploads at 24 to 30 frames per second. A robot logger sets its own rate, and the byte terms of the bill scale with it.

per cameracomputationresult
stored stream8 Mbit/s ÷ 81.000 MB/s, 3.60 GB an hour
decoded frame1920 × 1080 × 1.5 bytes (YUV 4:2:0: a brightness plane and two colour planes of a quarter the pixels)3,110,400 bytes
decoded stream3,110,400 bytes × 30 frames per second93.3 MB/s, 335.9 GB an hour
decoded ÷ stored12 bits per pixel ÷ 0.129 bits of stream per pixel93.3

Three cameras: 10.8 GB stored and 1,008 GB decoded an hour, in 324,000 frames; a corpus of 1,000 effective hours, the scale of this lesson's illustrations, is 10.8 TB stored and 1.008 PB decoded. DROID's stereo MP4 video is 8.7 TB for 350 hours: 24.9 GB per recorded hour from six 720p streams, 0.67 bits per pixel, at most 5.2 times the rate assumed here (failed episodes sit outside the 350 hours): the stored side is, if anything, light, and the decoded side is fixed by the pixels. Both are bytes, and the accelerator does not train on bytes: one H100 trains a 93 M model on this hour in 116.9 s, and the eight cores beside it need 337.5 s to decode it (§2 and §3 derive both). What does the accelerator eat, and how fast?

2 · What the accelerator eats

A model reads tokens. It resizes each camera frame to the side it was built for, 224 pixels here, and cuts it into patches of 14 × 14: floor(224 / 14)² = 16² = 256 tokens. At 384 pixels the same patch gives floor(384 / 14)² = 27² = 729, since 384 / 14 = 27.43 is floored, not rounded to 28. The decoder pays for the whole 1080p frame; the model keeps 256 tokens of it. Three cameras at 30 frames per second make 324,000 × 256 = 82.9 M tokens an hour.

A token passes through the P weights of the model. Going forward each weight does a multiply and an add, 2 FLOPs; going backward costs twice the forward pass (one pass for the gradient of the input, one for that of the weights). That is 6P FLOPs per token; attention across 256 tokens adds a few per cent and is ignored. An accelerator with dense BF16 peak Φ (A100: 312 TFLOPS, its datasheet; H100 SXM: 989.5, half of the 1,979 NVIDIA's page prints with sparsity) that reaches a share μ = 0.4 of it (assumed; only P / μ enters) trains, at n tokens a frame, on

D = μ Φ / (6 P n) frames per second.

For P = 93 M: 558 MFLOPs per token, 709,319 tokens a second, 2,771 frames a second on an H100.

model (every image token through every weight)Pframes per second, one H100H100 time for 324,000 frames
a small visuomotor network (assumed)25 M10,30731 s
Octo-Base (Octo Model Team, 2024)93 M2,771117 s
SigLIP so400m (Zhai et al., 2023): 27 layers of 4 × 1152² + 2 × 1152 × 4304 = 15.22 M weights411 M627517 s
the ViT-g encoder of V-JEPA 2 (Assran et al., 2025), "over 1 billion"1 B2581,257 s
a vision-language-action model of OpenVLA's size (Kim et al., 2024)7 B372.44 h

This is demand: the rate at which the machine would train if nothing made it wait. The smaller the model, the more it asks of what feeds it.

3 · What each token costs the decoders

A decoder works in pixels: to show a frame of an inter-coded stream it rebuilds the picture from the ones before it, every pixel, whatever the model keeps. To first order its cost is proportional to the pixels. One core, counting colour conversion and the resize, decodes 120 frames per second of 1080p, 248.8 Mpx/s (assumed: nothing this lesson relies on prints a measured figure). One NVDEC engine, the decoder on the accelerator, decodes 771 frames per second of 1080p H.264 in NVIDIA's application note: "indicative", per engine, in the Turing row that the A100 and H100 follow scaled by their video clocks. An H100 has 7 engines and an A100 5: at most 5,397 and 3,855 frames per second, upper bounds that add up only across simultaneous sessions.

The decoders pay for pixels and the accelerator eats tokens, so the exchange rate of the pipeline is

π = g W H / n pixels decoded per token,

for a stored frame of W × H, n tokens kept and g frames decoded for each frame used. Reading in order g = 1. A random frame of a stream with a keyframe every G frames needs the decoder to start at the last keyframe: (G + 1) / 2 = 15.5 frames on average for G = 30. At 1920 × 1080 and 256 tokens π = 8,100; a frame stored at the model's own 224 × 224 gives 14² = 196; random frames give 125,550.

Every stage can now be put in the accelerator's unit: its rate divided by its units per token is its capacity in tokens per second (a frame is 256 tokens). A token costs the decoder 1.5 × 8,100 = 12,150 bytes of pixels and the disk 130 bytes of stream, a 33,333-byte frame ÷ 256.

stage, one acceleratorits rateper tokentokens per second
accelerator (H100 at 40 %)395.8 TFLOPS6P = 558 MFLOPs709,319
decoders, 8 cores1,990.7 Mpx/sπ = 8,100 pixels245,760
read link, encoded video1.5 GB/s130 bytes11,520,000
read link, decoded frames1.5 GB/s12,150 bytes123,457

Reading encoded video the slowest stage is the decoders, 245,760 tokens a second against the accelerator's 709,319: busy 34.6 %. Reading decoded frames it is the link, 123,457: busy 17.4 % (§5).

How many cores. Little's law (Little, 1961): in a stable system the average number of items inside is their arrival rate times the time each stays, L = λτ, whatever the distributions. To keep the accelerator at full speed frames must enter the decoders at the rate it takes them, λ = D = 2,771 a second, and each stays τ = 1 / 120 s = 8.33 ms in a core. So L = 23.1 decoders are working: that is the cores one accelerator needs, D g / (frames per second of one core at this size). With c cores the decoders give c × 120 = 960 frames per second; with S the supply of the slowest stage, here the decoders', the accelerator is busy for

busy = min(1, S / D) = 960 / 2,771 = 34.6 %.

cores needed per H100, frames in order480p1080p2160p
25 M17.085.9343.6
93 M4.623.192.4
411 M1.05.220.9
1 B0.42.18.6

The baton's claim becomes a condition: a model that trains on video spends its time decoding it when it has fewer cores than D g / (core rate). At 1080p, in order, with 8 cores, that is every model below 268 M parameters: one of Octo-Base's size is below it and a 7 B model is far above. 4K raises the threshold to 1,074 M and random frames to 4,160 M. An H100's NVDEC engines lift the 93 M model's supply to 6,357 frames per second, enough at 1080p; the 25 M model reaches 61.7 % and the 93 M model at 4K 57.4 %.

4 · The pipeline as a queue

The formula is a claim about a queue, so the widget runs one. A read stage streams a frame every 1 / (read rate) seconds. A decoder takes the next frame, decodes it for g W H / (its rate) seconds and puts it in a buffer of 64 frames, waiting if the buffer is full. The accelerator takes a frame from the buffer and trains on it for 1 / D seconds. The simulated busy fraction agrees with the formula's to about a point in every state of this lesson.

It runs on these constants. Only their products enter, so the cores slider stands for cores × core speed, the model selector for P / μ, and the price selector for the byte prices.

assumptionvaluestatus
stream; model input8 Mbit/s at 1080p30, 0.129 bits per pixel per frame at any size; 224 × 224 in 14 × 14 patches, 256 tokensassumed; published (SigLIP so400m)
dense BF16 peak; share reachedH100 989.5, A100 312 TFLOPS; μ = 0.4published; assumed
decoders; read link8 cores per accelerator, 120 frames per second of 1080p each; an NVDEC engine 771, 7 on an H100; 1.5 GB/s per accelerator (100 Gbit/s shared by eight)assumed; the engine published, an upper bound
prices; epochs$0.02 per GB-month for 12 months, $0.02 per GB read, $0.04 per core-hour, $2.50 per accelerator-hour; 10 epochsassumed; the accelerator's is the ledger's

The widget

Run the pipeline: what the accelerator wants, what each stage gives
Top: what each stage gives one accelerator in camera frames per second (log axis) against what the accelerator wants (amber line); a red bar short of the line is where the machine waits. Middle: a simulated few milliseconds, time running right, a lane per decoder (the first eight) and one for the accelerator; its gaps are waiting. Third panel: the busy fraction for every model (larger downwards) and stored resolution (larger to the right), your state outlined.
stored per hour
-
decoded per hour
-
pixels decoded per token
-
accelerator wants, frames/s
-
decoders give, frames/s
-
read gives, frames/s
-
cores needed, without NVDEC
-
busy, min(1, S / D)
-
busy, simulated
-
binding stage
-
pipeline bill per effective hour
-
of it, idle accelerators
-
share of an own hour's price
-
read per epoch, 1,000 hours
-
Show the core JS
PL.demand = function (c) { return PL.A.mfu * PL.A.flops[c.acc] / (6 * c.P * PL.tokens(c)); };
PL.rates = function (c) {
  var A = PL.A, px = c.w * c.h, n = PL.tokens(c), D = PL.demand(c), g = c.read === 'rand' ? (A.gop + 1) / 2 : 1;
  var coreFps = A.coreFps * PX1080 / px, engFps = A.nvdecFps * PX1080 / px, eng = c.nvdec ? A.engines[c.acc] : 0;
  var sDec = c.read === 'dec' ? Infinity : (c.cores * coreFps + eng * engFps) / g;
  var bytes = c.read === 'dec' ? A.yuv * px : g * px * A.bpp / 8;
  var sRead = A.linkBps / bytes, S = Math.min(sDec, sRead);
  return { D: D, n: n, g: g, sDec: sDec, sRead: sRead, S: S, busy: Math.min(1, S / D), stage: S >= D ? 'accelerator' : (sRead < sDec ? 'read' : 'decode'),
...
  var secs = b.frames / Math.min(r.D, r.S), useful = b.frames / r.D;
  var cores = c.cores * secs / 3600 * A.coreH, idle = (secs - useful) / 3600 * gpu, learn = useful / 3600 * gpu;
...
    if (accEnd === Infinity && b > 0) { b--; accEnd = t + ta; as.push(t); if (t < W) acc.push([t, t + ta]); }
    for (j = 0; j < m && b < K; j++) if (held[j]) { held[j] = false; b++; }

What to try. Leave the defaults: three 1080p cameras at 30 fps, a 93 M model, 8 cores, frames in order. The accelerator wants 2,771 frames a second and the cores give 960: busy 34.6 %, decode binds, 23.1 cores are needed, each token costs 8,100 decoded pixels, and an hour costs $6.58 to move, 5.3 % of an own hour. At 2160p: 43.2 GB an hour, 92.4 cores needed, busy 8.7 %; at 480p: 4.6 needed, busy 100 %, 2.13 GB, $1.04. Back at 1080p, 23 cores give 99.6 % and 24 give 100 %, and the bill falls to $5.06. With the 411 M model 5.2 cores are needed, busy 100 %; with 25 M, 85.9 needed, busy 9.3 %, and 61.7 % with the card's NVDEC. Decoded frames on disk: 1,008 GB, busy 17.4 % with the read stage binding, bill $447.9. Random frames: 357.9 cores needed, busy 2.2 %. Byte prices ten times higher: $49.4, idle accelerators 3 % of it; a tenth: $2.31, 66 %.

5 · Three roads, rejected by computation

Three ways out of a stall are tempting. Each moves one number, and each fails by arithmetic already done.

optionwhat it costswhat goes wrong
decode once, keep the pixelsstorage × 93.3: $241.9 a year for one hour, 196 % of its $123.61 priceit saves $0.03 of decoding an epoch (break-even 665 epochs a month), and the stall moves: the link gives 1.5 GB/s of the 8.6 wanted (3.11 MB × 2,771 frames a second), busy 17.4 %
lowering: halve the frame ratehalf the frames of every second of motion; 5.4 GB and $3.29 an hourbusy stays 34.6 %: each frame costs the same decode and FLOPs, so a rate changes the bill and never the busy fraction
lowering: store 480pbytes ÷ 5.06; 4.6 cores needed, busy 100 %a feature under two pixels is taken as lost: a 1 mm clearance in a view 500 mm wide spans 3.8 pixels at 1080p, 2.6 at 720p, 1.7 at 480p, 0.4 at 224 (lesson 10's peg is decided below the sensor's resolution), so the task sets the size, here at least 1,000 pixels across, and not the decoder
lowering: make the model read 112 pixelsa coarser view, tokens 256 → 64the card trains on 4 times the frames, 11,083 a second, so 92.4 cores are needed and it is busy 8.7 %
more accelerators on the same cores4 cores each instead of 8: busy 17.3 %an epoch takes as long (the decoders set the rate) and the accelerators' bill for it doubles, $0.234 to $0.469 an hour of data; a faster card is the same move: with 8 cores an A100 is busy 100 % (7.3 needed), an H100, which trains 3.17 times the frames, 34.6 %
Road not taken · decode once, keep the pixels
It deletes the stage that stalls: pay for the decode once and every epoch reads ready frames. The price is the ratio of §1: 93.3 times the bytes, so 1,000 hours grow from 10.8 TB to 1.008 PB, and at the assumed byte prices a year of storage costs more than collecting the hours. The formula allows the other direction: decode is paid in π, so store a copy at the size the model reads (224 × 224: 0.6 cores needed, 0.26 GB an hour, $0.22 against $6.58, 30 times less) and keep the original encoded and cold. The encode of the copy is not priced here, and a model with a larger input needs the original.

Cores to the need are cheap: the 24 above cost $0.01 more over ten epochs, remove the $1.53 of idle accelerator time (the bill falls by $1.52) and bring an epoch over 1,000 hours on 8 accelerators from 11.7 to 4.1 hours.

6 · The bill, per hour and per source

per effective hour, three 1080p cameras, 10 epochscomputationdollars
storage, 12 months10.8 GB × $0.02 × 122.59
network10 × 10.8 GB × $0.022.16
cores10 × 8 × 337.5 s × $0.04 / 36000.30
idle accelerators10 × (337.5 − 116.9) s × $2.50 / 36001.53
the pipeline6.58
the learning itself, not a transport cost10 × 116.9 s × $2.50 / 36000.81

At these prices bytes are 72 % of the bill, the stall 23 % and the cores 5 %; the split moves with the prices, which are assumptions (the last selector).

A fixed share. Every term is linear in the hours moved: bytes in the hours, accelerator-seconds in the frames, cores in the accelerator-seconds. So t = $6.58 is a price per effective hour whatever the number of hours. A budget B spent on hours priced p buys B / p hours and moving them costs (B / p) t, a share t / p of B whatever B is: $100 buys 0.81 own hours and $5.33 of moving, $100,000 buys 809 and $5,326, both 5.33 %. Lesson 18's rule, equal marginal value per dollar, holds with ps + ts in the place of ps, and lesson 17's cost per useful hour becomes (ps + ts) / ρs, with ρs the exchange rate of lesson 16.

By kind of hour. The same arithmetic with each kind's format (an assumption for the last three), against lesson 17's price:

kind of hour, formatstored GB an hourpixels per tokenmoving it, $ an hourlesson 17's priceshare
own robot, 3 × 1080p at 30 fps10.88,1006.58$123.615.3 %
twin arm, the same10.88,1006.58$15.0044 %
older arm, 2 × 480p at 15 fps0.711,6010.35$15.002.3 %
simulator, rendered at the model's size, not stored0none0$0.18750 %
force-bearing demonstration: own format and a force channel of 6 × 4 bytes at 1 kHz10.98,1006.62$214.443.1 %

The twin's hour costs as much to move as the own robot's, 44 % of its price against 5 %. The force channel adds 0.086 GB, 0.8 % of a camera hour's bytes, to the dearest hour of the ledger. The older arm's small frames fit 8 cores (4.6 needed) and never stall. A simulated hour moves nothing, but stored as video it would cost 35 times its price to move. By bytes the kinds rank force-bearing hour, own = twin, older arm, simulator; by price, force-bearing hour, own, twin = older arm, simulator; by the share transport takes, twin, own, force-bearing hour, older arm.

What this lesson did not do
The Bench has no pixels, so this lesson prices the video a model of the sizes above would read; the ratios carry (cores needed, pixels per token, shares), and the totals, those of an illustration of 1,000 hours, do not. Core speed, share of peak, link and prices are assumptions and only their products enter; a small model reaches a smaller share of peak, which lowers the stall. Random access is a stand-in (a keyframe every 30 frames, no per-frame overhead, jitter, stragglers, shared storage or second node), and the NVDEC figures are upper bounds. The bill counts the effective hours stored; a corpus that keeps its duplicates pays the ratio of recorded to effective hours (lesson 22). The formats of the last three kinds are illustrations; setting the transport column beside the others is lesson 24.

Common mistakes / failure modes

"decode once, so the accelerators never wait"
Decoded frames are 93.3 times the stored bytes: a year of one hour costs $241.9, 196 % of its price, and the link binds, busy 17.4 % (§5).
"more accelerators finish the epoch sooner"
On the same cores the epoch takes as long and the accelerators' bill doubles: busy 17.3 % (§5).
"lower the frame rate until the machine stops waiting"
The bill halves and the busy fraction stays 34.6 %: every frame costs the same decode (§5).
"a model that trains on video is decode-bound"
Only when cores < D g / rate: a 411 M model in order at 1080p needs 5.2 cores, a 25 M model 85.9 (§3).

Checkpoint exercise

Try it
An A100 (312 TFLOPS dense) trains a 100 M model on 224 × 224 frames (256 tokens) at 40 % of peak. The data are two 1080p cameras at 30 fps, read in order; a core decodes 120 frames per second of 1080p and the accelerator has 4 cores. (a) How many frames per second does it train on? (b) How many cores does it need, and how busy is it with 4? (c) How long does it spend on one hour of these data, at full speed and as things are, and what does the waiting cost over 10 epochs at $2.50 an hour? Answer: (a) D = 0.4 × 312e12 / (6 × 100e6 × 256) = 812.5 frames per second. (b) 812.5 / 120 = 6.77 cores; 4 cores give 480 frames per second, so it is busy 59.1 %. (c) The hour is 2 × 30 × 3600 = 216,000 frames: 266 s at full speed, 450 s as things are; the 184 s of waiting cost $1.279 over 10 epochs, against $0.089 for three more cores (7 cover the need of 6.77) over the same epochs.

Where this points next

Moving an hour costs bytes, cores and the time an accelerator spends waiting: $6.58 an hour at the defaults, with the accelerator busy 34.6 % of the time, and a share that holds whatever is bought, 5.3 % of an own hour and 44 % of a twin hour. The roads that look free fail by the same arithmetic, and cores to the need cost cents. What no column can do alone is rank the kinds of data: a camera hour is 10.8 GB and its force channel adds 0.086, the hour that carries it is the dearest at $214.44, and transport takes 44 % of a twin hour. We can now price every source, its value, its shelf life, its duplicates and its transport. Put together, which kind of data does embodied training eat the most, by volume and by value, and what should be bought next?

Takeaway
An hour of video is moved in pixels and consumed in tokens, so every stage is compared in tokens per second and the accelerator is busy for min(1, supply / demand) of its time. The decoders pay π = g W H / n pixels per token, 8,100 at 1080p into 256 tokens, so a 93 M model needs 23.1 cores per accelerator and is busy 34.6 % with 8; read in order at 1080p, models below 268 M parameters are decode-bound and larger ones are not. Decoding once costs 93.3 times the bytes, a lower frame rate changes the bill and not the stall, a smaller model input makes the stall worse, and more accelerators on the same cores add idle ones. The bill is linear in the hours, so it is a fixed share t / p of any budget: $6.58 an hour, 5.3 % of an own hour, 44 % of a twin hour. The prices and the core speed are assumptions; the structure carries.

Interview prompts

Companion reads: System ML · 23 Profiling, MFU and finding the bottleneck (the share of peak a run reaches), System Design · 02 Latency, throughput and queueing (Little's law) and Lesson 10 · What cameras cannot see (a task decided below the sensor's resolution).