Replay, checks and leaks
Lesson 8 ended with a renderer that obeys the contract, a worker who wrote a quarter of a batch upside down while all twelve tests passed, and a factory that will run unattended for a million frames while nobody looks at a picture. This lesson gives the factory its own evidence: a manifest that makes every frame a replayable function of code, configuration and seeds; invariants that turn the contract into checks on every frame and batch; and splits by scene lineage that keep the exam from scoring data the model has already seen. The checks make the factory's number describe the program. They cannot say whether the program describes the street.
New idea: an unattended factory must carry its own evidence: a manifest that makes every frame a replayable function of code, configuration and seeds, invariants that every frame and batch must satisfy, and splits made by scene lineage. Each answers a different failure, drift, corruption and self-scoring, and each has a measurable chance of catching a given fault: of seventeen injected faults the invariants flag 14, the audit of the manifest 2 more, and one escapes everything inside the factory.
Forces next: A factory with a manifest that replays any frame, invariants that catch 14 of 17 injected bugs and an audit of the manifest two more, and splits by scene lineage that stop a memorizing model from scoring its own training data gives a number we can trust to describe the program: the naive program passes every check and its detector still misses 80% of the street's pedestrians. The program is still a model of the street, wrong in ways no check inside it can see, and a few real labelled frames can say how wrong. How should M real frames be spent, to grade, to calibrate the program or to train the detector, and what is a real frame worth in synthetic ones?
1 · Looking is a debugger, not a detector
Lesson 8's worker wrote a quarter of a batch upside down, image, mask and labels together. On the Street's own frames, over three seeds, that moves the exam by 0.5 points against a seed spread of 0.6 (Lesson 8's single seed on Blender's frames read 1.7): about a seed's worth either way, so the exam is not the check. Would looking be? Suppose a fault touches a fraction f of the frames a factory writes, and an inspector opens k of them at random, always recognising a corrupted frame when one is opened, which flatters the inspector. The inspector finds the fault with probability 1 − (1 − f)k, the chance that not every frame opened was clean. For the worker f = 25%, and 20 frames find it with probability 0.997; a smaller fault is another matter.
| Touched fraction f | k = 20 | k = 100 | k = 1,000 |
|---|---|---|---|
| 10% | 0.878 | 1.000 | 1.000 |
| 1% | 0.182 | 0.634 | 1.000 |
| 0.1% | 0.020 | 0.095 | 0.632 |
At a million frames a corruption of 0.1% is 1,000 bad frames, and a spot check of 100 finds it 9.5% of the time; 95% takes 2,995 frames, for one fault in one run. Looking is the right tool for a debugger: once a fault is suspected, one frame of the right kind shows it. It is the wrong tool for a detector, whose job is to find faults nobody suspects, in every run, at no cost per run. The check must be code on every frame, and each such check has a per-frame probability c of flagging a corrupted frame: c = 1 for a law the data must satisfy exactly, less when one frame holds thin evidence. Three failures need one: drift (the factory no longer makes what it made), corruption (a frame is not what the contract says) and self-scoring (the exam's frames are not independent of the training frames). The options, and what is wrong with each, order the lesson:
| Check | Cost per run | It can see | What goes wrong |
|---|---|---|---|
| look at a sample | attention, for 100 frames | what a person notices in a picture | the arithmetic above |
| compare with a stored copy | every frame twice | change | the copy is as wrong as the original |
| manifest and audits (§2) | 29 bytes a frame, a render per audited row | drift, unseeded noise, edited or reordered files | a bug that was there at the first run |
| invariants on every frame (§3) | a pass over each frame | whatever the contract states | what it does not state; thin evidence in one frame |
| a canary and the exam (§4) | a training run, a labelled street set | what training and the street care about | cost and resolution: not for every batch |
| a split by lineage (§5) | one join on the manifest | an exam scoring data it has seen | nothing, if the lineage key is true |
2 · A frame is a row: the manifest and two audits
In this lab a frame is a pure function of three things: the code, the configuration of the four stages, and the seeds (SV.sample draws the scene, the look and the sensor noise from three streams derived from one seed; §5 gives siblings different look and sensor seeds). A manifest row has to hold what is needed to re-make a frame and to know whether it was re-made.
| Field of a row | Bytes | Why |
|---|---|---|
| scene, look and sensor seeds | 12 | the whole randomness of the frame |
| split, and one bit: is a pedestrian in the scene | 1 | which side of the exam the frame feeds; the batch checks of §3 |
| checksum of the pixels | 8 | FNV-1a, 64 bits, over the 8-bit values |
| checksum of the labels | 8 | the same over label, areas, coverage mask, instance ids, depth and range (to a millimetre) |
Once per run the header holds the code version, a hash of the configuration and a golden digest, the digest of one fixed frame, which changes when anything in the generator does. A row is 29 bytes; the pixels of a frame alone are 6,912, so the manifest costs 0.42% of the data it describes and a million-frame run 29 MB. The checksums are 64 bits because a million frames make about 5·1011 pairs: a 32-bit digest would collide about 116 times among them, a 64-bit one once in 37 million runs. They are taken over quantised integers (depth and range to a millimetre, coverage to 1/255), so that a platform whose exp differs in the last bit cannot change them. Header and rows are the composition and collection-process part of the datasheet that Gebru et al. (2021) ask every dataset to ship, in a form a machine can check; they say how the data was made, not whether it was made right.
| Audit | What it does | It sees | It cannot see |
|---|---|---|---|
| integrity | re-hash every stored frame and compare with its row; no render | an edit, a truncation or label files reordered after the row was written | anything the generator wrote wrongly |
| replay | re-render k random rows from (configuration, seeds) with the code as it is now and compare | changed code or dependencies; a stored frame that is not a function of its seeds (unseeded noise, a stale buffer) | a bug the generator already had when the row was written |
The blind spot they share is the large one: a deterministic bug that was in the code at the first run. The row hashes what the generator wrote, and replay runs the same generator. §4 injects seventeen faults, and the audits sort them cleanly. Replaying every row finds a dependency drift (the blur of the sensor moved from 0.70 to 0.72 px, a change below anything an invariant can see) in 1,000 of 1,000 rows, a noise stream left unseeded for one frame in ten in 9.0% of rows, a stale buffer in 5.2%. Integrity finds label files listed in reverse order in 98 of the 100 frames of the batch (the other 2 are empty scenes whose labels are the same either way). Between them they find 0 of the 12 faults of the deterministic kind. A replay of k rows finds a fault that changes a fraction f of the rows with probability 1 − (1 − f)k, the arithmetic of §1: 29 rows for a 95% chance at f = 10%, 59 at 5%.
3 · The contract as code: invariants
An invariant is a pure function of one frame and of the program the factory declares it runs; it returns pass or fail. The contract of Lessons 7 and 8 already says what must hold, so the clauses are not invented, and each carries a tolerance under one rule: a tolerance comes from the lab's numerics and from clean frames, never from a bug it is meant to catch. An exact law is the strongest check there is, because its tolerance is the arithmetic precision of the data.
One clause needs a derivation, the camera's. In the sensor of Lesson 3 a pixel records electrons e′ = e + √e·n1 + r·n2 (n1, n2 unit normal draws, r the read noise), multiplies by the exposure gain g, clips at the full well fw and applies the gamma. The linear value q = g·e′/fw has variance (g/fw)·q̄ + (g·r/fw)² around its mean q̄ = g·e/fw. The road has one albedo and one shading, so two adjacent road pixels differ by noise alone, and half the variance of their difference estimates the pixel variance (in 8-bit values DN = q1/γ through the factor DN/(γq), plus the 1/12 of a grey level that rounding adds). Over n independent differences the ratio of observed to predicted variance is 1 ± √(2/n); the tolerance is 8 of those, and the check is skipped where the road records fewer than 20 electrons, where clipping at zero biases the law.
| Check | The clause | Tolerance, and where it comes from |
|---|---|---|
| pixels | every value is finite, in [0, 1] and a multiple of 1/255; the frame is not flat | the 8-bit grid; flat means a spread under one grey level |
| ground | the horizon is on row V0 in every frame, so no sky at or below it; depth at every ground pixel is fpx·hc / (v + ½ − V0) (focal length in pixels, camera height, row) | 10−6 relative: float32 spacing is 1.19×10−7; the worst clean frame is 5.0×10−8 |
| range | range = depth·√(1 + tan²φ + ρ²) for the pixel's ray (tan φ, ρ, 1), wherever a ray hit something | 10−6; the worst clean frame is 1.2×10−7 |
| label | label = [visible area ≥ max(kvis, rel·silhouette)] for the declared kvis and rel; area = sum of coverage ≤ silhouette | none: coverage is a multiple of 1/4 |
| mask | each pixel with coverage lies within a pixel of an instance-map pixel and each instance-map pixel has coverage; rows where the silhouette is at least two map pixels wide | the mask is sampled by four rays per pixel and the map by one; narrower rows can differ by two pixels, so they are skipped |
| gain | the logged gain is in [1, 8]; the mean linear level equals the exposure target while the gain is free | 2%; the worst clean frame deviates by 0.5% |
| sky | the top row's chromaticity is a declared sky palette at that elevation | 0.06; the worst clean frame is 0.027 |
| noise | variance of adjacent road pixels is the law above | 8σ of the ratio; the worst clean frame is 5.4σ |
| dup, lineage | no two rows share both checksums, the same record twice; no scene seed appears in two splits (two hash joins) | none: exact. The pixel checksum alone is not enough: the noise-free aaaa draws two scenes that differ by centimetres of depth identically, 1 pair in 1,000 frames |
| rate | pedestrian scenes in a batch of 1,000 lie in the 5σ binomial interval of the declared 0.5, [421, 579] | a clean batch leaves it with probability 4.7×10−7 (the exact tail) |
| pool | the noise ratio pooled over a batch of 100 frames | 3% (the law is good to about 2%) plus 6σ of the pooled sample |
False alarms were counted on frames the tolerances were never set on. 0 of 5,000 fresh frames of the exact-stage program bbba and 0 of 2,000 of the naive program aaaa trip any per-frame check (so do the 6,000 and 3,000 frames the tolerances were examined on), 5 clean batches of 1,000 pass dup, lineage and rate, and 40 clean batches of 100 pass pool, whose ratio is 1.021 ± 0.005.
4 · The silent-bug bench
To count what the checks catch the lab needs faults. The bench has seventeen: faults of convention and geometry (a mirrored image, swapped channels, depth that holds range, a wrong focal length, a mask one pixel off, a threshold off by a factor, pixel centres at the corner) and of the data path and pipeline (gamma applied twice, noise missing in one batch, a mistyped probability, a hard-coded palette, a drifting dependency, an unseeded noise stream, a stale buffer, test frames that reuse training seeds, label files in the wrong order). Each is injected into bbba and touches a fraction f of a run of 1,000 frames; the lab knows which, a project learns only from a flag. The bench is not a map of all faults. Injecting faults to see which checks notice them is mutation testing (DeMillo, Lipton and Sayward, 1978): it grades the checks, not the program.
A per-frame check that flags a touched frame with probability c flags a random frame with probability f·c, so k frames contain a flag with probability 1 − (1 − f·c)k, and 95% takes k95 = ln 0.05 / ln(1 − f·c) ≈ 3/(f·c) frames. c is measured on 1,000 frames the fault touches; the widget below simulates the same thing from the run.
| Fault | Touches | Per-frame c (check) | Frames for 95%, per-frame | Batch checks and audit |
|---|---|---|---|---|
| frame written upside down (Lesson 8's worker) | 25% | 1.00 (ground, range) | 11 | none needed |
| depth pass holds range | 100% | 1.00 (ground, range) | 1 | none needed |
| focal length 8% off in the renderer only | 100% | 1.00 (ground, range) | 1 | none needed |
| pixel centres at +0, not +½ | 100% | 1.00 (ground, range) | 1 | none needed |
| mask shifted one pixel | 100% | 0.39 (mask) | 7 | none needed |
| image mirrored, mask not | 50% | 0.12 (noise) | 49 | pool flags every batch |
| label files of one batch in reverse order | 10% | 0.04 (noise) | 748 | pool flags the batch; integrity 98 of 100 |
| red and blue swapped | 100% | 0.31 (sky) | 9 | none needed |
| gamma applied twice | 10% | 1.00 (gain) | 29 | none needed |
| read noise missing in one batch | 10% | 0.00 | never | pool: ratio 0.897 in that batch |
| visibility threshold 4 px² instead of 1 | 100% | 0.04 (label) | 76 | none needed |
| pedestrian probability 0.35, not 0.5 | 100% | 0.00 | never | rate: 344 in [421, 579] |
| clothes palette hard-coded to the bright one | 100% | 0.00 | never | none |
| dependency drift after the manifest | 100% | 0.00 | never | replay: 1,000 of 1,000 rows |
| noise stream unseeded in 10% of frames | 10% | 0.00 | never | replay: the touched rows |
| stale cache repeats the previous frame | 5% | 0.00 | never | dup: 52; replay: the touched rows |
| test frames reuse training seeds | 30% | 0.00 | never | dup: 64; lineage: 64 |
Read the table as classes. A fault of geometry or convention breaks an exact law on every frame, so one frame is enough (in every one of Lesson 8's upside-down frames the horizon is on the wrong row). A fault of registration, where pixels and labels no longer line up, is seen weakly because each frame is internally plausible: the mask clause catches a one-pixel shift in 0.39 of the frames, a mirrored image only through the road patch the maps call flat (0.12). Swapped channels are caught in 0.31 of the frames, because a blue sky with red and blue exchanged almost lands on the warm palette the look stage may also draw. Missing read noise is invisible in one frame and obvious in a hundred: the pooled ratio of the batch is 0.897 against 1.021 ± 0.005. The faults of reproducibility and of the splits produce valid frames; the joins and the audits see them, exactly.
Counting. In a run of 1,000 frames the per-frame and batch invariants flag 14 of the 17 faults with probability at least 0.95. The manifest audit flags 4, of which 2, the drift and the unseeded stream, no invariant sees; together 16. One fault, the hard-coded palette, produces frames that satisfy every clause, replay to the same digests and split cleanly: nothing inside the factory flags it.
The canary. For a fault that changes what the program draws without breaking a clause, use the data as it will be used: train the fixed detector on 160 frames of the run and grade it by AUC (the chance that a street frame with a pedestrian outscores one without; 0.5 is a coin) on 240 street frames that were labelled once, are never trained on and are reused by every run; 240 labelled frames, spent once. Thirty clean runs give 0.810 ± 0.015; a run below the mean minus three standard deviations, 0.766, is flagged, and 1 of 30 fresh clean runs was. Over three runs per fault, two or more below the threshold, the canary flags 1 of the seventeen, the swapped channels that the sky check already sees, and none of the three that no invariant flags (drift and the unseeded stream give valid frames; the palette lowers the AUC only to 0.782, 1.9 standard deviations). Its resolution, 0.045 of AUC, is set by the size of the run and not by the fault: it is a regression test of what training sees.
The exam. The exam is the program's grade: the miss rate over the street's pedestrians, at 10% false alarms, of the detector trained on the whole run of 1,600 frames. It costs three trainings and a labelled street set, so it belongs to a release and not to a batch. The clean run misses 46.8% ± 0.6 over three seeds (the exact-stage cell of Lesson 2's swap table: the same frames, the same procedure). The hard-coded palette adds 9.8 points, the swapped channels 19.4 and the one-pixel mask 2.4; the half-mirrored frames add 2.9 and the upside-down quarter 0.5, with spreads of 2.6 and 0.5 over seeds, which three seeds cannot tell from zero. What a check flags and what the exam punishes are different lists: the mask is flagged in 0.39 of its frames for 2.4 points, the upside-down frames in all of them for 0.5, and the palette, flagged by nothing, costs 9.8.
5 · Leaks: the exam scoring the data against itself
The last failure is about the exam. A factory renders each scene under several lights and noise draws, siblings that share the scene stream and differ in the look and sensor streams, and splits frames at random: siblings land on both sides. With 4 siblings and an 80/20 split a test frame has on average 2.4 siblings in the training set, and only 0.8% have none; a split by lineage, all siblings of a scene on one side, leaves 0. Lineage is the manifest's scene seed, so the split costs one join; it is what scikit-learn's GroupKFold does with the scene seed as the group.
Three learners, 600 scenes in 2,400 frames, three seeds, held-out AUC for "a pedestrian is visible". The two memorisers see a frame as a unit vector, the column profile of its horizontal edges (for each of the 96 columns the sum over rows of |I(u+1) − I(u−1)|, square-rooted), which keeps the scene and drops the colours, and answer with the share of positives among the k nearest training frames. The detector is the series' fixed one, 111 weights.
| Learner | Split by frame | Split by lineage | Leak |
|---|---|---|---|
| 1-nearest neighbour | 0.735 | 0.563 | 0.173 |
| 10-nearest neighbours | 0.742 | 0.594 | 0.147 |
| the fixed detector | 0.805 | 0.784 | 0.021 |
The leak is a property of capacity. For k = 1, 3, 10, 30 and 100 it is 0.173, 0.169, 0.147, 0.091 and 0.040: the fewer neighbours, the more the answer is a recollection. The detector cannot store 480 scenes in 111 weights: its gap is 0.021 on average and 0.068, −0.004 and 0.000 in the three seeds, within the noise of a 480-frame test set. That is the trap: graded by a low-capacity learner alone, a leaking split looks clean. On the random split the 1-nearest-neighbour learner even looks like a modest detector (0.735 against 0.805); only the lineage split shows that it learned the scenes and not the street.
Seed reuse is the same leak with identical siblings. When 30% of the test frames are copies of training frames, the 1-nearest-neighbour AUC rises from 0.563 to 0.701 and it recalls 100% of the duplicated positives; the detector gains 0.024. The check is not a model: the hash join of §3 finds every duplicate (64 of 64 in the bench), and the lineage join finds what a hash join cannot, siblings with different pixels.
What to try. With no fault every cell is green. Pick frame written upside down, Lesson 8's worker: the second row of frames has the sky at the bottom, the ground and range cells turn red, c is 1.000, k = 10 gives 0.944 (0.94 simulated from the run) and 95% takes 11 frames. Pick mask shifted one pixel: c is 0.385, so k = 1 gives 0.385 and k = 10 gives 0.992 (0.99 simulated from the run); 95% takes 7 frames. Pick image mirrored, mask not: f is 50%, c is 0.119, k = 50 gives 0.953 by the formula and 0.900 in this run, which holds 49 flagged frames where about 60 are expected; the noise and pool checks flag the run. Pick stale cache: the check curve stays at zero, the replay curve rises, and dup is red (52 repeated frames): no frame is wrong, the rows are. Pick clothes palette hard-coded: no cell is red and nothing inside the factory fires. In the lower panel choose 1-nearest neighbour and tick the lineage box: the AUC goes from 0.735 to 0.563, a leak of 0.173; with the fixed detector it goes from 0.805 to 0.784, a gap within the seed noise.
6 · What no check inside the factory can see
Run the naive program through the whole gate. Its aaaa frames carry no noise, one light and a white van, so its declared contract is simpler, and every clause holds: 0 of 1,000 frames trip a per-frame check, the batches pass dup, lineage and rate (487 pedestrian scenes in [421, 579]), the integrity audit and the replay of every row find 0 mismatches, and its canary is steady, AUC 0.651 ± 0.013 over eight runs, because a regression test compares a run with earlier runs of the same program. The frames are exactly what the pipeline says they are. And the detector trained on them misses 80.3% of the street's pedestrians, against 46.8% for the exact-stage program, while missing 0.0% on its own frames and alarming on 100% of unlabelled street frames. The checks establish that the number the factory reports describes the program. They cannot establish that the program describes the street, because every clause is a statement about the program's own declaration. The canary's level, 0.651 against 0.810 for the exact-stage program, and the exam do reflect the street, because labelled street frames enter there: they grade the program, which no gate on a run can. Only frames from the street can say how wrong the program is.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
The factory now carries its evidence: every frame is a row that can be re-made, the invariants flag 14 of seventeen injected faults and the audit of the manifest 2 more, and a split by lineage removes a leak of 0.173 AUC that the detector shows only as 0.021. All of it makes the number the factory reports describe the program, and the naive program passes every check while its detector misses 80.3% of the street's pedestrians. The program is a model of the street, wrong in ways nothing inside it can see, and only frames of the street can say how wrong. How should M real frames be spent, to grade, to calibrate the program or to train the detector, and what is a real frame worth in synthetic ones?
Interview prompts
- A worker writes a quarter of a batch upside down and all twelve renderer tests pass. Which checks fire, and after how many frames? (§3, §4 — the horizon and ground clauses on every touched frame: 11 frames for 95%; the exam moves a seed's worth.)
- What must a manifest row hold for a frame to be replayable, and what can replay not detect? (§2 — seeds, split, configuration hash, code version, two checksums; a deterministic bug present at the first run.)
- Where should an invariant's tolerance come from? (§3 — the numerics of the data and clean frames, never the bugs it should catch.)
- A check flags a touched frame with probability 0.12 and the fault touches half the frames: how many frames for 95%? (§4 — ln 0.05 / ln(1 − 0.06) ≈ 49.)
- What does a canary catch, what does it cost and what is its resolution? (§4 — a regression in what training sees; 240 labelled frames, once; about 0.045 of AUC.)
- Your detector scores the same on a random split and on a split by scene. Is the split clean? (§5 — not necessarily: a low-capacity learner cannot memorise, and a 1-nearest-neighbour learner leaks 0.173 on the same split.)
- Every check passes and the model fails on real data. Why? (§6 — every clause is about the program's own declaration; the naive program passes all of them and misses 80.3%.)
Companion reads: Data Engineering · 06 Dedup and decontamination, Data Engineering · 08 Quality validation and Computer Vision · 19 Evaluation and deployment.