all_lessons/Synthetic Vision Data/09 · Replay, checks and leakslesson 9 / 12

Replay, checks and leaks

Lesson 8 ended with a renderer that obeys the contract, a worker who wrote a quarter of a batch upside down while all twelve tests passed, and a factory that will run unattended for a million frames while nobody looks at a picture. This lesson gives the factory its own evidence: a manifest that makes every frame a replayable function of code, configuration and seeds; invariants that turn the contract into checks on every frame and batch; and splits by scene lineage that keep the exam from scoring data the model has already seen. The checks make the factory's number describe the program. They cannot say whether the program describes the street.

The thesis, here
An unattended factory has to carry its own evidence. Looking is a debugger and never a detector, so each failure needs a check that is code: drift is caught by replay, silent corruption by invariants derived from the contract, self-scoring by splitting on lineage. Each has a measurable chance of catching a given fault, and none can compare the program with the street.
Linear position
Forced by: A production renderer obeys the contract once each convention is set and checked: 9 of the 12 conventions we tested needed a setting other than the default, and each one failed silently before it was set. The tests check the renderer and nothing after it: a worker that wrote a quarter of a batch upside down passed all twelve and moved the exam by 1.7 points, hardly more than a new seed does. Once the factory runs unattended for a million frames nobody will look at the pictures. How do we know, automatically and every time it runs, that the data is the data we meant, and that the exam is not scoring the data against itself?
New idea: an unattended factory must carry its own evidence: a manifest that makes every frame a replayable function of code, configuration and seeds, invariants that every frame and batch must satisfy, and splits made by scene lineage. Each answers a different failure, drift, corruption and self-scoring, and each has a measurable chance of catching a given fault: of seventeen injected faults the invariants flag 14, the audit of the manifest 2 more, and one escapes everything inside the factory.
Forces next: A factory with a manifest that replays any frame, invariants that catch 14 of 17 injected bugs and an audit of the manifest two more, and splits by scene lineage that stop a memorizing model from scoring its own training data gives a number we can trust to describe the program: the naive program passes every check and its detector still misses 80% of the street's pedestrians. The program is still a model of the street, wrong in ways no check inside it can see, and a few real labelled frames can say how wrong. How should M real frames be spent, to grade, to calibrate the program or to train the detector, and what is a real frame worth in synthetic ones?
The plan
Six moves. (1) Show that looking cannot be the check. (2) Make a frame a row of a manifest and audit it twice. (3) Turn the contract into invariants with tolerances from the lab's numerics. (4) Inject seventeen faults, count what each check catches, add a canary and the exam. (5) Split by lineage and watch a memoriser read its training set. (6) Pass the naive program through the whole gate.

1 · Looking is a debugger, not a detector

Lesson 8's worker wrote a quarter of a batch upside down, image, mask and labels together. On the Street's own frames, over three seeds, that moves the exam by 0.5 points against a seed spread of 0.6 (Lesson 8's single seed on Blender's frames read 1.7): about a seed's worth either way, so the exam is not the check. Would looking be? Suppose a fault touches a fraction f of the frames a factory writes, and an inspector opens k of them at random, always recognising a corrupted frame when one is opened, which flatters the inspector. The inspector finds the fault with probability 1 − (1 − f)k, the chance that not every frame opened was clean. For the worker f = 25%, and 20 frames find it with probability 0.997; a smaller fault is another matter.

Touched fraction fk = 20k = 100k = 1,000
10%0.8781.0001.000
1%0.1820.6341.000
0.1%0.0200.0950.632

At a million frames a corruption of 0.1% is 1,000 bad frames, and a spot check of 100 finds it 9.5% of the time; 95% takes 2,995 frames, for one fault in one run. Looking is the right tool for a debugger: once a fault is suspected, one frame of the right kind shows it. It is the wrong tool for a detector, whose job is to find faults nobody suspects, in every run, at no cost per run. The check must be code on every frame, and each such check has a per-frame probability c of flagging a corrupted frame: c = 1 for a law the data must satisfy exactly, less when one frame holds thin evidence. Three failures need one: drift (the factory no longer makes what it made), corruption (a frame is not what the contract says) and self-scoring (the exam's frames are not independent of the training frames). The options, and what is wrong with each, order the lesson:

CheckCost per runIt can seeWhat goes wrong
look at a sampleattention, for 100 frameswhat a person notices in a picturethe arithmetic above
compare with a stored copyevery frame twicechangethe copy is as wrong as the original
manifest and audits (§2)29 bytes a frame, a render per audited rowdrift, unseeded noise, edited or reordered filesa bug that was there at the first run
invariants on every frame (§3)a pass over each framewhatever the contract stateswhat it does not state; thin evidence in one frame
a canary and the exam (§4)a training run, a labelled street setwhat training and the street care aboutcost and resolution: not for every batch
a split by lineage (§5)one join on the manifestan exam scoring data it has seennothing, if the lineage key is true

2 · A frame is a row: the manifest and two audits

In this lab a frame is a pure function of three things: the code, the configuration of the four stages, and the seeds (SV.sample draws the scene, the look and the sensor noise from three streams derived from one seed; §5 gives siblings different look and sensor seeds). A manifest row has to hold what is needed to re-make a frame and to know whether it was re-made.

Field of a rowBytesWhy
scene, look and sensor seeds12the whole randomness of the frame
split, and one bit: is a pedestrian in the scene1which side of the exam the frame feeds; the batch checks of §3
checksum of the pixels8FNV-1a, 64 bits, over the 8-bit values
checksum of the labels8the same over label, areas, coverage mask, instance ids, depth and range (to a millimetre)

Once per run the header holds the code version, a hash of the configuration and a golden digest, the digest of one fixed frame, which changes when anything in the generator does. A row is 29 bytes; the pixels of a frame alone are 6,912, so the manifest costs 0.42% of the data it describes and a million-frame run 29 MB. The checksums are 64 bits because a million frames make about 5·1011 pairs: a 32-bit digest would collide about 116 times among them, a 64-bit one once in 37 million runs. They are taken over quantised integers (depth and range to a millimetre, coverage to 1/255), so that a platform whose exp differs in the last bit cannot change them. Header and rows are the composition and collection-process part of the datasheet that Gebru et al. (2021) ask every dataset to ship, in a form a machine can check; they say how the data was made, not whether it was made right.

AuditWhat it doesIt seesIt cannot see
integrityre-hash every stored frame and compare with its row; no renderan edit, a truncation or label files reordered after the row was writtenanything the generator wrote wrongly
replayre-render k random rows from (configuration, seeds) with the code as it is now and comparechanged code or dependencies; a stored frame that is not a function of its seeds (unseeded noise, a stale buffer)a bug the generator already had when the row was written

The blind spot they share is the large one: a deterministic bug that was in the code at the first run. The row hashes what the generator wrote, and replay runs the same generator. §4 injects seventeen faults, and the audits sort them cleanly. Replaying every row finds a dependency drift (the blur of the sensor moved from 0.70 to 0.72 px, a change below anything an invariant can see) in 1,000 of 1,000 rows, a noise stream left unseeded for one frame in ten in 9.0% of rows, a stale buffer in 5.2%. Integrity finds label files listed in reverse order in 98 of the 100 frames of the batch (the other 2 are empty scenes whose labels are the same either way). Between them they find 0 of the 12 faults of the deterministic kind. A replay of k rows finds a fault that changes a fraction f of the rows with probability 1 − (1 − f)k, the arithmetic of §1: 29 rows for a 95% chance at f = 10%, 59 at 5%.

3 · The contract as code: invariants

An invariant is a pure function of one frame and of the program the factory declares it runs; it returns pass or fail. The contract of Lessons 7 and 8 already says what must hold, so the clauses are not invented, and each carries a tolerance under one rule: a tolerance comes from the lab's numerics and from clean frames, never from a bug it is meant to catch. An exact law is the strongest check there is, because its tolerance is the arithmetic precision of the data.

One clause needs a derivation, the camera's. In the sensor of Lesson 3 a pixel records electrons e′ = e + √e·n1 + r·n2 (n1, n2 unit normal draws, r the read noise), multiplies by the exposure gain g, clips at the full well fw and applies the gamma. The linear value q = g·e′/fw has variance (g/fw)·q̄ + (g·r/fw)² around its mean q̄ = g·e/fw. The road has one albedo and one shading, so two adjacent road pixels differ by noise alone, and half the variance of their difference estimates the pixel variance (in 8-bit values DN = q1/γ through the factor DN/(γq), plus the 1/12 of a grey level that rounding adds). Over n independent differences the ratio of observed to predicted variance is 1 ± √(2/n); the tolerance is 8 of those, and the check is skipped where the road records fewer than 20 electrons, where clipping at zero biases the law.

CheckThe clauseTolerance, and where it comes from
pixelsevery value is finite, in [0, 1] and a multiple of 1/255; the frame is not flatthe 8-bit grid; flat means a spread under one grey level
groundthe horizon is on row V0 in every frame, so no sky at or below it; depth at every ground pixel is fpx·hc / (v + ½ − V0) (focal length in pixels, camera height, row)10−6 relative: float32 spacing is 1.19×10−7; the worst clean frame is 5.0×10−8
rangerange = depth·√(1 + tan²φ + ρ²) for the pixel's ray (tan φ, ρ, 1), wherever a ray hit something10−6; the worst clean frame is 1.2×10−7
labellabel = [visible area ≥ max(kvis, rel·silhouette)] for the declared kvis and rel; area = sum of coverage ≤ silhouettenone: coverage is a multiple of 1/4
maskeach pixel with coverage lies within a pixel of an instance-map pixel and each instance-map pixel has coverage; rows where the silhouette is at least two map pixels widethe mask is sampled by four rays per pixel and the map by one; narrower rows can differ by two pixels, so they are skipped
gainthe logged gain is in [1, 8]; the mean linear level equals the exposure target while the gain is free2%; the worst clean frame deviates by 0.5%
skythe top row's chromaticity is a declared sky palette at that elevation0.06; the worst clean frame is 0.027
noisevariance of adjacent road pixels is the law above8σ of the ratio; the worst clean frame is 5.4σ
dup, lineageno two rows share both checksums, the same record twice; no scene seed appears in two splits (two hash joins)none: exact. The pixel checksum alone is not enough: the noise-free aaaa draws two scenes that differ by centimetres of depth identically, 1 pair in 1,000 frames
ratepedestrian scenes in a batch of 1,000 lie in the 5σ binomial interval of the declared 0.5, [421, 579]a clean batch leaves it with probability 4.7×10−7 (the exact tail)
poolthe noise ratio pooled over a batch of 100 frames3% (the law is good to about 2%) plus 6σ of the pooled sample

False alarms were counted on frames the tolerances were never set on. 0 of 5,000 fresh frames of the exact-stage program bbba and 0 of 2,000 of the naive program aaaa trip any per-frame check (so do the 6,000 and 3,000 frames the tolerances were examined on), 5 clean batches of 1,000 pass dup, lineage and rate, and 40 clean batches of 100 pass pool, whose ratio is 1.021 ± 0.005.

4 · The silent-bug bench

To count what the checks catch the lab needs faults. The bench has seventeen: faults of convention and geometry (a mirrored image, swapped channels, depth that holds range, a wrong focal length, a mask one pixel off, a threshold off by a factor, pixel centres at the corner) and of the data path and pipeline (gamma applied twice, noise missing in one batch, a mistyped probability, a hard-coded palette, a drifting dependency, an unseeded noise stream, a stale buffer, test frames that reuse training seeds, label files in the wrong order). Each is injected into bbba and touches a fraction f of a run of 1,000 frames; the lab knows which, a project learns only from a flag. The bench is not a map of all faults. Injecting faults to see which checks notice them is mutation testing (DeMillo, Lipton and Sayward, 1978): it grades the checks, not the program.

A per-frame check that flags a touched frame with probability c flags a random frame with probability f·c, so k frames contain a flag with probability 1 − (1 − f·c)k, and 95% takes k95 = ln 0.05 / ln(1 − f·c) ≈ 3/(f·c) frames. c is measured on 1,000 frames the fault touches; the widget below simulates the same thing from the run.

FaultTouchesPer-frame c (check)Frames for 95%, per-frameBatch checks and audit
frame written upside down (Lesson 8's worker)25%1.00 (ground, range)11none needed
depth pass holds range100%1.00 (ground, range)1none needed
focal length 8% off in the renderer only100%1.00 (ground, range)1none needed
pixel centres at +0, not +½100%1.00 (ground, range)1none needed
mask shifted one pixel100%0.39 (mask)7none needed
image mirrored, mask not50%0.12 (noise)49pool flags every batch
label files of one batch in reverse order10%0.04 (noise)748pool flags the batch; integrity 98 of 100
red and blue swapped100%0.31 (sky)9none needed
gamma applied twice10%1.00 (gain)29none needed
read noise missing in one batch10%0.00neverpool: ratio 0.897 in that batch
visibility threshold 4 px² instead of 1100%0.04 (label)76none needed
pedestrian probability 0.35, not 0.5100%0.00neverrate: 344 in [421, 579]
clothes palette hard-coded to the bright one100%0.00nevernone
dependency drift after the manifest100%0.00neverreplay: 1,000 of 1,000 rows
noise stream unseeded in 10% of frames10%0.00neverreplay: the touched rows
stale cache repeats the previous frame5%0.00neverdup: 52; replay: the touched rows
test frames reuse training seeds30%0.00neverdup: 64; lineage: 64

Read the table as classes. A fault of geometry or convention breaks an exact law on every frame, so one frame is enough (in every one of Lesson 8's upside-down frames the horizon is on the wrong row). A fault of registration, where pixels and labels no longer line up, is seen weakly because each frame is internally plausible: the mask clause catches a one-pixel shift in 0.39 of the frames, a mirrored image only through the road patch the maps call flat (0.12). Swapped channels are caught in 0.31 of the frames, because a blue sky with red and blue exchanged almost lands on the warm palette the look stage may also draw. Missing read noise is invisible in one frame and obvious in a hundred: the pooled ratio of the batch is 0.897 against 1.021 ± 0.005. The faults of reproducibility and of the splits produce valid frames; the joins and the audits see them, exactly.

Counting. In a run of 1,000 frames the per-frame and batch invariants flag 14 of the 17 faults with probability at least 0.95. The manifest audit flags 4, of which 2, the drift and the unseeded stream, no invariant sees; together 16. One fault, the hard-coded palette, produces frames that satisfy every clause, replay to the same digests and split cleanly: nothing inside the factory flags it.

The canary. For a fault that changes what the program draws without breaking a clause, use the data as it will be used: train the fixed detector on 160 frames of the run and grade it by AUC (the chance that a street frame with a pedestrian outscores one without; 0.5 is a coin) on 240 street frames that were labelled once, are never trained on and are reused by every run; 240 labelled frames, spent once. Thirty clean runs give 0.810 ± 0.015; a run below the mean minus three standard deviations, 0.766, is flagged, and 1 of 30 fresh clean runs was. Over three runs per fault, two or more below the threshold, the canary flags 1 of the seventeen, the swapped channels that the sky check already sees, and none of the three that no invariant flags (drift and the unseeded stream give valid frames; the palette lowers the AUC only to 0.782, 1.9 standard deviations). Its resolution, 0.045 of AUC, is set by the size of the run and not by the fault: it is a regression test of what training sees.

The exam. The exam is the program's grade: the miss rate over the street's pedestrians, at 10% false alarms, of the detector trained on the whole run of 1,600 frames. It costs three trainings and a labelled street set, so it belongs to a release and not to a batch. The clean run misses 46.8% ± 0.6 over three seeds (the exact-stage cell of Lesson 2's swap table: the same frames, the same procedure). The hard-coded palette adds 9.8 points, the swapped channels 19.4 and the one-pixel mask 2.4; the half-mirrored frames add 2.9 and the upside-down quarter 0.5, with spreads of 2.6 and 0.5 over seeds, which three seeds cannot tell from zero. What a check flags and what the exam punishes are different lists: the mask is flagged in 0.39 of its frames for 2.4 points, the upside-down frames in all of them for 0.5, and the palette, flagged by nothing, costs 9.8.

Road not taken · tighten the tolerances until the bench is caught
Mirrored images and swapped channels are caught in only 0.12 and 0.31 of the frames, and a 3σ tolerance on the noise ratio would catch more. What breaks is the one guarantee a tolerance has, the clean one: at 3σ the noise check flags 1.1% of clean frames, 10,900 alarms in a million, and at 4σ 0.2%. The road returns as a rule: keep the tolerance on the clean side (8σ here; the worst clean frame is 5.4σ) and raise the power by pooling frames, which is what the batch checks do.

5 · Leaks: the exam scoring the data against itself

The last failure is about the exam. A factory renders each scene under several lights and noise draws, siblings that share the scene stream and differ in the look and sensor streams, and splits frames at random: siblings land on both sides. With 4 siblings and an 80/20 split a test frame has on average 2.4 siblings in the training set, and only 0.8% have none; a split by lineage, all siblings of a scene on one side, leaves 0. Lineage is the manifest's scene seed, so the split costs one join; it is what scikit-learn's GroupKFold does with the scene seed as the group.

Three learners, 600 scenes in 2,400 frames, three seeds, held-out AUC for "a pedestrian is visible". The two memorisers see a frame as a unit vector, the column profile of its horizontal edges (for each of the 96 columns the sum over rows of |I(u+1) − I(u−1)|, square-rooted), which keeps the scene and drops the colours, and answer with the share of positives among the k nearest training frames. The detector is the series' fixed one, 111 weights.

LearnerSplit by frameSplit by lineageLeak
1-nearest neighbour0.7350.5630.173
10-nearest neighbours0.7420.5940.147
the fixed detector0.8050.7840.021

The leak is a property of capacity. For k = 1, 3, 10, 30 and 100 it is 0.173, 0.169, 0.147, 0.091 and 0.040: the fewer neighbours, the more the answer is a recollection. The detector cannot store 480 scenes in 111 weights: its gap is 0.021 on average and 0.068, −0.004 and 0.000 in the three seeds, within the noise of a 480-frame test set. That is the trap: graded by a low-capacity learner alone, a leaking split looks clean. On the random split the 1-nearest-neighbour learner even looks like a modest detector (0.735 against 0.805); only the lineage split shows that it learned the scenes and not the street.

Seed reuse is the same leak with identical siblings. When 30% of the test frames are copies of training frames, the 1-nearest-neighbour AUC rises from 0.563 to 0.701 and it recalls 100% of the duplicated positives; the detector gains 0.024. The check is not a model: the hash join of §3 finds every duplicate (64 of 64 in the bench), and the lineage join finds what a hash join cannot, siblings with different pixels.

Road not taken · split at random, then remove duplicates by hash
It is the standard cleaning step and exact for what it removes: copies. It removes none of the siblings (0 found), so the leak stays 0.173. The same failure is documented for public benchmarks: about 3% of the CIFAR-10 test images and about 9–10% of the CIFAR-100 test images have a near-duplicate in the training set (Barz and Denzler, 2020), and replacing them lowers the accuracy of popular networks by 9–14% relative; a survey found leakage in 294 papers in 17 fields of machine-learning-based science (Kapoor and Narayanan, 2023). A near-duplicate is a relation in the generator, not in the pixels, and a factory has the generator: key the split on its lineage.
The factory's console
Pick a fault and k, the number of frames checked. Top: the chance that k frames of a 1,000-frame run contain a flag, in closed form (lines) and simulated from the run (dots) for a looker who always recognises a bad frame once it is opened, the per-frame checks and a replay of k rows. Middle: the frame as drawn and as the faulty factory wrote it, rendered live (for faults of the data path or the manifest the frames are valid), every per-frame check run live on it, and the batch checks and audit of the run. Bottom: the leak, from the builder's table.
touched fraction f
—
per-frame c
—
P(found in k), closed form
—
P(found in k), simulated
—
frames for 95%
—
checks that flag the run
—
held-out AUC
—
leak (frame minus lineage)
—
Show the core JS
Hash.prototype.byte = function (b) {
  var lo = (this.lo ^ (b & 255)) >>> 0, t = lo * 0x1b3, hi = this.hi;
  this.lo = t % 4294967296;
  this.hi = (Math.imul(hi, 0x1b3) + Math.floor(t / 4294967296) + (lo << 8)) >>> 0;
L.row = function (rec, split) { var d = L.digest(rec); return { s: [rec.seeds.scene, rec.seeds.look, rec.seeds.sensor], split: split, hx: d.hx, hy: d.hy, ped: rec.full > 0 ? 1 : 0 }; };
    if (dp[q] > 0) { pred = D.f * D.hc / (v + 0.5 - D.V0); e = pred > 0 ? Math.abs(dp[q] - pred) / pred : 1; if (e > worst) worst = e; if (!(e <= L.TOL.f32)) bad++; }
  var lo = n * p - L.RATE_Z * sd, hi = n * p + L.RATE_Z * sd;
  function found(f, c, k) { return 1 - Math.pow(1 - f * c, k); }
  function k95(f, c) { var q = f * c; return q > 0 ? Math.max(1, Math.ceil(Math.log(0.05) / Math.log(1 - q))) : Infinity; }
    var sims = train.map(function (r, i) { var d = 0, q; for (q = 0; q < W; q++) d += t.f[q] * r.f[q]; return [d, i]; });

What to try. With no fault every cell is green. Pick frame written upside down, Lesson 8's worker: the second row of frames has the sky at the bottom, the ground and range cells turn red, c is 1.000, k = 10 gives 0.944 (0.94 simulated from the run) and 95% takes 11 frames. Pick mask shifted one pixel: c is 0.385, so k = 1 gives 0.385 and k = 10 gives 0.992 (0.99 simulated from the run); 95% takes 7 frames. Pick image mirrored, mask not: f is 50%, c is 0.119, k = 50 gives 0.953 by the formula and 0.900 in this run, which holds 49 flagged frames where about 60 are expected; the noise and pool checks flag the run. Pick stale cache: the check curve stays at zero, the replay curve rises, and dup is red (52 repeated frames): no frame is wrong, the rows are. Pick clothes palette hard-coded: no cell is red and nothing inside the factory fires. In the lower panel choose 1-nearest neighbour and tick the lineage box: the AUC goes from 0.735 to 0.563, a leak of 0.173; with the fixed detector it goes from 0.805 to 0.784, a gap within the seed noise.

6 · What no check inside the factory can see

Run the naive program through the whole gate. Its aaaa frames carry no noise, one light and a white van, so its declared contract is simpler, and every clause holds: 0 of 1,000 frames trip a per-frame check, the batches pass dup, lineage and rate (487 pedestrian scenes in [421, 579]), the integrity audit and the replay of every row find 0 mismatches, and its canary is steady, AUC 0.651 ± 0.013 over eight runs, because a regression test compares a run with earlier runs of the same program. The frames are exactly what the pipeline says they are. And the detector trained on them misses 80.3% of the street's pedestrians, against 46.8% for the exact-stage program, while missing 0.0% on its own frames and alarming on 100% of unlabelled street frames. The checks establish that the number the factory reports describes the program. They cannot establish that the program describes the street, because every clause is a statement about the program's own declaration. The canary's level, 0.651 against 0.810 for the exact-stage program, and the exam do reflect the street, because labelled street frames enter there: they grade the program, which no gate on a run can. Only frames from the street can say how wrong the program is.

What this lesson did not do
The bench is seventeen faults in one renderer and one program, not a map of the faults a real factory has. Blender frames would carry the same manifest (version, settings, seed, scene hash), but replay is bit-exact only under the conditions Lesson 8 found. A real fleet has no generator, so lineage there is inferred from drives, places and times, which this lesson does not do. How to spend labelled street frames is Lesson 10; what a changed braking policy does to the data is Lesson 11.

Common mistakes / failure modes

"We review a sample of frames every release"
A 100-frame look finds a fault in 0.1% of a million frames 9.5% of the time: the check must be code on every frame (§1).
"A checksum proves the data is right"
It proves the data is what was written: both audits find 0 of the 12 deterministic faults of the bench (§2).
"The detector shows no leak, so the split is clean"
The 111-weight detector's gap, 0.021, is within seed noise, while a 1-nearest-neighbour learner leaks 0.173 on the same split (§5).
"All checks green means a good simulator"
The naive program passes every check and misses 80.3% of the street's pedestrians (§6).

Checkpoint exercise

Try it
A factory writes 20,000 frames a night. A fault touches 0.5% of them, 100 frames. A per-frame invariant flags a touched frame with probability 0.2. How many frames must it see for a 95% chance, how many rows must a replay re-render if every touched row mismatches, and what does an inspector who opens 100 frames and always recognises a bad one achieve? What does the invariant achieve if it runs on all 20,000? Answer: the invariant flags a random frame with probability f·c = 0.005 · 0.2 = 0.0010, so k = ln 0.05 / ln(1 − 0.001) = 2,995 frames; the replay needs ln 0.05 / ln(1 − 0.005) = 598 rows; the inspector finds it with probability 1 − 0.995100 = 39%. On all 20,000 frames the invariant flags it with probability 1 − 0.99920000 = 100.0%.

Where this points next

The factory now carries its evidence: every frame is a row that can be re-made, the invariants flag 14 of seventeen injected faults and the audit of the manifest 2 more, and a split by lineage removes a leak of 0.173 AUC that the detector shows only as 0.021. All of it makes the number the factory reports describe the program, and the naive program passes every check while its detector misses 80.3% of the street's pedestrians. The program is a model of the street, wrong in ways nothing inside it can see, and only frames of the street can say how wrong. How should M real frames be spent, to grade, to calibrate the program or to train the detector, and what is a real frame worth in synthetic ones?

Takeaway
A factory nobody watches must carry its own evidence, because looking is a debugger: a 100-frame check finds a fault in 0.1% of a million frames 9.5% of the time. A manifest row of 29 bytes makes a frame replayable and auditable and finds drift, unseeded noise, stale buffers and reordered files, but not a deterministic bug present at the first run. Invariants turn the contract into code with tolerances taken from the numerics and from clean frames: of seventeen injected faults they flag 14 and the manifest 2 more, each with the chance 1 − (1 − f·c)k in k frames, and pooling frames raises a check's power without raising its false alarms. One fault, the hard-coded palette, is seen by nothing inside the factory and only by the exam. Splitting by scene lineage removes a leak of 0.173 AUC from a memorising learner; a low-capacity detector shows 0.021, within noise. Passing every check makes the number describe the program; the naive program passes them all while missing 80.3% of the street.

Interview prompts

Companion reads: Data Engineering · 06 Dedup and decontamination, Data Engineering · 08 Quality validation and Computer Vision · 19 Evaluation and deployment.