all_lessons/Synthetic Vision Data/index12 lessons · ~5h read

Synthetic Vision Data, from first principles

A program that draws a picture can write the answer beside it, so a vision model can be given as many labelled pictures as anyone cares to render. Whether the model has learned the street or only the program is a separate question, and this track is its answer. It builds one small laboratory, a street with a parked van, a shuttle's camera and a pedestrian who sometimes steps out, and uses it to price every part of a data-generating program in points of a single exam. Twelve lessons follow the gap: from the wall that free labels do not remove, through the stages that make it up and the evidence a project can collect about each, to the production renderer that has to obey the same contract, the checks that keep an unattended factory honest, what a few real frames are worth, what braking changes, and what counts as proof.

The seed question
A million labelled frames of a street cost an afternoon of compute; a thousand real ones cost a month. A detector trained on the million is perfect on the million and misses most of the pedestrians on the street. What exactly did the million teach it, what would make them teach the right thing, and how would you know?
Who this is for
An engineer or student who has trained a vision model, has been offered data from a simulator, and wants to know what it is worth before betting a project on it. The track assumes basic probability, logistic regression and the pinhole camera of Computer Vision · 04; what a renderer does is explained where it is first used. By the end you can say what a program's labels are exact about and what they are not; split the gap between a program and the world into stages and price each in points of a real exam; say which evidence a project can collect about each stage and what it costs; calibrate a camera model from flat frames; choose the range over which to randomize a nuisance and measure what that costs; count the cases a scene program never draws and spend renders on the rare ones without distorting what the model learns; write down the definitions a label rests on; carry the contract into a production renderer and test that it holds; keep a factory replayable and free of leakage; spend a few real frames well; use a program to say what a decision changes; and say how much real driving a claim of safety needs.
How this track is built
An original first-principles track, not a survey of the usual list of simulators, domain-randomization tricks and GAN refiners. Five things set it apart. (1) A relay, not a list. Every lesson opens with the problem the previous one left standing and closes by stating the next; that hand-off is one sentence, the baton, which appears word for word at the end of one lesson, at the start of the next and in the table below, and a script diffs them. (2) One exam. Every repair enters because the one before it left a measurable part of the same gap standing, so "better" always means points on the same ruler. (3) The mechanism runs. A small laboratory, the Street, a renderer, a camera model and a fixed detector in a few thousand lines of JavaScript, lets every widget train, grade and swap parts of a program in front of you; the "real street" is a second, richer program that the lessons may only sample, which is what lets us run swaps no project can. (4) Every number is recomputed. Each quoted number is tagged and an independent script re-derives it from scratch; facts about real systems come only from a verified list and are cited by name and year. (5) The production contract is executed. Lesson 8 runs Blender headless and compares it with the laboratory, rather than describing what a renderer should do.

The exam

A detector scores every frame. Take the frames of the world being graded that contain no pedestrian, set the threshold so that one in ten raises an alarm, and count the pedestrians (with at least six pixels of area showing) whose frames score below it: that fraction is the miss rate at 10% false alarms. The detector never changes; only its training data does. A detector trained on the naive program misses 80% of the street's pedestrians, one trained on the street itself misses 47%, and the 33 points between them are the thing the track divides up. Lesson 1 sets the exam up and shows that free labels cannot close the gap; each later lesson repairs one part of it and says how many points it bought.

Part I · The factory and the wall (lessons 01–02)

What a program that draws the answer buys, where its usefulness stops, and what each stage of the program costs.

01
Free labels, and the wall
One render yields image, mask, depth, range and silhouette; the Street, the fixed detector and the exam; learning curves on two worlds; the gap split into variance, which data removes, and bias, which it does not.
02
Where the gap lives
Camera, light, scene and label as a pipeline; the sixteen-way swap table; Shapley prices; why repairs compound; the alarm rate on unlabelled frames as a label-free gap meter; the evidence each stage answers to.

Part II · Pixels first (lessons 03–04)

Make the program's pixels the world's pixels: the camera that records the scene, then the light that falls on it.

03
The camera is a measurement
Photons to electrons to numbers; shot and read noise; photon-transfer calibration; auto-exposure and saturation; blur; motion during the exposure and the f·v·t/Z law; the detectability floor.
04
Same scene, other pixels
Look parameters as nuisance variables; randomization range against real accuracy and against data efficiency; probe pixels (sky, road) and exposure-invariant ratios; quantile checks of synthetic against real.

Part III · Cases (lessons 05–06)

Then make sure the right cases appear, and the rare ones in useful numbers.

05
Cover the cases
The scene program and its support; a coverage ruler built from a few labelled real frames and the ground-plane geometry; the chance a case of probability p is absent from N frames; what each missing case costs.
06
The rare case
Oversampling the step-out, prior shift and the logit correction, importance weights and effective sample size, allocation of renders across cases.

Part IV · Labels and the contract (lessons 07–08)

What the label means, and the same contract in a renderer nobody here wrote.

07
What the label means
Modal and amodal masks; the visibility rule and one detector's score under three conventions; depth along the axis against range along the ray; half-pixel alignment; label time.
08
The same contract in a real renderer
Blender 5.2 executed headless; axes and units; intrinsics from lens and sensor width; pixel centres; the depth pass; object-index masks and filter width; colour management; seeds and the manifest.

Part V · Trust and reality (lessons 09–10)

The checks that keep an unattended factory honest, and what a few real frames can still say.

09
Replay, checks and leaks
Seeds and provenance; invariants as tests of the contract; a bench of injected bugs; near-duplicate leakage and how it grows with model capacity; lineage splits.
10
A little reality
Grading, calibrating and fine-tuning with M real frames; mixing ratios; the exchange rate between real and synthetic frames; label drift when a refiner changes the picture.

Part VI · Action and proof (lessons 11–12)

What braking changes, and what counts as proof on the real street.

11
What braking changes
Observation against intervention; why logs of drivers mislead; matched branches and common random numbers; stopping margins from the detector's psychometric curve; positivity.
12
Proof
The rule of three; open against closed loop; rank agreement between simulated and real exams across the sixteen programs; the final ledger of gaps.

The whole derivation on one page

Read the second column of any row and the fourth column of the row above it: they are the same sentence. That is the linear thinking made visible. Each lesson's problem it inherits is exactly the previous lesson's problem it creates, and the middle column is the single move that resolves the first and manufactures the second. The first row inherits its problem from computer graphics, which taught the renderer as a function; the last row hands its problem to the tracks on training a robot model and world models, where the budget is spent.

#The problem it inheritsThe one moveThe problem it creates
Part I · The factory and the wall
01Computer graphics taught the renderer as a function: a scene goes in, an image comes out, and the program that draws the image knows exactly what is in it. A vision model needs the opposite service, images that arrive with the answer, by the thousand, for nothing, and a renderer provides exactly that. What does a model learn from pictures a program drew, and how would we find out whether it learned the street or only the program?Free labels, and the wall: A program that draws a picture also knows its answer, so labels are free and exact; train on them and the error on the program's own pictures falls to nothing while the error on the real street stops at a floor that no amount of data lowers.Free labels are exact and unlimited, and a detector trained on them stops making mistakes on the program's own pictures after about forty frames; on pictures of the real street it still misses 80% of the pedestrians, whether it saw forty frames or twenty-five hundred. The gap is not noise that more data averages away: it is a difference between the program and the world, and the program has four places to be wrong: how the camera records the scene, how the scene is lit, what exists in it, and what the label means. Which of them costs the most, and how could we tell without a real label for every frame?
02Free labels are exact and unlimited, and a detector trained on them stops making mistakes on the program's own pictures after about forty frames; on pictures of the real street it still misses 80% of the pedestrians, whether it saw forty frames or twenty-five hundred. The gap is not noise that more data averages away: it is a difference between the program and the world, and the program has four places to be wrong: how the camera records the scene, how the scene is lit, what exists in it, and what the label means. Which of them costs the most, and how could we tell without a real label for every frame?Where the gap lives: Treat the program as four stages, swap them one at a time between the program and the world to price each, and read off which evidence a project can collect about each and what it costs.Swapping the stages one at a time prices them: of the 33 points of miss rate between a detector trained on the program (80%) and one trained on the street itself (47%), the camera accounts for 16, the light and the scene for about 8 each, and the label for nothing, and the repairs compound, so the camera is worth 14.5 points while the other stages are wrong and 21 once they are right. The camera is therefore the place to start, because it is both the most expensive gap and the cheapest to measure: a few frames of a grey card pin down its noise without a single label. How does a camera turn light into numbers, and how do we measure it well enough to simulate it?
Part II · Pixels first
03Swapping the stages one at a time prices them: of the 33 points of miss rate between a detector trained on the program (80%) and one trained on the street itself (47%), the camera accounts for 16, the light and the scene for about 8 each, and the label for nothing, and the repairs compound, so the camera is worth 14.5 points while the other stages are wrong and 21 once they are right. The camera is therefore the place to start, because it is both the most expensive gap and the cheapest to measure: a few frames of a grey card pin down its noise without a single label. How does a camera turn light into numbers, and how do we measure it well enough to simulate it?The camera is a measurement: A pixel value is a measurement with a noise law: calibrate the full well and the read noise from flat frames, the blur from an edge and the exposure control from logs, add them to the program, and see which pedestrians the noise erases.Flat frames of a grey card gave the camera's full well and read noise, an edge gave its blur, and logs gave its exposure control; added to the program, they take the miss rate from 80% to 66%, within a point of what the street's own camera numbers give, and the program's pictures now carry the grain and softness of the real camera. A calibrated camera still sees only the light the program gives it: one sun behind the camera, white vans, bright shirts, a blue sky. The program's frames leave its exposure controller at a gain near 2, the street's run from 1.9 to the cap of 8 with a third of them at the cap, and the controller that is worth 1 point in the program's light is worth 7 in the street's. The real street shows the same scenes at dusk, in glare and under cloud. What should be allowed to vary between two pictures of the same scene, and what must not?
04Flat frames of a grey card gave the camera's full well and read noise, an edge gave its blur, and logs gave its exposure control; added to the program, they take the miss rate from 80% to 66%, within a point of what the street's own camera numbers give, and the program's pictures now carry the grain and softness of the real camera. A calibrated camera still sees only the light the program gives it: one sun behind the camera, white vans, bright shirts, a blue sky. The program's frames leave its exposure controller at a gain near 2, the street's run from 1.9 to the cap of 8 with a third of them at the cap, and the controller that is worth 1 point in the program's light is worth 7 in the street's. The real street shows the same scenes at dusk, in glare and under cloud. What should be allowed to vary between two pictures of the same scene, and what must not?Same scene, other pixels: Separate what the answer depends on from what it must ignore: randomize the look over the range the real street occupies, estimate that range from probe pixels in unlabelled frames, and pay the price in data that invariance costs.Randomising the light over the range read from the sky of unlabelled frames takes the miss rate from 66% to 61%, within a point of the 62% the street's exact light gives, and randomising wider than that buys no accuracy while the program's own error climbs from 7% to 17%. With the pixels right, what the program draws starts to matter in a way it did not: a scene repair that was worth a little over 5 points while the pictures were wrong is worth 15 now. And the scene is where the program is silent about whole kinds of case: every pedestrian it drew stood in the open and within 16 m. Which real cases does the scene never draw, and how do we count them?
Part III · Cases
05Randomising the light over the range read from the sky of unlabelled frames takes the miss rate from 66% to 61%, within a point of the 62% the street's exact light gives, and randomising wider than that buys no accuracy while the program's own error climbs from 7% to 17%. With the pixels right, what the program draws starts to matter in a way it did not: a scene repair that was worth a little over 5 points while the pictures were wrong is worth 15 now. And the scene is where the program is silent about whole kinds of case: every pedestrian it drew stood in the open and within 16 m. Which real cases does the scene never draw, and how do we count them?Cover the cases: A model can only learn cases its data contains: from a few outlined frames read each pedestrian's distance (the foot row) and whether they step out (an outline touching a nearer vehicle's), count the cases the scene program never draws, size the evidence each case needs, and widen the program until the missing mass is gone.Counting cases on a small labelled sample, with each pedestrian's distance read off the foot row of its outline and a nearer van touching it as the sign of a step-out, shows that the scene program draws cells holding only 42% of the street's pedestrians: 51% of them stand beyond the 16 m it ever drew and 17% step out from behind a van it never lets hide anyone. Setting the distance range and the step-out probability from 200 outlined frames takes the miss rate from 62% to about 52%, as far as the street's own pedestrian settings do, and the clutter, once the pedestrians are right, takes it to 47%. But a cell is a box, not a frequency: the step-outs nearest the shuttle, which leave the least time to stop, occur in one frame in 122, so a training set of 1,600 frames holds about 13 of them, and the detector trained with the street's scene misses 57% of them against 10% of the pedestrians as near in the open. How should a generator spend its renders on cases that rarely happen, and what does that do to the statistics the model learns?
06Counting cases on a small labelled sample, with each pedestrian's distance read off the foot row of its outline and a nearer van touching it as the sign of a step-out, shows that the scene program draws cells holding only 42% of the street's pedestrians: 51% of them stand beyond the 16 m it ever drew and 17% step out from behind a van it never lets hide anyone. Setting the distance range and the step-out probability from 200 outlined frames takes the miss rate from 62% to about 52%, as far as the street's own pedestrian settings do, and the clutter, once the pedestrians are right, takes it to 47%. But a cell is a box, not a frequency: the step-outs nearest the shuttle, which leave the least time to stop, occur in one frame in 122, so a training set of 1,600 frames holds about 13 of them, and the detector trained with the street's scene misses 57% of them against 10% of the pedestrians as near in the open. How should a generator spend its renders on cases that rarely happen, and what does that do to the statistics the model learns?The rare case: Draw the rare case on purpose and undo the distortion afterwards: importance weights restore the true frequencies, the effective sample size says what the oversampling bought, and the allocation follows from the variance.Drawing the rare case on purpose, 12 times as often as it occurs, gives the detector 160 examples of it instead of 13 for the same 1,600 renders, and what it makes of them depends on the weight each carries: as drawn they say the case is common, and it misses 8 points fewer of the rare case and 1 point more of everything else; weighted by p/q they restore the street, and the detector is the natural one at an effective sample size of 1,463 frames. The pictures, the cases and their frequencies are now right, and the training data and the exam agree about what is in a frame. They may still disagree about what counts: a pedestrian with two pixels showing is a pedestrian to the program and not to an exam that asks for a visible fraction, and the miss rate on the rare case moves by 19 points with the rule. What does a label mean, and where can the program's definition differ from the exam's?
Part IV · Labels and the contract
07Drawing the rare case on purpose, 12 times as often as it occurs, gives the detector 160 examples of it instead of 13 for the same 1,600 renders, and what it makes of them depends on the weight each carries: as drawn they say the case is common, and it misses 8 points fewer of the rare case and 1 point more of everything else; weighted by p/q they restore the street, and the detector is the natural one at an effective sample size of 1,463 frames. The pictures, the cases and their frequencies are now right, and the training data and the exam agree about what is in a frame. They may still disagree about what counts: a pedestrian with two pixels showing is a pedestrian to the program and not to an exam that asks for a visible fraction, and the miss rate on the rare case moves by 19 points with the rule. What does a label mean, and where can the program's definition differ from the exam's?What the label means: A label is a definition and the program's definitions are exact only for itself: write down which pixels, how much counts, how far, aligned how, counted how finely, at which instant, and measure each clause twice, as the target a detector is taught and as the rule that scores it, because the two doors move different numbers.The labels are definitions, and the places where two definitions disagree are now written down and measured: visible mask or whole silhouette, what counts as a pedestrian at all (the same detector scores 44.7% or 51.8% on the exam and 40.1% or 65.1% on the rare case, depending on the rule), depth along the axis or range along the ray (17% of the close step-outs by depth are farther than 12 m by range), a mask half a pixel off. The clauses that move a reported number are not the ones that move a detector: the visibility rule moves what is reported by up to 25 points and what is learned by about one, while a mask half a pixel off costs 2.3 points of training and the hidden part of a silhouette 1.6. All of it was settled in a renderer we wrote ourselves, where we chose every convention. A production pipeline uses a production renderer whose conventions we did not choose, and its defaults decide, silently, how a sliver is counted. How do we make a renderer we did not write obey the same contract?
08The labels are definitions, and the places where two definitions disagree are now written down and measured: visible mask or whole silhouette, what counts as a pedestrian at all (the same detector scores 44.7% or 51.8% on the exam and 40.1% or 65.1% on the rare case, depending on the rule), depth along the axis or range along the ray (17% of the close step-outs by depth are farther than 12 m by range), a mask half a pixel off. The clauses that move a reported number are not the ones that move a detector: the visibility rule moves what is reported by up to 25 points and what is learned by about one, while a mask half a pixel off costs 2.3 points of training and the hidden part of a silhouette 1.6. All of it was settled in a renderer we wrote ourselves, where we chose every convention. A production pipeline uses a production renderer whose conventions we did not choose, and its defaults decide, silently, how a sliver is counted. How do we make a renderer we did not write obey the same contract?The same contract in a real renderer: A production renderer has conventions nobody here chose: run Blender headless on the Street, compare it with the program pixel by pixel, and set every convention until the contract holds.A production renderer obeys the contract once each convention is set and checked: 9 of the 12 conventions we tested needed a setting other than the default, and each one failed silently before it was set. The tests check the renderer and nothing after it: a worker that wrote a quarter of a batch upside down passed all twelve and moved the exam by 1.7 points, hardly more than a new seed does. Once the factory runs unattended for a million frames nobody will look at the pictures. How do we know, automatically and every time it runs, that the data is the data we meant, and that the exam is not scoring the data against itself?
Part V · Trust and reality
09A production renderer obeys the contract once each convention is set and checked: 9 of the 12 conventions we tested needed a setting other than the default, and each one failed silently before it was set. The tests check the renderer and nothing after it: a worker that wrote a quarter of a batch upside down passed all twelve and moved the exam by 1.7 points, hardly more than a new seed does. Once the factory runs unattended for a million frames nobody will look at the pictures. How do we know, automatically and every time it runs, that the data is the data we meant, and that the exam is not scoring the data against itself?Replay, checks and leaks: Make the factory replayable and self-checking: a manifest from which any frame can be redrawn, invariants that catch silent bugs, and splits by scene lineage so that a model cannot be graded on what it has already seen.A factory with a manifest that replays any frame, invariants that catch 14 of 17 injected bugs and an audit of the manifest two more, and splits by scene lineage that stop a memorizing model from scoring its own training data gives a number we can trust to describe the program: the naive program passes every check and its detector still misses 80% of the street's pedestrians. The program is still a model of the street, wrong in ways no check inside it can see, and a few real labelled frames can say how wrong. How should M real frames be spent, to grade, to calibrate the program or to train the detector, and what is a real frame worth in synthetic ones?
10A factory with a manifest that replays any frame, invariants that catch 14 of 17 injected bugs and an audit of the manifest two more, and splits by scene lineage that stop a memorizing model from scoring its own training data gives a number we can trust to describe the program: the naive program passes every check and its detector still misses 80% of the street's pedestrians. The program is still a model of the street, wrong in ways no check inside it can see, and a few real labelled frames can say how wrong. How should M real frames be spent, to grade, to calibrate the program or to train the detector, and what is a real frame worth in synthetic ones?A little reality: Spend a few real labelled frames where they are worth most: to grade, to calibrate the program's parameters, or to train the detector, and price a real frame in synthetic frames.A real labelled frame is worth 26 frames of the calibrated program, and 400 of them are best spent calibrating and fine-tuning on 100 and grading with the other 300: the detector now misses 49% of the street's pedestrians where it first missed 80%, and the grade certifies at most 57%. That is a rate per frame, and the shuttle meets a pedestrian once per approach: first seen at 16 m, a pedestrian leaves 8 frames before the shuttle can no longer stop, and missing all 8 has probability 49% to the eighth power, 0.3%, if misses are independent and 49% if they repeat, which no set of single frames can say. Nor can any log of human driving say what braking would have changed, because the drivers who braked early are the ones who saw the pedestrian. Can the program say what would have happened had the shuttle braked?
Part VI · Action and proof
11A real labelled frame is worth 26 frames of the calibrated program, and 400 of them are best spent calibrating and fine-tuning on 100 and grading with the other 300: the detector now misses 49% of the street's pedestrians where it first missed 80%, and the grade certifies at most 57%. That is a rate per frame, and the shuttle meets a pedestrian once per approach: first seen at 16 m, a pedestrian leaves 8 frames before the shuttle can no longer stop, and missing all 8 has probability 49% to the eighth power, 0.3%, if misses are independent and 49% if they repeat, which no set of single frames can say. Nor can any log of human driving say what braking would have changed, because the drivers who braked early are the ones who saw the pedestrian. Can the program say what would have happened had the shuttle braked?What braking changes: A program can branch where a log cannot: replay the same step-out with and without braking, share the random numbers between the branches, and say what a decision changes.Because a program can branch, the same step-out can be replayed with and without braking, and with the random numbers shared between the branches the difference between two detectors needs about 7 times fewer runs than two independent campaigns; the stopping margin then separates two detectors that 1,000 episodes of collisions cannot (2.9 standard errors against 1.6), while the exam, which grades frames, cannot see the decision rule that moves one detector's collisions from 9.3% to 20.0%. Every number so far is still a statement about the program, and the program is the thing under test: change only the pedestrian's walking speed, keeping its mean, and the closest band's collision rate moves by 19 points. What would count as proof that the system is safe on the real street, and how much real driving does that proof cost?
12Because a program can branch, the same step-out can be replayed with and without braking, and with the random numbers shared between the branches the difference between two detectors needs about 7 times fewer runs than two independent campaigns; the stopping margin then separates two detectors that 1,000 episodes of collisions cannot (2.9 standard errors against 1.6), while the exam, which grades frames, cannot see the decision rule that moves one detector's collisions from 9.3% to 20.0%. Every number so far is still a statement about the program, and the program is the thing under test: change only the pedestrian's walking speed, keeping its mean, and the closest band's collision rate moves by 19 points. What would count as proof that the system is safe on the real street, and how much real driving does that proof cost?Proof: Decide what counts as evidence on the real street: the power of a real test for a rare event, the validity of a simulator as an evaluator, and the ledger that says which of the gaps remain.A factory that makes labelled frames, a ledger that prices each stage's gap, a pipeline that can be trusted, a handful of real frames to calibrate with, and an exam only the real street can grade: what remains is a budget. A claim that a step-out ends in a collision less than once in a thousand takes 2,995 failure-free real step-outs, 29,950 hours at an assumed one step-out in ten, and no program shortens that; a program with the street's light and scene ranks twelve candidate systems as the street does (rank correlation 0.99 and 1.00), and showing so takes about 450 real step-outs for each of them. Of all the real observations the series' ledger asks of the street, 85% are such step-outs, and 99% if the failure is ten times rarer, while the frames and labels stay where they were. Every real hour, rendered frame and human label has a price and a value, and the proof of safety sets how many real hours cannot be avoided. How should the budget be spent, and how is a model of the street trained when frames are the cheap part and consequences the expensive one?
How to read this
Straight through, in order: the track is one argument, and each lesson's "Linear position" box names the problem it inherits, the single idea it adds and the problem it hands on. If you have time only for the spine, read 01 → 02 → 03 → 05 → 07 → 10 → 11 → 12. Each lesson has one widget that runs its mechanism, and a "What to try" paragraph that tells you which control to move first. To re-check every quoted number yourself, run python3 tools/chain/validate_chain.py --series syn --widgets --oracles from the repository root: it re-derives each one with an independent script.

Where this track sits

Before it: Computer Graphics, from first principles is the renderer run forwards, and Computer Vision · 04 Cameras and projection geometry is the camera model the Street uses. Beside it: 3D Vision, from first principles uses the same rays to reconstruct a scene rather than to draw one. After it: World Models, from first principles asks how to train a model that predicts what happens next when an agent acts, with synthetic frames as one of its data sources, and Training a Robot Model spends the budget this track sizes: what a real hour, a rendered frame and a human label each cost, and what each is worth.