The same contract in a real renderer
Lesson 7 settled the contract clause by clause in a renderer we wrote. A production team builds its scenes in one it did not write. This lesson runs Blender 5.2.2 headless on the Street's own scenes: one analytic scene per clause, rendered with the setting at Blender's default and then set, then whole frames, then trained detectors and the exam. 9 of the 12 defaults fail, none with a warning. With all twelve set, the production renderer reproduces the Street's pixel to 0.039% of full scale, and 1,600 of its frames train a detector that misses 46.5% of the street's pedestrians, where the Street's own renderer gives 46.8%. What no clause can say is whether the frames a factory wrote are the frames Blender drew.
New idea: a renderer obeys a contract only where each clause is set explicitly and checked against a truth known in closed form. The defaults are someone else's contract: 9 of 12 clauses failed at Blender's default, the worst costing 28.3 points of miss rate, and 4 of the 8 the exam can grade costing less than its seed wobble.
Forces next: A production renderer obeys the contract once each convention is set and checked: 9 of the 12 conventions we tested needed a setting other than the default, and each one failed silently before it was set. The tests check the renderer and nothing after it: a worker that wrote a quarter of a batch upside down passed all twelve and moved the exam by 1.7 points, hardly more than a new seed does. Once the factory runs unattended for a million frames nobody will look at the pictures. How do we know, automatically and every time it runs, that the data is the data we meant, and that the exam is not scoring the data against itself?
1 · One scene, two renderers
The Street describes a scene as a list of parts: a box or a vertical cylinder with its position, size, height range [y0, y1] and colour, plus the direction of the sun, the ambient term and the sky. That description is all a renderer needs, so Blender can draw what the Street draws. Take 24 scenes of the exact-stage program bbba (the street's scene, light and camera, the program's own label rule: Lesson 5) and build them in Blender with every setting at its default, except two that nobody can leave: the camera stands 1.4 m up and looks along +z (a new camera looks straight down), and the pixel buffer is read top row first (Blender's buffers start at the bottom). Each surface is shaded by the Street's own radiance formula, evaluated by shader nodes from the geometry's normal, so that whatever differs between the renderers is a convention and not a lighting model. The sensor stays ours: Blender supplies radiance and SV.sense turns it into pixel values.
Call the error of a frame the mean over its pixels of |LBlender − LStreet|, averaged over the three colour channels, in units of full-scale radiance, with the Street's side drawn at 16 × 16 rays per pixel. The default build is off by 7.2% on average and 33% at the 99th percentile, and 68% of its pixels are more than 1% wrong. The differences have names: a field of view of 39.6° where the Street has 60.0°, the horizon on row 12 where the Street has it on row 9, every part half sunk into the road, a tone curve that squeezes the highlights. Trained on 1,600 frames of this build (64 samples each instead of 4,096, to keep a batch affordable), the fixed detector misses 64.9% of the street's pedestrians, against 46.8% for the Street's own renderer.
The pictures are plausible: a street, seen a little closer. Some of the differences are conventions, some are mistakes of our own scene builder, and a picture does not say which. How do we know when none is left?
2 · A renderer is done when closed forms agree
The answer cannot be a look at the pictures: nobody will look at a million, and the clauses that hurt may not show. It is a conformance suite: for each clause, one scene whose result can be computed on paper, rendered with the setting at Blender's default and with it set, and a number that is zero when the clause holds. A clause holds when that error is under a tolerance the camera would not notice: positions 0.05 px (a 14th of the sensor's blur of 0.7 px), radiance 1% of full scale (a quarter of the shot noise of the brightest pixel, 1/√(600·0.95) = 4.2%), a focal length 0.5%, an area 5%. For example a post 0.3 m wide at x = +2 m, z = 10 m has the corners of its silhouette at columns 63.15 and 66.14, so its centre must be at column 64.64.
| Decision | Why |
|---|---|
| Radiance, not pictures: scene-linear float output | a picture is what colour management made of the radiance; the camera of Lesson 3 is calibrated on radiance, so Blender supplies radiance and SV.sense does the rest |
| Shading by formula: every surface an emitter whose colour is the Street's radiance formula, evaluated from the normal | renderers differ in lighting models by design, and a test must see conventions only. These frames are not photorealistic and are not meant to be |
| Masks by construction: the visible mask is the alpha of a render in which every other object is a holdout; the silhouette is the alpha of the pedestrian alone | a mask pass is a renderer's own convention (clause 11); a holdout render is coverage by definition |
One more thing is ours to choose: how the Street's frame maps into Blender's. The Street's axes (x right, y up, z forward) form a left-handed frame and Blender's world (X right, Y forward, Z up) a right-handed one, so the map (x, y, z) → (X, Y, Z) = (x, z, y), which swaps two axes, is a reflection of determinant −1: it changes the hand without mirroring the picture. A rotation in its place would mirror the street.
3 · Twelve clauses, twelve tests
Each row is a clause of the contract, what Blender does by default, the scene that tests it, and the error measured at the default and with the setting set, in the unit of the row. Three clauses hold at the default: pixel centres, the depth pass and units. The other 9 fail, each silently: no error, no warning, a valid frame.
| Clause | Blender's default | Test scene: what it measures | Error: default → set | At the default |
|---|---|---|---|---|
| 1 axes, pose | camera looks down | posts at x = +2 m (z = 10 m) and x = −2 m (z = 5 m): their centre columns | not in view → 0.003 px | fails |
| 2 row order | buffer starts at the bottom row | sky over ground: the boundary row, seen from the top of the buffer | 9.00 px → 0.001 px | fails |
| 3 focal length | 50 mm, 36 mm, auto fit | slab edges at x = ±1.5 m, z = 10 m: their separation is 3f/10 | +60.4% → 0.05% | fails |
| 4 principal point | no shift | sky over ground: the horizon row | 3.00 px → 0.001 px | fails |
| 5 pixel centres | corner convention, centre at +½ | a slab edge placed at u = 40.25: where the pixels put it | 0.000 px | holds |
| 6 shapes | origin at the centre, flat normals, 32 sides | a cylinder 1.7 m tall at 10 m (its top row); a wide cylinder's shading | 7.03 px → 0.02 px; 4.4% → 0.59% | fails |
| 7 what a pixel holds | 8-bit PNG through AgX | flat frames of radiance 0.02 to 0.95, read back as radiance | 42.7% → 0.0001% | fails |
| 8 pixel filter | Blackman-Harris, 1.5 px | a step edge at 16 sub-pixel offsets: its edge-spread function | 0.117 of the step leaks → 0 | fails |
| 9 sampler | 4096 samples, adaptive, denoiser on | an edge and a 0.3 px sliver, rendered both ways | 3.2% → 0.3% | fails |
| 10 depth pass | distance along the axis | a ground plane and a wall at z = 20 m against the ground-plane law | 0.19% | holds |
| 11 object masks | Object Index pass (the usual pick), one sample per pixel | slivers 0.3, 0.3 and 2.3 px wide, 3 m tall, at 20 m: their area in px² | 248% → 3.2% | fails |
| 12 units | metre, scale 1 | a wall at 10 m: the depth of the centre pixel | 0.0000002 m | holds |
Where things are (1–5). The mirrored map puts the post at x = +2 m at column 31.36 instead of 64.64: a picture of the street, reversed. The camera is a pinhole whose focal length in pixels is f = lens·W/sensor width for a horizontal fit, so f = 83.1 px needs a lens of 83.1·36/96 = 31.16 mm; the default 50 mm gives 133.3 px. The fit AUTO puts the sensor width on the longer side: it equals HORIZONTAL in a 96 × 24 frame but not in a 24 × 96 one, where the same lens gives 83.09 px under AUTO and 20.82 px under HORIZONTAL (measured), which is why camera.angle is not "the horizontal field of view". The shift is in units of the longer side and positive moves the picture down, so the horizon sits on row H/2 + shift·96; the default gives row 12.00 and row 9 needs shift = −3/96 (the last 0.001 px is the end of the ground plane, 100 km out). Blender counts pixels from their corners, pixel i covering [i, i + 1), as the Street does: the edge placed at 40.25 is found at 40.250. A toolbox that put pixel centres on integers would be half a pixel off.
How a pixel is made (6–9). A step edge placed at 16 sub-pixel offsets, read through one pixel, is the filter's edge-spread function; its derivative is the line-spread function and its spread the filter's σ. The box filter gives 0.288 px (1/√12 = 0.289) whatever width is asked: the BOX type ignored 1.5 and gave 0.288. The default gives 0.417 px: Cycles' Blackman-Harris window spans twice the width asked, 3 px for 1.5, and a window of that span has σ = 0.416 (integrated here). That is an extra 0.30 px, which with the sensor's blur of 0.7 px makes 0.815 px where the contract has 0.757; at a pixel boundary it moves 0.117 of the step into the neighbour, and a 0.5 px sliver's total reads 0.516 where the box gives 0.504.
The colour clause is the sharpest: AgX leaves mid-grey 0.18 at 0.180 and bends the ends, −42.7% at 0.95 and −28.2% at 0.02, so a test on grey would pass; the Standard transform in 8 bits is within 0.07%, float output exact. The denoiser is the one default that moves pixels: beside an edge it overshoots by 3.2%, and the default sampler takes 7 times as long; adaptive sampling alone changes nothing.
With the contract's settings (64 samples, no adaptive sampling, no denoiser) a render is reproducible: the same twice, one thread against eight, identical to the last bit, and the same twice even with the defaults, on this machine; a new seed changes a pixel by up to 0.031 of full scale.
Shapes (6). Blender's cylinder has 2 units of depth, the full height, its origin at the centre, 32 sides and flat normals; placed with its origin at its base it is half buried and a pedestrian is 7.5 rows tall instead of 14.5. Flat normals err by 4.4% across a wide cylinder, smooth ones by 0.59%; 32 sides leave the width 0.20% short (the worst case is 0.48%).
What the extra channels mean (10–12). The Depth pass is distance along the axis: it follows the ground-plane law z = f·hc/(v + ½ − v0) to 0.19% over 864 ground pixels. Read as range it would be wrong by up to 14.2% at the corner of the field (range = depth·√(1 + tan²φ + ρ²), Lesson 7). The Mist pass is the range: it follows the quadratic of the range to within 0.003 and the quadratic of the depth only to 0.165, on a factor that runs from 0 to 1. Where nothing is hit Depth reads 1010 where the Street's reads 0. The Object Index pass takes one sample at each pixel centre: three slivers whose true areas are 3.74, 3.74 and 28.67 px² read 0, 13 and 39: the first is missed whole, the second inflated to a column of 13 pixels. The holdout alpha reads 3.63, 3.86 and 28.81, and Cryptomatte, the coverage-weighted pass, gives the same numbers. The unit scale changes what the interface calls a unit and not the number in the pass: the wall reads 9.95 at scale 1 and at 0.01.
4 · From the defaults to conformance
With all twelve set, the 24 scenes agree with the Street's pixel to 0.039% on average (99th percentile 0.66%, 0.3% of the pixels over 1%), 183 times closer than the default build. Against the Street's own pixel of 2 × 2 rays the error is 0.28%, because that pixel is itself 0.27% from the area integral the contract asks of a pixel (Lesson 7): the renderer we wrote is the less accurate one. What is left is small. Two sample patterns of the same Blender frame differ by 0.023%; at 1,024 samples the frames are still 0.035% from the area integral, so what remains is not sampling but what the two geometries and the reference's own 16 × 16 grid disagree about. Mask areas agree to 1.4% on average (3% at worst) over the 11 pedestrians that count, silhouettes to 1.1%.
The widget walks two of those scenes from the defaults to conformance, setting the clauses in the order they fail, loudest first.
What to try. (1) Leave k at 0 and read scene A: the camera looks at the ground and the error is 4.412%. Set one convention (the axes): the street appears and the error rises to 7.767%, every post and the pedestrian drawn in the wrong place. (2) Slide to 3, the focal length now right: 7.533%, no nearer than at 2 (7.338%). A field of view of 60° with the horizon on row 12 is no closer to the Street's picture than 39.6° was, so the mean error is not a count of the conventions that hold. (3) At 4 the horizon is on row 9 and the error falls to 4.116% with 67.4% of the pixels moving; at 5 nothing moves, because the pixel centres were right at the default. (4) At 6, the shapes, it falls to 0.467%: 34.1% of the pixels change and the 99th percentile drops from 19.99% to 1.06%. (5) Steps 7, 8 and 9 take it to 0.089, 0.058 and 0.036%; the output moves 85.8% of the pixels, the filter 21.7%, the sampler 13.1%. Steps 10 to 12 move nothing: all three held. (6) Switch to scene B, the step-out behind the van: its error is 9.506% at step 2, 0.838% at step 6 with a 99th percentile of 5.63%, five times scene A's, and 0.043% at step 12. (7) Set the scale to 0.2% and slide through 7, 8 and 9: what is left is on the edges, 67% of the remaining error at step 7 on the 28% of the pixels that are edges.
5 · Train on production frames, and price each silent failure
The lab can say what a failed clause costs, as a detector's miss rate. Take the exact-stage program bbba and replace only its renderer: the same 1,600 scene draws as the Street-rendered cell, Blender's radiance with all twelve set, the sensor of Lesson 3 (the same noise stream as the Street's own frame of that seed), and the label from Blender's holdout alpha (the program's rule: a frame has a pedestrian when at least one pixel of the silhouette shows). The fixed detector is trained on those frames and graded on the street as every cell of the series is: the threshold lets 10% of the real validation frames without pedestrians alarm, and the miss rate is taken over the 1,323 real-test pedestrians with at least 6 visible pixels. Three seeds give 46.1, 46.2 and 47.1%, mean 46.5%, against 46.7, 46.3 and 47.5% (46.8%) for the Street's own renderer: a difference of −0.4 points, inside the seed range of either (1.0 and 1.1 points). Set by the contract, the production renderer is the street's program to within a seed.
Now leave one clause at its default and set the other eleven (one seed; the all-set cell of the same seed, 46.1%, is the reference). Clause 1 has no row: with the camera looking down there is no pedestrian to grade.
| Clause left at its default | Its error (§3) | Miss rate | Points over the reference |
|---|---|---|---|
| focal length (50 mm, auto fit) | +60.4% | 74.4% | +28.3 |
| row order (bottom row first) | 9.00 px | 52.1% | +6.0 |
| shapes: origin at the centre | 7.03 px | 48.4% | +2.3 |
| principal point (no shift) | 3.00 px | 47.6% | +1.5 |
| output (8-bit AgX) | 42.7% | 46.6% | +0.5 |
| masks (Object Index) | 248% | 46.6% | +0.5 |
| pixel filter (Blackman-Harris 1.5) | 0.117 | 46.4% | +0.3 |
| shapes: flat normals | 4.4% | 46.1% | +0.0 |
| sampler: denoiser on | 3.2% | 46.0% | −0.1 |
| all of them (pose and row order set, nothing else) | §1 | 64.9% | +18.8 |
The price is not in proportion to the error. The focal length alone costs 28.3 points and the row order 6.0; four clauses, the output, the masks, the filter and the denoiser, leave the exam inside its seed range of 1.0 point, though each fails its own test by a wide margin (the masks by 248% of an area). A team that watched only the exam would have seen the lens, the row order, the horizon and the shapes' origin, four of the eight it can grade, and called the rest seed noise. The prices do not add: all of them together cost 18.8 points, less than the lens alone, because each wrong clause changes what the others do to the picture.
6 · What a frame costs
A batch of 1,600 conforming frames (radiance and the pedestrian's alpha of each) takes about a minute in Cycles on this CPU: 63, 108 and 65 s for the three seeds, the middle one while the machine was busy. The same batch with the denoiser left on takes 301 s, 4.8 times as long. Timed frame by frame, best of several passes, a conforming frame takes 19 ms and one with every setting at its default, 4,096 samples included, 410 ms, 21 times as long. Part of that has nothing to do with a convention: with a sky built from nodes Cycles builds an importance map of the world at every change of the scene, and a frame with it takes 278 ms, 14 times as long as without. The contract sets the world's sampling method to none, which changes no pixel (the largest difference over six scenes is 0).
With more samples the error falls almost as 1/spp, faster than the 1/√spp of independent samples: between 4 and 256 samples the rms error against a 4,096-sample render of the same frame (another seed) has a log-log slope of −0.93, 50 times lower for 64 times the samples. At 64 samples it is 0.102% of full scale, 41 times under the sensor's own shot noise of 4.2%, so more samples buy nothing the camera would keep. With one sample a pixel is a point sample at its centre: the edge window of clause 9 is off by 70%, the 1 − 0.3 that a centre sample predicts for the sliver.
EEVEE, the real-time engine, in one measured paragraph. It renders on the GPU here, and a frame takes 258 ms, 13 times a conforming Cycles frame at this size. It delivers Cryptomatte, depth and mist, but no Object Index. Its pixel filter (filter_size, default 1.5) leaks like Cycles' default: 0.123 of a step at a pixel boundary, 0.040 at 1.0, none at 0, a point sample. Its frame differs from the 4,096-sample Cycles frame of the same scene by 0.65% on average and 13.3% at the 99th percentile.
A manifest sketch, the entry to the next lesson. To replay a frame one records the Blender version and build, the twelve settings, a hash of the scene description and the seed; not the thread count, since one thread and eight gave identical pixels. Files are compared by their decoded pixels, because an EXR file carries time stamps. The claim reaches as far as what was measured: this build, this machine, these settings.
7 · What no clause tests
The twelve tests check the renderer. A factory is more than a renderer: scenes go in, a worker writes arrays, another reads them, a detector trains. Take the 1,600 conforming frames of §5 and let a worker with another habit write a share of them bottom row first, image and mask together: the first 5% of every hundred (80 frames), then the first 25% (400 frames). None of the twelve tests runs through that writer, so all twelve pass. The frames are still streets, upside down, with the sky at the bottom and the horizon on row 15 where it belongs on row 9; their labels were flipped with them, so the same 781 frames hold a pedestrian.
Train and grade as in §5. With 5% written upside down the exam reads 46.1%, the all-set cell's number (0.0 points). With a quarter it reads 47.8%, 1.7 points worse, hardly more than the seed range of 1.0. A team that watched the exam would have called a quarter of its data upside down a bad seed.
The twelve tests passed, the exam did not notice, and the data are not the data we meant. Nothing in the pipeline looks at a frame. A check that looked at every frame would see it at once: the horizon is on row 9 in every frame the camera takes, and this batch breaks that in 400 of 1,600. But one check for one failure is a patch, and the same pipeline has other ways to be wrong: a stale file, a mask shifted by a pixel, scenes of the exam that the detector had already seen. Once the factory runs unattended for a million frames nobody will look at the pictures.
Common mistakes / failure modes
camera.angle is the horizontal field of view"Checkpoint exercise
shift_y puts the horizon on row 24? What is the horizontal field of view if the same lens renders a 64 × 256 portrait frame? Answer: AUTO puts the sensor width on the longer side, 256 px, so f = 20·256/36 = 142.2 px. The horizon sits on row H/2 + shift·S with S = 256, so 24 = 32 + 256·shift and shift = −0.03125 (negative moves the picture up). In the portrait frame the longer side is 256 again, so f is the same 142.2 px and the horizontal field of view is 2 atan(32/142.2) = 25.4°, not the 2 atan(18/20) = 84.0° that camera.angle reports. Under the HORIZONTAL fit f would be 20·64/36 = 35.6 px.Where this points next
The production renderer now obeys the contract: nine conventions set, three that held, and a detector trained on its frames misses 46.5% of the street's pedestrians where the Street's own renderer gives 46.8%. The tests that got it there check the renderer and nothing after it. A worker that wrote a quarter of a batch upside down passed all twelve and moved the exam by 1.7 points, hardly more than a new seed does. Once the factory runs unattended for a million frames nobody will look at the pictures. How do we know, automatically and every time it runs, that the data is the data we meant, and that the exam is not scoring the data against itself?
Interview prompts
- Your team moves data generation from a renderer it wrote to Blender. What do you check before you trust a frame, and why not look at pictures? (§2 — a conformance suite: one scene per convention whose answer is known in closed form, run at the default and set; pictures hide the quiet clauses, and nobody looks at a million.)
- A detector trained on the new renderer's frames misses 28.3 points more than one trained on the old. How do you find the cause? (§3, §5 — compare whole frames for the loud clauses, then set one clause at a time and re-grade; here the focal length alone cost that.)
- How do you tell whether a pass is coverage or a point sample? (§3 — render a sliver narrower than a pixel that misses its centre: the Object Index pass loses it, a holdout alpha or Cryptomatte reads its area.)
- What is the difference between depth and range, and which does each pass give? (§3 — depth is along the axis (Z), range along the ray (Mist); they differ by up to 14.2% at the corner of a 60° field.)
- Why did the exam not see four of the eight failures it could grade? (§5 — each moved the miss rate by under a point, the seed wobble; the tests have no noise floor and found them.)
- What must a team record to replay a Blender frame? (§6 — the version and build, the twelve settings, the scene description and the seed; not the thread count; compare decoded pixels, not file bytes.)
- A worker writes a quarter of a batch upside down. Which of your checks fire? (§7 — none of the renderer's twelve tests, and the exam moves 1.7 points; only a check run on every frame can.)
Companion reads: Computer Graphics · 06 Sampling and antialiasing (the pixel filter), Computer Graphics · 14 Color, HDR and tone mapping (what a view transform does to radiance) and Computer Vision · 04 Cameras and projection geometry (the pinhole whose conventions the first five clauses pin down).