all_lessons/Synthetic Vision Data/08 · The same contract in a real rendererlesson 8 / 12

The same contract in a real renderer

Lesson 7 settled the contract clause by clause in a renderer we wrote. A production team builds its scenes in one it did not write. This lesson runs Blender 5.2.2 headless on the Street's own scenes: one analytic scene per clause, rendered with the setting at Blender's default and then set, then whole frames, then trained detectors and the exam. 9 of the 12 defaults fail, none with a warning. With all twelve set, the production renderer reproduces the Street's pixel to 0.039% of full scale, and 1,600 of its frames train a detector that misses 46.5% of the street's pedestrians, where the Street's own renderer gives 46.8%. What no clause can say is whether the frames a factory wrote are the frames Blender drew.

The thesis, here
A renderer you did not write answers a list of questions with defaults somebody else chose, and a wrong answer still draws a picture. The contract is that list of questions. A renderer obeys it only where each answer is set explicitly and then tested against a scene whose correct result is known in closed form. The tests tell, and the pictures and the exam do not: half of the failures the exam can grade move it by less than its seed wobble.
Linear position
Forced by: The labels are definitions, and the places where two definitions disagree are now written down and measured: visible mask or whole silhouette, what counts as a pedestrian at all (the same detector scores 44.7% or 51.8% on the exam and 40.1% or 65.1% on the rare case, depending on the rule), depth along the axis or range along the ray (17% of the close step-outs by depth are farther than 12 m by range), a mask half a pixel off. The clauses that move a reported number are not the ones that move a detector: the visibility rule moves what is reported by up to 25 points and what is learned by about one, while a mask half a pixel off costs 2.3 points of training and the hidden part of a silhouette 1.6. All of it was settled in a renderer we wrote ourselves, where we chose every convention. A production pipeline uses a production renderer whose conventions we did not choose, and its defaults decide, silently, how a sliver is counted. How do we make a renderer we did not write obey the same contract?
New idea: a renderer obeys a contract only where each clause is set explicitly and checked against a truth known in closed form. The defaults are someone else's contract: 9 of 12 clauses failed at Blender's default, the worst costing 28.3 points of miss rate, and 4 of the 8 the exam can grade costing less than its seed wobble.
Forces next: A production renderer obeys the contract once each convention is set and checked: 9 of the 12 conventions we tested needed a setting other than the default, and each one failed silently before it was set. The tests check the renderer and nothing after it: a worker that wrote a quarter of a batch upside down passed all twelve and moved the exam by 1.7 points, hardly more than a new seed does. Once the factory runs unattended for a million frames nobody will look at the pictures. How do we know, automatically and every time it runs, that the data is the data we meant, and that the exam is not scoring the data against itself?
The plan
Seven moves. (1) Build one scene in both renderers and see what the defaults do. (2) Say when a renderer is done: closed-form tests, radiance and not pictures, masks by construction. (3) Run the twelve tests, each with the setting at its default and set. (4) Close the loop on whole frames, along the path from the defaults to conformance. (5) Train on the production frames and price each silent failure in points. (6) Count what a frame costs. (7) Find what no clause tests.

1 · One scene, two renderers

The Street describes a scene as a list of parts: a box or a vertical cylinder with its position, size, height range [y0, y1] and colour, plus the direction of the sun, the ambient term and the sky. That description is all a renderer needs, so Blender can draw what the Street draws. Take 24 scenes of the exact-stage program bbba (the street's scene, light and camera, the program's own label rule: Lesson 5) and build them in Blender with every setting at its default, except two that nobody can leave: the camera stands 1.4 m up and looks along +z (a new camera looks straight down), and the pixel buffer is read top row first (Blender's buffers start at the bottom). Each surface is shaded by the Street's own radiance formula, evaluated by shader nodes from the geometry's normal, so that whatever differs between the renderers is a convention and not a lighting model. The sensor stays ours: Blender supplies radiance and SV.sense turns it into pixel values.

Call the error of a frame the mean over its pixels of |LBlender − LStreet|, averaged over the three colour channels, in units of full-scale radiance, with the Street's side drawn at 16 × 16 rays per pixel. The default build is off by 7.2% on average and 33% at the 99th percentile, and 68% of its pixels are more than 1% wrong. The differences have names: a field of view of 39.6° where the Street has 60.0°, the horizon on row 12 where the Street has it on row 9, every part half sunk into the road, a tone curve that squeezes the highlights. Trained on 1,600 frames of this build (64 samples each instead of 4,096, to keep a batch affordable), the fixed detector misses 64.9% of the street's pedestrians, against 46.8% for the Street's own renderer.

The pictures are plausible: a street, seen a little closer. Some of the differences are conventions, some are mistakes of our own scene builder, and a picture does not say which. How do we know when none is left?

2 · A renderer is done when closed forms agree

The answer cannot be a look at the pictures: nobody will look at a million, and the clauses that hurt may not show. It is a conformance suite: for each clause, one scene whose result can be computed on paper, rendered with the setting at Blender's default and with it set, and a number that is zero when the clause holds. A clause holds when that error is under a tolerance the camera would not notice: positions 0.05 px (a 14th of the sensor's blur of 0.7 px), radiance 1% of full scale (a quarter of the shot noise of the brightest pixel, 1/√(600·0.95) = 4.2%), a focal length 0.5%, an area 5%. For example a post 0.3 m wide at x = +2 m, z = 10 m has the corners of its silhouette at columns 63.15 and 66.14, so its centre must be at column 64.64.

DecisionWhy
Radiance, not pictures: scene-linear float outputa picture is what colour management made of the radiance; the camera of Lesson 3 is calibrated on radiance, so Blender supplies radiance and SV.sense does the rest
Shading by formula: every surface an emitter whose colour is the Street's radiance formula, evaluated from the normalrenderers differ in lighting models by design, and a test must see conventions only. These frames are not photorealistic and are not meant to be
Masks by construction: the visible mask is the alpha of a render in which every other object is a holdout; the silhouette is the alpha of the pedestrian alonea mask pass is a renderer's own convention (clause 11); a holdout render is coverage by definition

One more thing is ours to choose: how the Street's frame maps into Blender's. The Street's axes (x right, y up, z forward) form a left-handed frame and Blender's world (X right, Y forward, Z up) a right-handed one, so the map (x, y, z) → (X, Y, Z) = (x, z, y), which swaps two axes, is a reflection of determinant −1: it changes the hand without mirroring the picture. A rotation in its place would mirror the street.

3 · Twelve clauses, twelve tests

Each row is a clause of the contract, what Blender does by default, the scene that tests it, and the error measured at the default and with the setting set, in the unit of the row. Three clauses hold at the default: pixel centres, the depth pass and units. The other 9 fail, each silently: no error, no warning, a valid frame.

ClauseBlender's defaultTest scene: what it measuresError: default → setAt the default
1 axes, posecamera looks downposts at x = +2 m (z = 10 m) and x = −2 m (z = 5 m): their centre columnsnot in view → 0.003 pxfails
2 row orderbuffer starts at the bottom rowsky over ground: the boundary row, seen from the top of the buffer9.00 px → 0.001 pxfails
3 focal length50 mm, 36 mm, auto fitslab edges at x = ±1.5 m, z = 10 m: their separation is 3f/10+60.4% → 0.05%fails
4 principal pointno shiftsky over ground: the horizon row3.00 px → 0.001 pxfails
5 pixel centrescorner convention, centre at +½a slab edge placed at u = 40.25: where the pixels put it0.000 pxholds
6 shapesorigin at the centre, flat normals, 32 sidesa cylinder 1.7 m tall at 10 m (its top row); a wide cylinder's shading7.03 px → 0.02 px; 4.4% → 0.59%fails
7 what a pixel holds8-bit PNG through AgXflat frames of radiance 0.02 to 0.95, read back as radiance42.7% → 0.0001%fails
8 pixel filterBlackman-Harris, 1.5 pxa step edge at 16 sub-pixel offsets: its edge-spread function0.117 of the step leaks → 0fails
9 sampler4096 samples, adaptive, denoiser onan edge and a 0.3 px sliver, rendered both ways3.2% → 0.3%fails
10 depth passdistance along the axisa ground plane and a wall at z = 20 m against the ground-plane law0.19%holds
11 object masksObject Index pass (the usual pick), one sample per pixelslivers 0.3, 0.3 and 2.3 px wide, 3 m tall, at 20 m: their area in px²248% → 3.2%fails
12 unitsmetre, scale 1a wall at 10 m: the depth of the centre pixel0.0000002 mholds

Where things are (1–5). The mirrored map puts the post at x = +2 m at column 31.36 instead of 64.64: a picture of the street, reversed. The camera is a pinhole whose focal length in pixels is f = lens·W/sensor width for a horizontal fit, so f = 83.1 px needs a lens of 83.1·36/96 = 31.16 mm; the default 50 mm gives 133.3 px. The fit AUTO puts the sensor width on the longer side: it equals HORIZONTAL in a 96 × 24 frame but not in a 24 × 96 one, where the same lens gives 83.09 px under AUTO and 20.82 px under HORIZONTAL (measured), which is why camera.angle is not "the horizontal field of view". The shift is in units of the longer side and positive moves the picture down, so the horizon sits on row H/2 + shift·96; the default gives row 12.00 and row 9 needs shift = −3/96 (the last 0.001 px is the end of the ground plane, 100 km out). Blender counts pixels from their corners, pixel i covering [i, i + 1), as the Street does: the edge placed at 40.25 is found at 40.250. A toolbox that put pixel centres on integers would be half a pixel off.

How a pixel is made (6–9). A step edge placed at 16 sub-pixel offsets, read through one pixel, is the filter's edge-spread function; its derivative is the line-spread function and its spread the filter's σ. The box filter gives 0.288 px (1/√12 = 0.289) whatever width is asked: the BOX type ignored 1.5 and gave 0.288. The default gives 0.417 px: Cycles' Blackman-Harris window spans twice the width asked, 3 px for 1.5, and a window of that span has σ = 0.416 (integrated here). That is an extra 0.30 px, which with the sensor's blur of 0.7 px makes 0.815 px where the contract has 0.757; at a pixel boundary it moves 0.117 of the step into the neighbour, and a 0.5 px sliver's total reads 0.516 where the box gives 0.504.

The colour clause is the sharpest: AgX leaves mid-grey 0.18 at 0.180 and bends the ends, −42.7% at 0.95 and −28.2% at 0.02, so a test on grey would pass; the Standard transform in 8 bits is within 0.07%, float output exact. The denoiser is the one default that moves pixels: beside an edge it overshoots by 3.2%, and the default sampler takes 7 times as long; adaptive sampling alone changes nothing.

With the contract's settings (64 samples, no adaptive sampling, no denoiser) a render is reproducible: the same twice, one thread against eight, identical to the last bit, and the same twice even with the defaults, on this machine; a new seed changes a pixel by up to 0.031 of full scale.

Shapes (6). Blender's cylinder has 2 units of depth, the full height, its origin at the centre, 32 sides and flat normals; placed with its origin at its base it is half buried and a pedestrian is 7.5 rows tall instead of 14.5. Flat normals err by 4.4% across a wide cylinder, smooth ones by 0.59%; 32 sides leave the width 0.20% short (the worst case is 0.48%).

What the extra channels mean (10–12). The Depth pass is distance along the axis: it follows the ground-plane law z = f·hc/(v + ½ − v0) to 0.19% over 864 ground pixels. Read as range it would be wrong by up to 14.2% at the corner of the field (range = depth·√(1 + tan²φ + ρ²), Lesson 7). The Mist pass is the range: it follows the quadratic of the range to within 0.003 and the quadratic of the depth only to 0.165, on a factor that runs from 0 to 1. Where nothing is hit Depth reads 1010 where the Street's reads 0. The Object Index pass takes one sample at each pixel centre: three slivers whose true areas are 3.74, 3.74 and 28.67 px² read 0, 13 and 39: the first is missed whole, the second inflated to a column of 13 pixels. The holdout alpha reads 3.63, 3.86 and 28.81, and Cryptomatte, the coverage-weighted pass, gives the same numbers. The unit scale changes what the interface calls a unit and not the number in the pass: the wall reads 9.95 at scale 1 and at 0.01.

4 · From the defaults to conformance

With all twelve set, the 24 scenes agree with the Street's pixel to 0.039% on average (99th percentile 0.66%, 0.3% of the pixels over 1%), 183 times closer than the default build. Against the Street's own pixel of 2 × 2 rays the error is 0.28%, because that pixel is itself 0.27% from the area integral the contract asks of a pixel (Lesson 7): the renderer we wrote is the less accurate one. What is left is small. Two sample patterns of the same Blender frame differ by 0.023%; at 1,024 samples the frames are still 0.035% from the area integral, so what remains is not sampling but what the two geometries and the reference's own 16 × 16 grid disagree about. Mask areas agree to 1.4% on average (3% at worst) over the 11 pedestrians that count, silhouettes to 1.1%.

The widget walks two of those scenes from the defaults to conformance, setting the clauses in the order they fail, loudest first.

From the defaults to conformance
Slide the number of conventions set (the first k of the twelve; the rest stay at Blender's defaults). Top to bottom: the Street's frame at 16 × 16 rays per pixel (drawn here), the Blender frame recorded with k conventions set, and their difference. The bars give each clause's test error at its default (red) and set (teal) in units of its tolerance; the clause just set is highlighted. Blender's frames and conformance scenes are recorded input; every error, statistic and bar is computed here.
clause just set
—
mean error
—
99th percentile
—
pixels moved by this step
—
its test: default → set
—
at the default
—
Show the core JS
L.boxCentre = function (x, z, hx, hz) { var us = [], a, b; for (a = -1; a <= 1; a += 2) for (b = -1; b <= 1; b += 2) us.push(W / 2 + F * (x + a * hx) / (z + b * hz)); return (Math.min.apply(null, us) + Math.max.apply(null, us)) / 2; };
  function f(v, x0) { var Wd = v.red.length, dl = sum(v.red), dr = Wd - sum(v.green); return (dr - dl) * 10 / (2 * x0); }
M.shift = function (T) { var t = L.horizon(), d = sum(T.default), s = sum(T.set); return { d: Math.abs(d - t), s: Math.abs(s - t), vDefault: d, vSet: s, truth: t, tol: L.TOL.px, unit: 'px' }; };
  for (i = 1; i < pts.length; i++) { var dG = pts[i][1] - pts[i - 1][1], x0 = pts[i - 1][0], x1 = pts[i][0]; tot += dG; m1 += dG * (x0 + x1) / 2; m2 += dG * (x0 * x0 + x0 * x1 + x1 * x1) / 3; }
  return { sigma: Math.sqrt(m2 - m1 * m1), leak: esf[0][5], mass: tot };
function winErr(v) { var m = 0, k; for (k = 0; k < 9; k++) m = Math.max(m, Math.abs(v.w40[k] - W40[k]), Math.abs(v.w20[k] - W20[k])); return m; }
L.errMap = function (a, b) { var e = new Float32Array(HW), q; for (q = 0; q < HW; q++) e[q] = (Math.abs(a[q] - b[q]) + Math.abs(a[HW + q] - b[HW + q]) + Math.abs(a[2 * HW + q] - b[2 * HW + q])) / 3; return e; };
L.fails = function (m) { return m.d === null || m.d > m.tol; };

What to try. (1) Leave k at 0 and read scene A: the camera looks at the ground and the error is 4.412%. Set one convention (the axes): the street appears and the error rises to 7.767%, every post and the pedestrian drawn in the wrong place. (2) Slide to 3, the focal length now right: 7.533%, no nearer than at 2 (7.338%). A field of view of 60° with the horizon on row 12 is no closer to the Street's picture than 39.6° was, so the mean error is not a count of the conventions that hold. (3) At 4 the horizon is on row 9 and the error falls to 4.116% with 67.4% of the pixels moving; at 5 nothing moves, because the pixel centres were right at the default. (4) At 6, the shapes, it falls to 0.467%: 34.1% of the pixels change and the 99th percentile drops from 19.99% to 1.06%. (5) Steps 7, 8 and 9 take it to 0.089, 0.058 and 0.036%; the output moves 85.8% of the pixels, the filter 21.7%, the sampler 13.1%. Steps 10 to 12 move nothing: all three held. (6) Switch to scene B, the step-out behind the van: its error is 9.506% at step 2, 0.838% at step 6 with a 99th percentile of 5.63%, five times scene A's, and 0.043% at step 12. (7) Set the scale to 0.2% and slide through 7, 8 and 9: what is left is on the edges, 67% of the remaining error at step 7 on the 28% of the pixels that are edges.

Road not taken · compare the pictures
The tempting check is a picture metric: render both, subtract, stop when the mean error is small. It rewards the loudest clauses and hides the quiet ones. In scene A the mean error is 0.467% after step 6, under 1%, and a team that stopped there would leave the output, the filter and the sampler at their defaults: three of the nine failures, and none of them visible at that scale. A mean over pixels is also not what a detector reads: a sliver's area, wrong by 248% under the Object Index, is a handful of pixels in 2,304 and moves the mean by nothing. The road returns as the whole-frame comparison of this section, kept as the last check after the tests.

5 · Train on production frames, and price each silent failure

The lab can say what a failed clause costs, as a detector's miss rate. Take the exact-stage program bbba and replace only its renderer: the same 1,600 scene draws as the Street-rendered cell, Blender's radiance with all twelve set, the sensor of Lesson 3 (the same noise stream as the Street's own frame of that seed), and the label from Blender's holdout alpha (the program's rule: a frame has a pedestrian when at least one pixel of the silhouette shows). The fixed detector is trained on those frames and graded on the street as every cell of the series is: the threshold lets 10% of the real validation frames without pedestrians alarm, and the miss rate is taken over the 1,323 real-test pedestrians with at least 6 visible pixels. Three seeds give 46.1, 46.2 and 47.1%, mean 46.5%, against 46.7, 46.3 and 47.5% (46.8%) for the Street's own renderer: a difference of −0.4 points, inside the seed range of either (1.0 and 1.1 points). Set by the contract, the production renderer is the street's program to within a seed.

Now leave one clause at its default and set the other eleven (one seed; the all-set cell of the same seed, 46.1%, is the reference). Clause 1 has no row: with the camera looking down there is no pedestrian to grade.

Clause left at its defaultIts error (§3)Miss ratePoints over the reference
focal length (50 mm, auto fit)+60.4%74.4%+28.3
row order (bottom row first)9.00 px52.1%+6.0
shapes: origin at the centre7.03 px48.4%+2.3
principal point (no shift)3.00 px47.6%+1.5
output (8-bit AgX)42.7%46.6%+0.5
masks (Object Index)248%46.6%+0.5
pixel filter (Blackman-Harris 1.5)0.11746.4%+0.3
shapes: flat normals4.4%46.1%+0.0
sampler: denoiser on3.2%46.0%−0.1
all of them (pose and row order set, nothing else)§164.9%+18.8

The price is not in proportion to the error. The focal length alone costs 28.3 points and the row order 6.0; four clauses, the output, the masks, the filter and the denoiser, leave the exam inside its seed range of 1.0 point, though each fails its own test by a wide margin (the masks by 248% of an area). A team that watched only the exam would have seen the lens, the row order, the horizon and the shapes' origin, four of the eight it can grade, and called the rest seed noise. The prices do not add: all of them together cost 18.8 points, less than the lens alone, because each wrong clause changes what the others do to the picture.

Road not taken · let the exam find the bad clauses
The exam is the number the project cares about, so let it do the finding: change one thing, retrain, re-grade. It finds the focal length and the row order. It misses four of the eight it can grade, each inside the seed range: the filter (0.3 points), the output (0.5), the masks (0.5) and the denoiser (−0.1). It also costs a retraining per clause and per experiment, against a noise floor of a point; a test costs a render and has no floor. The road returns as a price list: the tests say whether a clause failed, the exam says what the failure cost this detector on this exam, and another detector would price it differently.

6 · What a frame costs

A batch of 1,600 conforming frames (radiance and the pedestrian's alpha of each) takes about a minute in Cycles on this CPU: 63, 108 and 65 s for the three seeds, the middle one while the machine was busy. The same batch with the denoiser left on takes 301 s, 4.8 times as long. Timed frame by frame, best of several passes, a conforming frame takes 19 ms and one with every setting at its default, 4,096 samples included, 410 ms, 21 times as long. Part of that has nothing to do with a convention: with a sky built from nodes Cycles builds an importance map of the world at every change of the scene, and a frame with it takes 278 ms, 14 times as long as without. The contract sets the world's sampling method to none, which changes no pixel (the largest difference over six scenes is 0).

With more samples the error falls almost as 1/spp, faster than the 1/√spp of independent samples: between 4 and 256 samples the rms error against a 4,096-sample render of the same frame (another seed) has a log-log slope of −0.93, 50 times lower for 64 times the samples. At 64 samples it is 0.102% of full scale, 41 times under the sensor's own shot noise of 4.2%, so more samples buy nothing the camera would keep. With one sample a pixel is a point sample at its centre: the edge window of clause 9 is off by 70%, the 1 − 0.3 that a centre sample predicts for the sliver.

EEVEE, the real-time engine, in one measured paragraph. It renders on the GPU here, and a frame takes 258 ms, 13 times a conforming Cycles frame at this size. It delivers Cryptomatte, depth and mist, but no Object Index. Its pixel filter (filter_size, default 1.5) leaks like Cycles' default: 0.123 of a step at a pixel boundary, 0.040 at 1.0, none at 0, a point sample. Its frame differs from the 4,096-sample Cycles frame of the same scene by 0.65% on average and 13.3% at the 99th percentile.

A manifest sketch, the entry to the next lesson. To replay a frame one records the Blender version and build, the twelve settings, a hash of the scene description and the seed; not the thread count, since one thread and eight gave identical pixels. Files are compared by their decoded pixels, because an EXR file carries time stamps. The claim reaches as far as what was measured: this build, this machine, these settings.

7 · What no clause tests

The twelve tests check the renderer. A factory is more than a renderer: scenes go in, a worker writes arrays, another reads them, a detector trains. Take the 1,600 conforming frames of §5 and let a worker with another habit write a share of them bottom row first, image and mask together: the first 5% of every hundred (80 frames), then the first 25% (400 frames). None of the twelve tests runs through that writer, so all twelve pass. The frames are still streets, upside down, with the sky at the bottom and the horizon on row 15 where it belongs on row 9; their labels were flipped with them, so the same 781 frames hold a pedestrian.

Train and grade as in §5. With 5% written upside down the exam reads 46.1%, the all-set cell's number (0.0 points). With a quarter it reads 47.8%, 1.7 points worse, hardly more than the seed range of 1.0. A team that watched the exam would have called a quarter of its data upside down a bad seed.

The twelve tests passed, the exam did not notice, and the data are not the data we meant. Nothing in the pipeline looks at a frame. A check that looked at every frame would see it at once: the horizon is on row 9 in every frame the camera takes, and this batch breaks that in 400 of 1,600. But one check for one failure is a patch, and the same pipeline has other ways to be wrong: a stale file, a mask shifted by a pixel, scenes of the exam that the detector had already seen. Once the factory runs unattended for a million frames nobody will look at the pictures.

What this lesson did not do
The frames are emission-shaded geometry: Blender's materials, shadows and light transport were not exercised, so nothing here says how a photoreal Blender frame differs from the street's, and these are not meant to be. The suite asks twelve questions of one version of one renderer on one CPU (Blender 5.2.2 LTS, Cycles, 64 samples); another version, a GPU or another denoiser must be asked again, and reproducibility was measured on this machine only. EEVEE got a paragraph and no suite. The price list is one detector on one exam, one seed per row. What a manifest must hold, what to check on every frame, and how the exam can score data it has already seen: Lesson 9. What a few real frames do to the program: Lesson 10.

Common mistakes / failure modes

"The picture looks like the street, so the camera is right"
The default build was a plausible street, 7.2% off on average, and its focal length alone cost 28.3 points of miss rate (§1, §5).
"camera.angle is the horizontal field of view"
It depends on the lens and the sensor only. In a 24 × 96 frame the same lens gives 83.09 px under AUTO and 20.82 px under HORIZONTAL (§3).
"The object-index pass is a mask"
It is a point sample at the pixel centre: slivers of 3.74, 3.74 and 28.67 px² read 0, 13 and 39. A holdout alpha or Cryptomatte reads their area (§3).
"A PNG is what Blender rendered, so it is radiance"
An 8-bit AgX picture is display-referred: radiance 0.95 comes back 42.7% low while mid-grey comes back exact, so a test on grey passes (§3).
"Leave the denoiser on, it only cleans"
Beside an edge it overshoots by 3.2%, and the exam does not see it (−0.1 points); a test with a known edge does (§3, §5).
"Change the handedness with a rotation"
A rotation cannot change the hand; the one that keeps determinant +1 mirrors the street. The map that swaps y and z has determinant −1 and keeps the picture (§2).
"If the exam did not move, the data is fine"
Four failing clauses moved it by under 1.0 point, and a quarter of the batch upside down by 1.7 (§5, §7).

Checkpoint exercise

Try it
A frame is 256 × 64 pixels, the lens is 20 mm on the default 36 mm sensor, and the fit is AUTO. What is the focal length in pixels, and what shift_y puts the horizon on row 24? What is the horizontal field of view if the same lens renders a 64 × 256 portrait frame? Answer: AUTO puts the sensor width on the longer side, 256 px, so f = 20·256/36 = 142.2 px. The horizon sits on row H/2 + shift·S with S = 256, so 24 = 32 + 256·shift and shift = −0.03125 (negative moves the picture up). In the portrait frame the longer side is 256 again, so f is the same 142.2 px and the horizontal field of view is 2 atan(32/142.2) = 25.4°, not the 2 atan(18/20) = 84.0° that camera.angle reports. Under the HORIZONTAL fit f would be 20·64/36 = 35.6 px.

Where this points next

The production renderer now obeys the contract: nine conventions set, three that held, and a detector trained on its frames misses 46.5% of the street's pedestrians where the Street's own renderer gives 46.8%. The tests that got it there check the renderer and nothing after it. A worker that wrote a quarter of a batch upside down passed all twelve and moved the exam by 1.7 points, hardly more than a new seed does. Once the factory runs unattended for a million frames nobody will look at the pictures. How do we know, automatically and every time it runs, that the data is the data we meant, and that the exam is not scoring the data against itself?

Takeaway
A renderer you did not write has defaults you did not choose, and a wrong default still draws a plausible street. 9 of the 12 conventions of the contract failed at Blender's defaults, silently, and the way to find them was not to look: it was a scene per clause whose answer is known in closed form, run at the default and set, with tolerances taken from the camera's own noise. Set by the contract, the production renderer reproduces the Street's pixel to 0.039% of full scale and trains a detector within a seed of the Street's own, 46.5% against 46.8%. The exam prices a failure and is a poor alarm: the focal length cost 28.3 points, and four clauses cost less than its seed range of 1.0, because the price is that of one detector on one exam. A batch written upside down passed all twelve tests and moved the exam by 1.7 points, so the data path needs checks of its own, run on every frame.

Interview prompts

Companion reads: Computer Graphics · 06 Sampling and antialiasing (the pixel filter), Computer Graphics · 14 Color, HDR and tone mapping (what a view transform does to radiance) and Computer Vision · 04 Cameras and projection geometry (the pinhole whose conventions the first five clauses pin down).