all_lessons/ World Models/ 27 · the mixture is the modellesson 27 / 31

The mixture is the model

Mixture weights get set by intuition in an afternoon and then decide more about the final model than the architecture did. The reason they are hard is not that the trade-offs are subtle. It is that the thing you want to up-weight does not exist in the quantity you want.

Where we are
Lesson 26 gave us five stages, each wanting different data. Now we specify what "different data" means, and hit the wall lesson 17 predicted: the source that carries the handle has a hundred hours, the source that carries coverage has a million. A mixture weight is a request, and availability is what decides whether it can be honoured.
Forced by 26Five stages, each with a different data requirement. This stepDesign the mixture against real availability: up-weighting a scarce source only re-reads it. Forces 28Broad pretraining buys coverage and loses control — hence post-training.

1 · What each source can and cannot teach

Think of it as a shopping list where two aisles have different problems. One aisle is enormous, nearly free, and stocks only half of what the recipe needs. The other stocks exactly the missing ingredient and has three hundred and fifty grams of it, total, worldwide. You can write any proportions you like on the list. The shop still only has what it has — and that gap between what you asked for and what exists is the whole subject.

Four sources, and the honest columns for each. Note that no source is good at everything, and the one that is good at the thing we need most is the one there is least of.

sourceavailablecoveragehandlereal physics
broad internet video~10⁶ hexcellentnoneyes
in-domain video (your scenes, no actions)~10⁴ hnarrow but relevantnoneyes
action-grounded interaction~10²–10³ hvery narrowfullyes
simulationunboundedwhatever you authorfullapproximate

Two observations do most of the work.

First, coverage and handle are anti-correlated across sources, for the same reason as lesson 17: the handle requires an instrumented actor, and instrumenting an actor is what makes data scarce. This is the mixture problem's whole structure.

Second, simulation is the only source that breaks the anti-correlation — unbounded quantity and a perfect action column — and it pays for that with the one column it cannot fake. Which is why the mixture question is largely "how much sim can I get away with," and why synthetic_vision/10 exists to answer it properly rather than by feel.

2 · Availability turns a weight into an epoch count

Here is the mechanism people miss. A mixture weight wi and a total budget B specify how many hours you will draw from source i. But you can only draw what exists. If the request exceeds the supply, the sampler simply returns the same data again:

epochsi = B · wi / availi

Put numbers in. Budget 50,000 hours, and you decide the grounded source deserves 8% because the handle matters:

50,000 × 0.08 = 4,000 hours requested  ÷  350 available = 11.4 epochs

You did not get 4,000 hours of grounded data. You got 350 hours, eleven times, plus eleven times the opportunity to memorise it. Up-weighting a scarce source does not add information; it adds repetition, and repetition pays on a completely different curve.

Model the diminishing return explicitly. The first pass over a source delivers its full information; subsequent passes deliver progressively less, roughly logarithmically:

effectivei = availi · (1 + ln(epochsi))   for epochsi > 1

At 11.4 epochs that gives 350 × (1 + 2.43) = 1,201 effective hours from a 4,000-hour request — a 3.3× discount. And the discount worsens as you push harder, which means there is a ceiling on how much handle you can buy with mixture weights alone. Past that ceiling the only remaining moves are to collect more (lessons 16 to 24 of the robot series) or to substitute simulation.

The mixture allocator, priced against availability
Weights are normalised, then converted to epoch counts against real availability (10⁶ / 10⁴ / 350 h / unbounded), then discounted logarithmically for repeats. Try raising the grounded share to chase controllability and watch the epoch warning fire.
Interactive four-way mixture allocator.
fidelity
—
controllability
—
real-world relevance
—
worst repetition
—
diagnosis
—

Now do the experiment that makes the point. Set grounded to 0 and simulation to 40. Controllability recovers — because sim has a perfect action column — but real-world relevance drops, because sim's physics is discounted. There is no setting of four sliders that maximises all three. That is not a limitation of the widget; it is the actual situation, and mixture design is choosing which of the three you are willing to be worst at.

3 · The three-way trade, stated as a decision

maximise fidelity
Lean broad video
Best-looking model, best generalisation to novel scenes. Ignores your controller. Use when the model is a perception prior, not a simulator.
maximise controllability
Lean sim + grounded
Obeys actions, works in the authored world, brittle outside it. Use when the deployment distribution is genuinely narrow — a fixed cell, a known task.
maximise relevance
Lean in-domain real
Right scenes, right physics, little coverage and no handle. Use for the final post-training pass, never for the bulk.

The recipe that actually ships is not one of these three, it is the sequence of them — which is exactly lesson 26's staging, now with data attached: bulk coverage from broad video early, handle from sim in the middle, relevance from a small in-domain grounded set at the end. Mixture is a function of training time, not a constant. That is the single most useful correction to how most projects set it.

4 · How to actually find the weights

You cannot grid-search a four-way mixture at full scale. Three things work, in ascending cost.

  1. Fit per-source scaling curves at small scale. Train small models on each source alone across several data scales, and fit an exponent and a ceiling per source. Each source has its own curve — and the ceiling matters more than the exponent, because it tells you where adding hours stops helping. This is cheap and reorders your priorities immediately.
  2. Sweep along a line, not over a grid. Fix the ratio between the two abundant sources and sweep only the scarce-source share. In practice one dimension carries most of the variance because the abundant sources are near-substitutes for each other.
  3. Match the mixture to the metric that will gate release. If controllability is the gate, evaluate mixtures on the paired-intervention test, not on validation loss. Optimising a mixture against loss reliably selects the mixture with the most broad video, because that is what minimises loss — and it is the mixture with the least handle. This mistake is common and it is invisible unless you look for it.
The sim-share trap
Simulation is unbounded, free, and perfectly labeled, so every optimiser you point at the mixture — human or automated — will push its share up. The check is not a feeling about realism; it is a held-out real evaluation set that never contains synthetic data, gating release. synthetic_vision/10 prices it in frames (grading is the largest bill, paid in real frames the model never trains on) and synthetic_vision/09 builds the split that keeps synthetic frames out of it; adopt both rather than reinventing them.
Takeaway
Coverage and handle are anti-correlated across sources because the handle requires an instrumented actor, and simulation is the only source that escapes the trade — at the cost of approximate physics. A mixture weight is a request: epochsi=B·wi/availi, and beyond one epoch the return is roughly avail·(1+ln epochs), so up-weighting a 350-hour source to 8% of a 50,000-hour budget yields 11.4 epochs and a 3.3× discount — repetition, not information. Fit per-source scaling curves, sweep the scarce share along a line, and never select a mixture on validation loss, which reliably picks the mixture with the least handle.

Where this points next

Bulk training on the broad mixture gives a model with excellent coverage and weak control — precisely the anti-correlation above, expressed in the weights. The standard remedy is a final pass on a small, curated, action-diverse set. Lesson 28 asks what that pass can actually accomplish, and finds a hard boundary: post-training moves behaviour freely, and cannot move the ceiling that lesson 20 set.

Interview prompts