The mixture is the model
Mixture weights get set by intuition in an afternoon and then decide more about the final model than the architecture did. The reason they are hard is not that the trade-offs are subtle. It is that the thing you want to up-weight does not exist in the quantity you want.
1 · What each source can and cannot teach
Think of it as a shopping list where two aisles have different problems. One aisle is enormous, nearly free, and stocks only half of what the recipe needs. The other stocks exactly the missing ingredient and has three hundred and fifty grams of it, total, worldwide. You can write any proportions you like on the list. The shop still only has what it has — and that gap between what you asked for and what exists is the whole subject.
Four sources, and the honest columns for each. Note that no source is good at everything, and the one that is good at the thing we need most is the one there is least of.
| source | available | coverage | handle | real physics |
|---|---|---|---|---|
| broad internet video | ~10⁶ h | excellent | none | yes |
| in-domain video (your scenes, no actions) | ~10⁴ h | narrow but relevant | none | yes |
| action-grounded interaction | ~10²–10³ h | very narrow | full | yes |
| simulation | unbounded | whatever you author | full | approximate |
Two observations do most of the work.
First, coverage and handle are anti-correlated across sources, for the same reason as lesson 17: the handle requires an instrumented actor, and instrumenting an actor is what makes data scarce. This is the mixture problem's whole structure.
Second, simulation is the only source that breaks the anti-correlation — unbounded quantity and a perfect action column — and it pays for that with the one column it cannot fake. Which is why the mixture question is largely "how much sim can I get away with," and why synthetic_vision/10 exists to answer it properly rather than by feel.
2 · Availability turns a weight into an epoch count
Here is the mechanism people miss. A mixture weight wi and a total budget B specify how many hours you will draw from source i. But you can only draw what exists. If the request exceeds the supply, the sampler simply returns the same data again:
epochsi = B · wi / availiPut numbers in. Budget 50,000 hours, and you decide the grounded source deserves 8% because the handle matters:
50,000 × 0.08 = 4,000 hours requested ÷ 350 available = 11.4 epochsYou did not get 4,000 hours of grounded data. You got 350 hours, eleven times, plus eleven times the opportunity to memorise it. Up-weighting a scarce source does not add information; it adds repetition, and repetition pays on a completely different curve.
Model the diminishing return explicitly. The first pass over a source delivers its full information; subsequent passes deliver progressively less, roughly logarithmically:
effectivei = availi · (1 + ln(epochsi)) for epochsi > 1At 11.4 epochs that gives 350 × (1 + 2.43) = 1,201 effective hours from a 4,000-hour request — a 3.3× discount. And the discount worsens as you push harder, which means there is a ceiling on how much handle you can buy with mixture weights alone. Past that ceiling the only remaining moves are to collect more (lessons 16 to 24 of the robot series) or to substitute simulation.
Now do the experiment that makes the point. Set grounded to 0 and simulation to 40. Controllability recovers — because sim has a perfect action column — but real-world relevance drops, because sim's physics is discounted. There is no setting of four sliders that maximises all three. That is not a limitation of the widget; it is the actual situation, and mixture design is choosing which of the three you are willing to be worst at.
3 · The three-way trade, stated as a decision
The recipe that actually ships is not one of these three, it is the sequence of them — which is exactly lesson 26's staging, now with data attached: bulk coverage from broad video early, handle from sim in the middle, relevance from a small in-domain grounded set at the end. Mixture is a function of training time, not a constant. That is the single most useful correction to how most projects set it.
4 · How to actually find the weights
You cannot grid-search a four-way mixture at full scale. Three things work, in ascending cost.
- Fit per-source scaling curves at small scale. Train small models on each source alone across several data scales, and fit an exponent and a ceiling per source. Each source has its own curve — and the ceiling matters more than the exponent, because it tells you where adding hours stops helping. This is cheap and reorders your priorities immediately.
- Sweep along a line, not over a grid. Fix the ratio between the two abundant sources and sweep only the scarce-source share. In practice one dimension carries most of the variance because the abundant sources are near-substitutes for each other.
- Match the mixture to the metric that will gate release. If controllability is the gate, evaluate mixtures on the paired-intervention test, not on validation loss. Optimising a mixture against loss reliably selects the mixture with the most broad video, because that is what minimises loss — and it is the mixture with the least handle. This mistake is common and it is invisible unless you look for it.
Where this points next
Bulk training on the broad mixture gives a model with excellent coverage and weak control — precisely the anti-correlation above, expressed in the weights. The standard remedy is a final pass on a small, curated, action-diverse set. Lesson 28 asks what that pass can actually accomplish, and finds a hard boundary: post-training moves behaviour freely, and cannot move the ceiling that lesson 20 set.
Interview prompts
- Why are coverage and handle anti-correlated across data sources? The handle requires an instrumented actor, and instrumenting an actor is exactly what limits quantity — so the sources with an action column are the narrow ones. (§1)
- What does simulation uniquely offer, and what does it cost? Unbounded quantity with a perfect action column, at the price of approximate physics — the one column it cannot fake. (§1)
- Convert an 8% mixture weight on a 350-hour source at a 50,000-hour budget into what you actually get. 4,000 hours requested over 350 available is 11.4 epochs; discounted as avail·(1+ln epochs) that is about 1,200 effective hours — a 3.3× penalty. (§2)
- Why does up-weighting a scarce source stop working past a point? Because it buys repetition rather than information, and repeats pay logarithmically, so there is a ceiling on how much handle mixture weights alone can purchase. (§2)
- Why should the mixture change over training rather than stay constant? Each stage has a different requirement — coverage early, handle in the middle, relevance last — so a single constant mixture is a compromise none of the stages wanted. (§3)
- Why is selecting a mixture on validation loss actively harmful? Loss is minimised by the most broad video, which is the mixture with the least action handle; select on the metric that gates release, such as paired-intervention controllability. (§4)
- What guards against sim-share creep? A held-out real evaluation set that never contains synthetic data and gates release, split by lineage as in synthetic_vision/09 and priced as in synthetic_vision/10. (§4)