An hour is not an hour: exchange rates
Lesson 15 closed on a programme with one robot, scarce robot time and data of several kinds to spend it on. Counting hours does not compare them: on the Bench, 8 layouts of your own plus the same 0.66 hours from each of five sources leave one policy succeeding on 5.8 % to 98.8 % of fresh layouts. This lesson builds the unit that does, the exchange rate: how many hours of your own demonstrations an hour of the source replaces, read off your own success curve. An arm of the same make is worth about one, an older arm a third, footage a sixth, a simulator from nearly one to less than nothing as its error grows, and all the demonstrations without force together less than one demonstration with it. A rate is not a price.
New idea: an hour of any source is put in one unit, own-demonstration hours, by reading the pooled policy's success off your own curve: ρ = (Neq − n) / m. The rate belongs to an operating point, can be negative, and falls to zero as the hours grow when the source lacks a column the task needs.
Forces next: Measured on the Bench, an hour of another arm's demonstrations replaces a fraction of an hour of your own, an hour of labelled video a smaller fraction, an hour of simulation a fraction that depends on the gap, and a source that lacks a column the task needs replaces none of it, however many hours there are. Exchange rates put every source in one unit, but not in dollars. What does an hour of each source cost?
1 · The same hours, different data
The programme that lesson 15 left must spend scarce robot time on data of several kinds. Before asking how much of each, we need a way to count them; the cheapest place to test one is the layouts task of lessons 7 and 8: a two-link arm carries a cup past five posts to a mat, a layout is a shift and a tilt of the row, and the policy copies the stored demonstration of the nearest layout. Success is the share of 1,000 fresh layouts, one run each under gusts, on which the cup reaches the mat untouched. One demonstration is one layout, 9.3 s of motion in 186 frames of 0.05 s, so an hour holds 387.1 demonstrations; resets, discards and operator time are lesson 17's.
Your arm holds n layouts, and a source adds m attempted layouts, of which the laboratory keeps those that succeeded. The sources: an arm of the same make (the twin); an older arm from another laboratory, slower and less precise (lesson 8); footage of a person, whose hand wobbles more than the arm (lesson 9) and is tracked to 3 mm; simulators whose arm model is off by a fraction of a link and that label frames in joint angles (lessons 8 and 11). The tally: 8 own layouts plus 256 attempted layouts of each, which at 9.3 s a layout is 39.7 minutes.
| 8 own layouts plus 256 attempted layouts of | kept | frames stored | success on 1,000 fresh layouts |
|---|---|---|---|
| nothing (your eight alone) | 0 | 0 | 21.6 % |
| an arm of the same make | all 256 | 47,479 | 98.8 % |
| an older arm from another laboratory | 197 | 44,816 | 71.7 % |
| footage of a person | 180 | 39,495 | 56.0 % |
| a simulator, arm model off by 1 % | all 256 | 47,475 | 82.3 % |
| a simulator, arm model off by 3 % | all 256 | 47,460 | 5.8 % |
With the same arm, policy, test layouts and hours, success runs from 5.8 % to 98.8 %. The worst row is below the 21.6 % that the 8 layouts reached alone: those hours did harm.
2 · What a unit has to be
A unit must be (i) the same for every source; (ii) measured on this policy and task, since it states what hours do for this learner; (iii) spendable, so that the programme can convert the hours it buys into hours it would otherwise record; (iv) valid where the programme stands, because the next hour adds less as the policy improves. Four candidates fail:
| unit | what goes wrong |
|---|---|
| hours of motion | The tally of §1: the same hours, 5.8 % to 98.8 %. |
| frames stored | Frames are hours again. At 64 attempted layouts the simulators off by 0.1 % and 3 % store 11,873 and 11,867 frames and reach 80.3 % and 12.5 %. |
| error on held-out frames | Lesson 1: the copy error stayed between 2.9 and 4.4 % while success fell from 87.5 to 42.0 %; lesson 9: the label with the smallest error did not decide. |
| dollars per hour | Not yet. A price says what an hour costs to make, not what it does: two simulators run at one speed whatever the error of their arm model. Lesson 17 prices hours; the unit must first say what they are worth. |
One thing survives all four: your own demonstrations. A programme can always measure their success curve on its own robot, so the unit is measured on this policy and task; they are what the money would buy otherwise, so it is spendable; and the curve is a ruler on which the success of any source can be located, wherever the programme stands. So ask of any pooled set: how many own layouts would have given this success?
3 · Reading a rate off your own curve
Your own curve is the success of a policy built from N own layouts, on the same 1,000 test layouts as everything below. It reads 22.9 % at one layout, 23.1 % at 2, 8.9 % at 4, 21.6 % at 8, 54.9 % at 32, 69.8 % at 48 and 99.0 % at 256. It saturates, gaining 24.1 points from 32 to 64 layouts, 14.3 from 64 to 128 and 5.7 from 128 to 256, and it is not smooth at tiny N: four layouts do worse than two, since a few can leave a gap wherever they fall, so the inversion reads its running maximum, which can only rise.
Pool n own layouts with m attempted layouts of a source and measure the success S. Let Neq be the own-layout count whose own success is S, found by straight lines between the measured points and, below the first, by a line from the origin. The pooled set did what Neq own layouts do; n of them were already yours, so the source supplied Neq − n, and per attempted layout, that is per source hour in the unit of §1, the exchange rate is
ρ = (Neq − n) / m
Take the older arm with n = 8 and m = 64, which is 9.9 minutes of its motion. The pooled policy succeeds on 55.1 % of the test layouts; your own curve passes 54.9 % at 32 layouts and 69.8 % at 48, so it reaches 55.1 % at 32.2. The pooled set does what 32.2 own layouts do, the source supplied 24.2 of them, and ρ = 24.2 / 64 = 0.38: 9.9 minutes of the older arm did the work of 3.8 minutes of yours.
At ρ = 1 the source is as good as your own, at 0 it is worthless (S is what your n layouts reach on the ruler), below 0 it did harm. There is a floor, because Neq cannot be below 0: ρ ≥ −n / m, here −0.125. Near n = 8 the ruler is flat at 23.1 % from 2 layouts to 8, so even the 21.6 % of your eight alone reads Neq = 0.9 and ρ = −0.11: below 23.1 % the rate says harm and the success says how much.
The test layouts give S a Wilson interval, 52.0 to 58.2 %. The running maximum is monotone, so the interval passes through the inversion: Neq from 29.1 to 35.5, ρ from 0.33 to 0.43. That is the sampling error of 1,000 test layouts and nothing more: it leaves out the noise of the own curve, read at one point of one draw, and the draw of the layouts, which §4 measures.
The widget
What to try. Leave the defaults, the older arm with 8 own layouts and 64 attempted: success 55.1 %, your own curve reaches it at 32.2 layouts, ρ = 0.38. Slide the hours to 16 and to 256: ρ = 0.76 and 0.17, a fall the right panel shows. Choose 32 own layouts: 0.17 at 64. Choose the arm of the same make: 1.01 at 64, and the teal curve lies on the grey one. Step through the simulators from 0.1 to 3 %: ρ = 0.93, 0.83, 0.42 and −0.12, with joint-label gaps of 0.08, 0.23, 0.76 and 2.27 cm; at 3 % the dashed line runs to the left edge, because Neq is 0.5 layouts, below the 8 you hold; with hand labels the same simulator gives 1.02. Last, set the task to the peg: with force one demonstration gives 82.5 % and ten give 98.5 %; without it forty give 63.0 %, which the with-force curve passes at 0.76 of a demonstration, so ρ = 0.019.
4 · What the rates are
The table gives each source's rate at 64 attempted layouts, on the draw of layouts the widget uses and averaged over twelve other draws of both sets of layouts, own and source, on the same test layouts.
| source | 8 own layouts | 32 own layouts | |
|---|---|---|---|
| this draw | twelve draws | twelve draws | |
| an arm of the same make | 1.01 | 1.05 | 1.10 |
| an older arm | 0.38 | 0.29 | 0.16 |
| footage of a person | 0.16 | 0.14 | 0.01 |
| simulator off by 0.1 % | 0.93 | 0.97 | 1.03 |
| simulator off by 1 % | 0.42 | 0.39 | 0.28 |
| simulator off by 3 % | −0.12 | −0.08 | −0.29 |
| the same, hand labels | 1.02 | 1.05 | 1.09 |
Read the instrument first. The twin is a control: its demonstrations come from the same distribution as yours, so a sound instrument reads about 1. It reads 1.01 on this draw, with an interval of 0.90 to 1.12 from the test layouts, and 0.81 to 1.29 over twelve draws, 1.05 ± 0.24. That spread comes from which layouts were drawn and is wider than the interval from the test layouts: rates closer together are not ordered by one draw. These orderings hold on every draw, at 8 own layouts and at 32: twin above older arm above footage, and simulators falling as the gap grows.
The older arm's rate falls as you take more of it: 0.76 at 16 attempted layouts, 0.38 at 64, 0.17 at 256. Its curve flattens below yours, ending at 71.7 % where yours reaches 99.0 %, and at 64 attempted layouts 22 % of its attempts are discarded. Footage is worth 0.16: a person's hand wobbles more than the arm, the tracker adds 3 mm, and 36 % of the 64 attempts are discarded, so the rate is 0.26 per kept layout and 0.16 per attempted layout, which the laboratory pays for.
A simulator never discards, and its rate follows its gap. A joint-angle label carries the simulator's link lengths (lesson 8): one link longer by δ, the other shorter. When your arm tracks the labels its hand lands off the simulator's by 2|δ| sin(q2 / 2), q2 being the elbow angle, since the two link directions differ by q2: 0.08, 0.23, 0.76 and 2.27 cm for the simulators off by 0.1, 0.3, 1 and 3 %, against the 1.32 cm by which the demonstrated paths clear the posts. At 0.57 of that clearance the rate is 0.42; at 1.72 times it the 64 simulated layouts leave the policy at 12.5 % against the 21.6 % of its 8 alone, where 0.5 of one own layout would, and ρ = −0.12 is next to the floor of −8 / 64 and to the −0.11 of no change (§3). With hand positions as labels the same frames are worth 1.02: the gap enters through the label.
The operating point. On this draw every rate is lower at 32 own layouts, but over twelve draws that is three effects. A source as good as your own is worth the same however much you hold (the twin, 1.05 at 8 own and 1.10 at 32). A mediocre source is worth less the more you hold (the older arm 0.29 and 0.16). A harmful source harms more (the 3 % simulator −0.08 and −0.29), because there is more to destroy and the floor moves from −0.125 to −0.5. A rate belongs to a source, a policy, a task, the n you hold and the m you take.
5 · A column that no number of hours buys
Everything so far was a source with every column the task needs, differing in quality. Lesson 10 met the other case: a peg in a hole with a millimetre of clearance, where holding position inserts 63.5 % of attempts and only a policy that feels the force and gives way does better. A clone learns that only from demonstrations that recorded force; drop the column and the clone is the position policy.
| demonstrations m | success with force | success without force | with-force demonstrations with the same success |
|---|---|---|---|
| 1 | 82.5 % | 51.5 % | 0.62 |
| 5 | 82.5 % | 66.5 % | 0.81 |
| 10 | 98.5 % | 64.0 % | 0.78 |
| 40 | 98.5 % | 63.0 % | 0.76 |
Both clones come from the same m demonstrations of the yielding policy, on 200 attempts. The with-force curve is the ruler (a line from the origin below its first point); the clone without force is used alone, since it cannot read force from frames that carry none and nothing can be pooled. Forty demonstrations without force reach 63.0 %, which the ruler passes at 0.76 of a with-force demonstration, so ρ = 0.76 / 40 = 0.019. At no m on the grid (1 to 40) do they reach the 82.5 % that one demonstration with force gives (the best is 66.5 %), so together they replace at most 0.81 of a demonstration and ρ(m) ≤ 0.81 / m goes to zero.
That is a structural zero, and it differs in kind from a low rate. Footage is worth little because its hours are worse than yours, and its curve still climbs: 38.0, 44.8 and 56.0 % at 64, 128 and 256 attempted layouts with 8 own. The clone without force is flat at 63.0 to 66.5 % from two demonstrations to forty, its ceiling the position policy's, which no frame without force corrects; ten demonstrations with force give 98.5 %. A column the task needs cannot be bought from a source that lacks it.
6 · A rate is not a price
A rate is a ratio of two quantities of one kind, own motion hours per source motion hour: it has no currency and no clock. What it says about money is a ceiling. An hour of the source replaces ρ own hours, so it is worth paying at most ρ · pown, where pown is what an own hour costs; pay more and you do better to record your own.
| source | most an hour is worth paying (share of an own hour's price) | what an hour costs |
|---|---|---|
| an arm of the same make | 1.01 | ? |
| an older arm | 0.38 | ? |
| footage of a person | 0.16 | ? |
| simulator off by 3 % | −0.12 | no price is worth paying |
The right-hand column is what the table lacks, and without it ranking by rate cannot choose. The twin ranks first, but an hour of it is worth 2.7 hours of the older arm, so it is the better buy only if it costs less than 2.7 times as much; footage is worth buying only below a sixth of an own hour's price; the 3 % simulator's ceiling is negative, as if you had to be paid 0.12 of an own hour's price for each hour you took.
Common mistakes / failure modes
Checkpoint exercise
Where this points next
Measured on the Bench, an hour of an older arm's demonstrations replaces 0.38 of an hour of your own, an hour of footage 0.16, an hour of simulation 0.93 down to −0.12 as its arm model's error grows from 0.1 to 3 %, and an hour of demonstrations without the force column 0.019 at forty and less with more. The unit says nothing about cost: an hour of the twin is worth 2.7 hours of the older arm, and whether that makes it the better buy depends on two prices the rate does not contain; ρ · pown is only the most an hour is worth paying. Exchange rates put every source in one unit, but not in dollars. What does an hour of each source cost?
Interview prompts
- Why can't two data sources be compared by their hours or their frames? (§1, §2 — the same hours gave 5.8 to 98.8 %, and equal frames hide rates a whole own hour apart.)
- How would you measure what an hour of another source is worth to your policy? (§3 — pool it with your own, read the success off your own curve, and take (Neq − n) / m.)
- A source measures at 1.2. Is it better than your own data? (§4 — not necessarily: the twin, which cannot be better, reads 0.81 to 1.29 over twelve draws.)
- Can an exchange rate be negative, and how low can it go? (§3, §4 — yes, when the source does harm; Neq cannot fall below 0, so ρ ≥ −n / m, which the 3 % simulator sits next to.)
- Why does a mediocre source's rate fall as m grows, and as n grows? (§4 — its curve flattens below yours, so its next layout is worth less than its first.)
- What separates a low exchange rate from a structural zero? (§5 — a low rate keeps rising with hours; a missing column caps the total below one demonstration with force, so ρ falls like 1 / m.)
- Why is the exchange rate not enough to decide a purchase? (§6 — it gives only the ceiling ρ · pown; the asking price and pown are separate numbers.)
Companion reads: Lesson 8 · Other bodies (the first exchange rate) and Lesson 10 · What cameras cannot see (the peg and force).