all_lessons/Robot Model Training/16 · Exchange rateslesson 16 / 24

An hour is not an hour: exchange rates

Lesson 15 closed on a programme with one robot, scarce robot time and data of several kinds to spend it on. Counting hours does not compare them: on the Bench, 8 layouts of your own plus the same 0.66 hours from each of five sources leave one policy succeeding on 5.8 % to 98.8 % of fresh layouts. This lesson builds the unit that does, the exchange rate: how many hours of your own demonstrations an hour of the source replaces, read off your own success curve. An arm of the same make is worth about one, an older arm a third, footage a sixth, a simulator from nearly one to less than nothing as its error grows, and all the demonstrations without force together less than one demonstration with it. A rate is not a price.

The thesis, here
The same hours are worth anything from less than nothing to everything, so hours are not a unit. Every source can be held against a programme's own demonstrations: how many of yours would the same success have cost? That count per hour of the source is its exchange rate, and it belongs to a policy, a task and a point on the curve.
Linear position
Forced by: Twenty trials cannot tell seventy-six percent from sixty-eight, telling a ten-point improvement apart takes hundreds of trials per policy, and a programme that evaluates honestly spends more robot time on evaluation than on training. The programme also needs data of several kinds, each from a source with a different price, a different usefulness and a different shelf life. How much of each kind do we need, in what unit can they even be compared, and what does each cost?
New idea: an hour of any source is put in one unit, own-demonstration hours, by reading the pooled policy's success off your own curve: ρ = (Neq − n) / m. The rate belongs to an operating point, can be negative, and falls to zero as the hours grow when the source lacks a column the task needs.
Forces next: Measured on the Bench, an hour of another arm's demonstrations replaces a fraction of an hour of your own, an hour of labelled video a smaller fraction, an hour of simulation a fraction that depends on the gap, and a source that lacks a column the task needs replaces none of it, however many hours there are. Exchange rates put every source in one unit, but not in dollars. What does an hour of each source cost?
The plan
Six moves. (1) Put the programme's tally on the Bench and watch hours fail. (2) Rule out frames, held-out error, points of success and dollars as units. (3) Read a rate off your own curve, with an interval. (4) Measure it for every source. (5) Find the column that no number of hours buys. (6) Say what a rate cannot say.

1 · The same hours, different data

The programme that lesson 15 left must spend scarce robot time on data of several kinds. Before asking how much of each, we need a way to count them; the cheapest place to test one is the layouts task of lessons 7 and 8: a two-link arm carries a cup past five posts to a mat, a layout is a shift and a tilt of the row, and the policy copies the stored demonstration of the nearest layout. Success is the share of 1,000 fresh layouts, one run each under gusts, on which the cup reaches the mat untouched. One demonstration is one layout, 9.3 s of motion in 186 frames of 0.05 s, so an hour holds 387.1 demonstrations; resets, discards and operator time are lesson 17's.

Your arm holds n layouts, and a source adds m attempted layouts, of which the laboratory keeps those that succeeded. The sources: an arm of the same make (the twin); an older arm from another laboratory, slower and less precise (lesson 8); footage of a person, whose hand wobbles more than the arm (lesson 9) and is tracked to 3 mm; simulators whose arm model is off by a fraction of a link and that label frames in joint angles (lessons 8 and 11). The tally: 8 own layouts plus 256 attempted layouts of each, which at 9.3 s a layout is 39.7 minutes.

8 own layouts plus 256 attempted layouts ofkeptframes storedsuccess on 1,000 fresh layouts
nothing (your eight alone)0021.6 %
an arm of the same makeall 25647,47998.8 %
an older arm from another laboratory19744,81671.7 %
footage of a person18039,49556.0 %
a simulator, arm model off by 1 %all 25647,47582.3 %
a simulator, arm model off by 3 %all 25647,4605.8 %

With the same arm, policy, test layouts and hours, success runs from 5.8 % to 98.8 %. The worst row is below the 21.6 % that the 8 layouts reached alone: those hours did harm.

Scale: what carries over
The Bench is small: its whole own curve, one layout to 256, is 0.66 hours of motion. Rates, ratios between rates, curve shapes and thresholds carry to a real programme; totals do not, since a real scene needs far more than two numbers to describe and each further number multiplies the layouts needed by about 5.8 at the Bench's factor (lesson 7).

2 · What a unit has to be

A unit must be (i) the same for every source; (ii) measured on this policy and task, since it states what hours do for this learner; (iii) spendable, so that the programme can convert the hours it buys into hours it would otherwise record; (iv) valid where the programme stands, because the next hour adds less as the policy improves. Four candidates fail:

unitwhat goes wrong
hours of motionThe tally of §1: the same hours, 5.8 % to 98.8 %.
frames storedFrames are hours again. At 64 attempted layouts the simulators off by 0.1 % and 3 % store 11,873 and 11,867 frames and reach 80.3 % and 12.5 %.
error on held-out framesLesson 1: the copy error stayed between 2.9 and 4.4 % while success fell from 87.5 to 42.0 %; lesson 9: the label with the smallest error did not decide.
dollars per hourNot yet. A price says what an hour costs to make, not what it does: two simulators run at one speed whatever the error of their arm model. Lesson 17 prices hours; the unit must first say what they are worth.

One thing survives all four: your own demonstrations. A programme can always measure their success curve on its own robot, so the unit is measured on this policy and task; they are what the money would buy otherwise, so it is spendable; and the curve is a ruler on which the success of any source can be located, wherever the programme stands. So ask of any pooled set: how many own layouts would have given this success?

Road not taken · points of success per hour
Measure a source by what it adds, success with it minus success without it. That needs no ruler and fails requirement (iv): the same 64 twin layouts add 60.4 points to 8 own layouts and 31.0 to 32. The source did not change; the point where you stand did. It is also capped, since a policy at 98 % can gain nothing. A count of own layouts has neither fault: the twin's rate stays near 1 at both points (§4). Points per hour return in lesson 18, as a slope at a stated point.

3 · Reading a rate off your own curve

Your own curve is the success of a policy built from N own layouts, on the same 1,000 test layouts as everything below. It reads 22.9 % at one layout, 23.1 % at 2, 8.9 % at 4, 21.6 % at 8, 54.9 % at 32, 69.8 % at 48 and 99.0 % at 256. It saturates, gaining 24.1 points from 32 to 64 layouts, 14.3 from 64 to 128 and 5.7 from 128 to 256, and it is not smooth at tiny N: four layouts do worse than two, since a few can leave a gap wherever they fall, so the inversion reads its running maximum, which can only rise.

Pool n own layouts with m attempted layouts of a source and measure the success S. Let Neq be the own-layout count whose own success is S, found by straight lines between the measured points and, below the first, by a line from the origin. The pooled set did what Neq own layouts do; n of them were already yours, so the source supplied Neq − n, and per attempted layout, that is per source hour in the unit of §1, the exchange rate is

ρ = (Neq − n) / m

Take the older arm with n = 8 and m = 64, which is 9.9 minutes of its motion. The pooled policy succeeds on 55.1 % of the test layouts; your own curve passes 54.9 % at 32 layouts and 69.8 % at 48, so it reaches 55.1 % at 32.2. The pooled set does what 32.2 own layouts do, the source supplied 24.2 of them, and ρ = 24.2 / 64 = 0.38: 9.9 minutes of the older arm did the work of 3.8 minutes of yours.

At ρ = 1 the source is as good as your own, at 0 it is worthless (S is what your n layouts reach on the ruler), below 0 it did harm. There is a floor, because Neq cannot be below 0: ρ ≥ −n / m, here −0.125. Near n = 8 the ruler is flat at 23.1 % from 2 layouts to 8, so even the 21.6 % of your eight alone reads Neq = 0.9 and ρ = −0.11: below 23.1 % the rate says harm and the success says how much.

The test layouts give S a Wilson interval, 52.0 to 58.2 %. The running maximum is monotone, so the interval passes through the inversion: Neq from 29.1 to 35.5, ρ from 0.33 to 0.43. That is the sampling error of 1,000 test layouts and nothing more: it leaves out the noise of the own curve, read at one point of one draw, and the draw of the layouts, which §4 measures.

The widget

The exchange desk: what is an hour of this source worth in your own hours?
Left: success (up) against layouts, own plus attempted (right, log scale): grey your own curve, dots as measured and line the running maximum the inversion reads; teal your n (black dot) plus the first m layouts of the source; amber the chosen m with its interval, a dashed line across to your own curve and down to the axis, where it reads Neq. Right: ρ (up) against m with 8 own layouts (solid) and 32 (dashed), bars from the test layouts, dotted line the floor. A simulator's percentage is the error of its arm model. The task selector swaps in the peg of lesson 10: your curve is then the clone with force, the source the same demonstrations without.
success, own + source
—
95 % interval
—
own layouts alone
—
own layouts with the same success
—
exchange rate ρ
—
ρ interval (test layouts)
—
source layouts kept
—
frames from the source
—
joint-label gap (margin 1.3 cm)
—
Show the core JS
EX.cell = function (key, na, m) {
  ...
  var cur = BL.ownCurve(S.space), r = BL.setting({ na: na, body: S.body, m: m, w: 1, space: S.space });
  var at = function (s) { return m > 0 ? (BL.equivalentOwn(cur, s) - na) / m : NaN; };
  ...
    neq: BL.equivalentOwn(cur, r.succ), rho: at(r.succ), rlo: at(r.ci[0]), rhi: at(r.ci[1]) });
};
...
BL.equivalentOwn = function (curve, sPooled) {
  var N = [0].concat(curve.N), s = [0], i;
  for (i = 0; i < curve.s.length; i++) s.push(Math.max(s[i], curve.s[i]));
  if (sPooled <= 0) return 0;
  for (i = 1; i < N.length; i++) if (sPooled <= s[i]) return N[i - 1] + (N[i] - N[i - 1]) * (sPooled - s[i - 1]) / Math.max(1e-12, s[i] - s[i - 1]);
  var a = N.length - 1, sl = (N[a] - N[a - 1]) / Math.max(1e-12, s[a] - s[a - 1]);
  return N[a] + sl * (sPooled - s[a]);
};

What to try. Leave the defaults, the older arm with 8 own layouts and 64 attempted: success 55.1 %, your own curve reaches it at 32.2 layouts, ρ = 0.38. Slide the hours to 16 and to 256: ρ = 0.76 and 0.17, a fall the right panel shows. Choose 32 own layouts: 0.17 at 64. Choose the arm of the same make: 1.01 at 64, and the teal curve lies on the grey one. Step through the simulators from 0.1 to 3 %: ρ = 0.93, 0.83, 0.42 and −0.12, with joint-label gaps of 0.08, 0.23, 0.76 and 2.27 cm; at 3 % the dashed line runs to the left edge, because Neq is 0.5 layouts, below the 8 you hold; with hand labels the same simulator gives 1.02. Last, set the task to the peg: with force one demonstration gives 82.5 % and ten give 98.5 %; without it forty give 63.0 %, which the with-force curve passes at 0.76 of a demonstration, so ρ = 0.019.

4 · What the rates are

The table gives each source's rate at 64 attempted layouts, on the draw of layouts the widget uses and averaged over twelve other draws of both sets of layouts, own and source, on the same test layouts.

source8 own layouts32 own layouts
this drawtwelve drawstwelve draws
an arm of the same make1.011.051.10
an older arm0.380.290.16
footage of a person0.160.140.01
simulator off by 0.1 %0.930.971.03
simulator off by 1 %0.420.390.28
simulator off by 3 %−0.12−0.08−0.29
the same, hand labels1.021.051.09

Read the instrument first. The twin is a control: its demonstrations come from the same distribution as yours, so a sound instrument reads about 1. It reads 1.01 on this draw, with an interval of 0.90 to 1.12 from the test layouts, and 0.81 to 1.29 over twelve draws, 1.05 ± 0.24. That spread comes from which layouts were drawn and is wider than the interval from the test layouts: rates closer together are not ordered by one draw. These orderings hold on every draw, at 8 own layouts and at 32: twin above older arm above footage, and simulators falling as the gap grows.

The older arm's rate falls as you take more of it: 0.76 at 16 attempted layouts, 0.38 at 64, 0.17 at 256. Its curve flattens below yours, ending at 71.7 % where yours reaches 99.0 %, and at 64 attempted layouts 22 % of its attempts are discarded. Footage is worth 0.16: a person's hand wobbles more than the arm, the tracker adds 3 mm, and 36 % of the 64 attempts are discarded, so the rate is 0.26 per kept layout and 0.16 per attempted layout, which the laboratory pays for.

A simulator never discards, and its rate follows its gap. A joint-angle label carries the simulator's link lengths (lesson 8): one link longer by δ, the other shorter. When your arm tracks the labels its hand lands off the simulator's by 2|δ| sin(q2 / 2), q2 being the elbow angle, since the two link directions differ by q2: 0.08, 0.23, 0.76 and 2.27 cm for the simulators off by 0.1, 0.3, 1 and 3 %, against the 1.32 cm by which the demonstrated paths clear the posts. At 0.57 of that clearance the rate is 0.42; at 1.72 times it the 64 simulated layouts leave the policy at 12.5 % against the 21.6 % of its 8 alone, where 0.5 of one own layout would, and ρ = −0.12 is next to the floor of −8 / 64 and to the −0.11 of no change (§3). With hand positions as labels the same frames are worth 1.02: the gap enters through the label.

The operating point. On this draw every rate is lower at 32 own layouts, but over twelve draws that is three effects. A source as good as your own is worth the same however much you hold (the twin, 1.05 at 8 own and 1.10 at 32). A mediocre source is worth less the more you hold (the older arm 0.29 and 0.16). A harmful source harms more (the 3 % simulator −0.08 and −0.29), because there is more to destroy and the floor moves from −0.125 to −0.5. A rate belongs to a source, a policy, a task, the n you hold and the m you take.

5 · A column that no number of hours buys

Everything so far was a source with every column the task needs, differing in quality. Lesson 10 met the other case: a peg in a hole with a millimetre of clearance, where holding position inserts 63.5 % of attempts and only a policy that feels the force and gives way does better. A clone learns that only from demonstrations that recorded force; drop the column and the clone is the position policy.

demonstrations msuccess with forcesuccess without forcewith-force demonstrations with the same success
182.5 %51.5 %0.62
582.5 %66.5 %0.81
1098.5 %64.0 %0.78
4098.5 %63.0 %0.76

Both clones come from the same m demonstrations of the yielding policy, on 200 attempts. The with-force curve is the ruler (a line from the origin below its first point); the clone without force is used alone, since it cannot read force from frames that carry none and nothing can be pooled. Forty demonstrations without force reach 63.0 %, which the ruler passes at 0.76 of a with-force demonstration, so ρ = 0.76 / 40 = 0.019. At no m on the grid (1 to 40) do they reach the 82.5 % that one demonstration with force gives (the best is 66.5 %), so together they replace at most 0.81 of a demonstration and ρ(m) ≤ 0.81 / m goes to zero.

That is a structural zero, and it differs in kind from a low rate. Footage is worth little because its hours are worse than yours, and its curve still climbs: 38.0, 44.8 and 56.0 % at 64, 128 and 256 attempted layouts with 8 own. The clone without force is flat at 63.0 to 66.5 % from two demonstrations to forty, its ceiling the position policy's, which no frame without force corrects; ten demonstrations with force give 98.5 %. A column the task needs cannot be bought from a source that lacks it.

6 · A rate is not a price

A rate is a ratio of two quantities of one kind, own motion hours per source motion hour: it has no currency and no clock. What it says about money is a ceiling. An hour of the source replaces ρ own hours, so it is worth paying at most ρ · pown, where pown is what an own hour costs; pay more and you do better to record your own.

sourcemost an hour is worth paying (share of an own hour's price)what an hour costs
an arm of the same make1.01?
an older arm0.38?
footage of a person0.16?
simulator off by 3 %−0.12no price is worth paying

The right-hand column is what the table lacks, and without it ranking by rate cannot choose. The twin ranks first, but an hour of it is worth 2.7 hours of the older arm, so it is the better buy only if it costs less than 2.7 times as much; footage is worth buying only below a sixth of an own hour's price; the 3 % simulator's ceiling is negative, as if you had to be paid 0.12 of an own hour's price for each hour you took.

What this lesson did not do
It priced nothing: pown and every asking price are lesson 17's. It measured an average over a block of m layouts at one n, not the value of the next hour (lesson 18 takes the slope), and no source that expires (lesson 19) or fleet (lesson 20). It used a lookup policy and plain pooling at weight 1; a network trained on a mixture can weight sources, read pictures and share across bodies, so its rates can differ. One draw of layouts gives the numbers, twelve more the error bar. The contact block pools nothing. The task is small: see §1.

Common mistakes / failure modes

"an hour is an hour: count the hours"
The same hours from five sources left one policy anywhere from 5.8 to 98.8 %, and equal frames hide rates from 0.93 to −0.12 (§1, §2).
"the rate is a property of the source"
The older arm is worth 0.29 with 8 own layouts and 0.16 with 32, and 0.76 at 16 attempted layouts against 0.17 at 256 (§4).
"enough hours of the cheap source close the gap"
Without force forty demonstrations reach 63.0 %; one with force reaches 82.5 % (§5).
"the exchange rate is the price"
It is the ceiling ρ · pown; which source to buy depends on two prices the rate does not contain (§6).

Checkpoint exercise

Try it
Your own curve passes through 54.9 % at 32 layouts, 69.8 % at 48 and 79.0 % at 64, and its first point is 22.9 % at one layout. (a) Eight own layouts plus 64 attempted layouts of a source succeed on 62 % of fresh layouts. What is the rate? (b) The same 64 layouts added to 32 own layouts give 66 %. What is the rate now? (c) Another source gives 12 % with 8 own plus 64 attempted. What is its rate, and what is the lowest rate possible? Answer: (a) 62 % lies between 54.9 and 69.8 %, so Neq = 32 + 16 × (62 − 54.9) / (69.8 − 54.9) = 39.6 and ρ = (39.6 − 8) / 64 = 0.49. (b) Neq = 32 + 16 × (66 − 54.9) / 14.9 = 43.9 and ρ = (43.9 − 32) / 64 = 0.19: the same hours did less because you already held more. (c) 12 % is below the first point, so Neq lies on the line from the origin: 12 / 22.9 = 0.52 layouts and ρ = (0.52 − 8) / 64 = −0.12. The lowest possible is Neq = 0, ρ = −8 / 64 = −0.125: this source sits at the bottom of the scale, with the pooled policy worse than one own layout (22.9 %).

Where this points next

Measured on the Bench, an hour of an older arm's demonstrations replaces 0.38 of an hour of your own, an hour of footage 0.16, an hour of simulation 0.93 down to −0.12 as its arm model's error grows from 0.1 to 3 %, and an hour of demonstrations without the force column 0.019 at forty and less with more. The unit says nothing about cost: an hour of the twin is worth 2.7 hours of the older arm, and whether that makes it the better buy depends on two prices the rate does not contain; ρ · pown is only the most an hour is worth paying. Exchange rates put every source in one unit, but not in dollars. What does an hour of each source cost?

Takeaway
An hour of data is not a unit: the same 0.66 hours from five sources left one policy anywhere from 5.8 to 98.8 %. What survives is your own demonstrations. Pool n own layouts with m attempted layouts of a source, read the success off your own curve, and the rate is ρ = (Neq − n) / m, with an interval from the test layouts and, from the draw of layouts, an error bar that reaches ±0.24 near 1. A source of the same make is worth about 1, an older arm about a third, footage a sixth, a simulator from nearly 1 to below 0 as its arm model's error grows; mediocre sources fall with m and with n. A source without a column the task needs replaces at most 0.81 of a demonstration in all, so its rate falls like 0.81 / m and no number of hours closes the gap. A rate is a ceiling on a price, not a price.

Interview prompts

Companion reads: Lesson 8 · Other bodies (the first exchange rate) and Lesson 10 · What cameras cannot see (the peg and force).