Sums and comparison

The pairs stand still and the order spreads

The share of pairs of game values the order can compare falls from 77.5 per cent on day two to 60 on day three and then stops: built twenty times under two recipes and calibrated, day four sits within two standard errors of day three. That looked like a floor under the order. Measured beyond pairs it is not. Chains of three keep falling, triples with no comparable pair keep rising, and the largest set of mutually confused values in a sample of 120 grows from about 20 to about 30 — every one of forty seeds wider than day three's mean. The floor belongs to one statistic, not to the order it summarises.

Assumes: Twenty draws and a second recipe · A floor, and not a decline

Two game values need not be comparable. The order that says one position is at least as good for Left as another is a partial order: some pairs are ordered one way, some the other, and some are confused — their difference is a first-player win, so neither is at least the other. Comparing positions is where the order is defined, and confused is not the same as unknown is where the third answer is shown to be an answer.

How often a pair is confused depends on how old the values are. A floor, and not a decline counted it on the values born by each day. On day two, 77.5 per cent of pairs compare; on day three, 60.2. Day four has too many values to list, so a sample was built — day-three values put together as options, kept when the result is genuinely four days deep — and corrected by the bias the same recipe shows one day lower, where the truth is known. Twenty draws and a second recipe repeated that twenty times under two recipes and found the corrected day-four figure within a standard error of day three’s. The decline had stopped.

That essay named the question the number could not settle. A share of comparable pairs that stops falling could mean the order stops changing as values age — a property of the order — or that one statistic happens to level while the order goes on changing underneath it. The test is to measure things about the order that the pair share does not determine.

Three measures the pair share does not fix

Three are natural, and each asks about more than two values at once.

A triple of values is a chain when the order compares all three pairs: the three can be lined up from worst to best for Left. It is wholly confused when the order compares none of them. And the width of a set of values is the size of its largest wholly confused subset — the most values that can be chosen so that no two compare. The width is found exactly rather than searched for, by a theorem of Dilworth’s: in any finite partial order it equals the number of elements less the largest matching between each element and one strictly below it.

Days two and three, beyond pairs. The order on the values born by day two (22) and day three (1,474), measured by the share of comparable pairs, of triples that are chains, of triples with no comparable pair, and by the width. Day three: 60.2% of pairs, 24.6% chains, 11.6% wholly confused triples, width 86.
Fig. 1 The values born by day two and by day three, measured four ways: the share of pairs the order compares, of triples it compares completely, of triples it compares not at all, and the width. The last row is random samples of 120 day-three values, the figure a built sample of that size has to be calibrated against.

On the enumerated days the four measures all move together. From day two to day three the pair share falls from 77.5 per cent to 60.2; the share of chains among triples halves, from 48.7 to 24.6; the share of wholly confused triples multiplies eightfold, from 1.4 to 11.6; and the width goes from 4 of 22 values to 86 of 1,474. Every measure says the order is getting less orderly, and none of them is redundant with another, since the triple shares are not functions of the pair share and the width is not a function of either.

The last row is the one the rest of this essay leans on. The width of a set depends on how many values are in it, so day four cannot be compared with day three’s 86 except through samples of the same size. Twenty random samples of 120 day-three values give a width of 19.7 on average, a chain share of 23.7 per cent and a wholly confused share of 12.2 — and a pair share of 59.2, within the sampling error of the whole day’s 60.2.

The widest confused set on day two

Day two is small enough to see the width whole.

The widest confused set on day two. A largest set of day-two values in which no two are comparable: 0, ∗, 1 | −1, ∗2. Day two's width is four and day three's is 86.
Fig. 2 A largest set of day-two values no two of which compare: nought, star, the switch ±1 and star two. Day two’s width is four, so no five of its twenty-two values are pairwise confused; day three’s is eighty-six.

Nought, ∗, ±1 and ∗2. Each pair’s difference is a first-player win: ∗ and ∗2 differ by ∗3, nought differs from each of the others by that other, and ±1 differs from ∗ and ∗2 by a switch with a star in it that whoever moves first can win. They are four values any reader of these essays has met, and they are four things the order cannot rank. No fifth day-two value is confused with all four. The numbers it is confused with measures the same relation from the side of a single hot value — the interval of numbers a switch cannot be told apart from — and ±1’s place in this set is that interval containing nought.

The day-three set that realises the width is eighty-six values long, and it is found the same way: a matching between each value and one strictly below it, as large as possible, and König’s construction to read the wholly confused set off the matching. The count 86 is exact, not a lower bound from a search that stopped.

What each recipe gets wrong, one day lower

A built day-four sample is not a random sample, and the calibration that made the pair floor believable has to be repeated for each new measure.

Each recipe's bias, one day lower. Built day-three samples under two recipes against random samples of the enumerated day, on pairs, chains, wholly confused triples and width. Both recipes over-compare and under-spread: the narrow recipe's samples have width 13.0 against the honest 19.7.
Fig. 3 Built day-three samples of 120 under the two recipes, twenty seeds each, beside random samples of the enumerated day. Both recipes build samples that are more comparable and narrower than the day really is, the narrow recipe by more; the difference is the bias subtracted at day four.

The narrow recipe chooses one or two values a side; the wide recipe one to four. Both favour forms with few options, and forms with few options compare more often — a form with one option a side is often a number, and any two numbers compare — so both overstate the pair share, the narrow recipe by eight points, the wide by two. The same preference shows in the other measures, and more strongly in the width. A sample built from day two by the narrow recipe has width 13.0 where a random sample of the day has 19.7; the wide recipe’s has 14.3. Both recipes build orders a third narrower than the real one.

That bias is measured, not assumed, which is the whole of what makes a correction possible. Each day-four sample built by a recipe is corrected by the amount that recipe was wrong at day three, measure by measure, seed by seed. If a recipe’s bias at day four differed from its bias at day three, the two recipes would disagree once corrected. For the pair share the earlier essay found they agree. They agree here too, on every measure, to within their errors.

The pairs level and the order does not

The pairs level and the order does not. Corrected day-four figures under two recipes beside the honest day-three figures, for samples of 120. The share of comparable pairs is unchanged within its error; chains of three fall, wholly confused triples rise from 12.2% to about 14.8%, and the width rises from 19.7 to 29.4 and 30.3.
Fig. 4 Corrected day-four figures under both recipes beside the honest day-three figures, means and standard deviations over twenty seeds. The pair share is unchanged within its error; chains of three fall, wholly confused triples rise, and the width of a sample of 120 rises from about 20 to about 30.

The pair share stays where it was. Corrected, day four’s share of comparable pairs is 57.6 per cent under the narrow recipe and 58.7 under the wide, against day three’s 59.2 — within 1.7 and 0.6 standard errors. That is the floor of the earlier essay, reproduced by a different sampling of the day-three yardstick.

Nothing else stays. The share of chains among triples falls to 19.2 and 21.3 per cent, from 23.7. The share of wholly confused triples rises to 14.8 and 14.5, from 12.2. And the width moves most of all: a sample of 120 day-four values holds 29.4 or 30.3 mutually confused values, where a sample of 120 day-three values holds 19.7. Half as many again, in one day.

Which statistics stand still. The change from day three to corrected day four in each of four statistics of the order, under each recipe, in absolute terms and in standard errors. The share of comparable pairs moves by under two; the width moves by many.
Fig. 5 The change from day three to corrected day four in each measure, in absolute terms and in standard errors. The pair share moves by under two standard errors under both recipes; the confused triples move by four, and the width by thirteen and sixteen.

Put in standard errors, the difference between the measures is not subtle. The pair share moves by 1.7 and 0.6. Chains move by 3.6 and 2.5, confused triples by 4.4 and 3.7, and the width by 13.0 and 16.5. A floor is a statistic that moves by less than its error from one day to the next, and by that definition the pair share has one and the width does not have anything like one.

The triples need one more sentence of care, because two of the three measures are less independent of the pairs than they look. If every pair compared independently with the same chance pp, a triple would be a chain with chance p3p^3 and wholly confused with chance (1−p)3(1 - p)^3. Against that baseline the chains are unremarkable: they run at 1.14 times the independent figure on day three and at 1.01 and 1.05 times on day four. Most of the fall in chains is the small fall in the pair share, cubed. A chain needs three comparable pairs, so a one-point drop in pairs becomes a larger drop in chains without any change in how the pairs are arranged.

The wholly confused triples are different. They run at 1.8 times the independent figure on day three and at 1.9 and 2.1 times on day four. Confusion clusters: values confused with one value tend to be confused with each other, about twice as often as chance would have it, and slightly more so at day four than at day three. That is a change in arrangement, and it is the change the width measures directly.

What the width is counting

Dilworth’s theorem says something about the width from the other side, and it makes the rise easier to picture. The largest wholly confused set in a partial order has exactly as many elements as the fewest chains that cover it — every value lies on one of the chains, and no two values of the confused set can share a chain.

So a width of 19.7 in a sample of 120 day-three values means those values can be arranged into about twenty chains, each an ordered run from worst to best for Left, averaging six values long. A width of 30 in a sample of 120 day-four values means thirty chains averaging four. The order at day four is made of shorter chains, more of them, side by side. The share of pairs that compare can stay where it was while that happens, because two values on different chains may still compare — the chains are a covering, not a partition of comparability — and what shrinks is the length of the longest ordered runs, not the number of ordered pairs.

That is the picture the pair floor could not give. A reader told that three in five pairs compare on both days would imagine the same order, a little larger. The chain covers say it is a different order: at day three a value sits in a column of six it can be ranked against, at day four in a column of four, with half as many columns again beside it.

Every seed is wider

A mean can hide a spread, and the earlier essay’s lesson was that a single draw can say almost anything.

Every seed is wider. The range over 20 seeds of the corrected day-four pair share and width under each recipe. The pair shares fall on both sides of day three's value; every width lies above day three's mean and 39 of 40 above its widest draw.
Fig. 6 The least and greatest corrected day-four value over the twenty seeds of each recipe, and how many seeds lie above day three’s mean. The pair shares fall on both sides of it; every one of the forty widths is above it, and thirty-nine are above the widest of twenty honest day-three draws.

Seed by seed, the two statistics behave differently in exactly the way a floor and a rise would. The corrected pair share runs from 51.1 to 63.6 per cent across the forty seeds, and eighteen of the forty lie above day three’s mean — the scatter a noisy estimate of an unchanged quantity has. The corrected width runs from 24.7 to 34.7, and all forty lie above day three’s mean of 19.7. Thirty-nine lie above the widest of the twenty random day-three samples, which reached 25.

So the rise in width is not a matter of a few seeds or of one recipe. It is the same under both, it is there in every seed, and it is larger than the whole spread of the day-three yardstick.

What the samples cannot show

Day four is built, not listed. Every day-four figure here is a corrected sample of 120 values, and the correction assumes a recipe’s bias is the same at day four as at day three. The two recipes agreeing after correction is the evidence for that assumption, on every measure; it is not a proof of it.

Both recipes lean the same way. Each favours forms with few options, and the calibration removes the part of that preference which shows at day three. A recipe weighted the other way — towards forms with many options on each side — would have a bias of the opposite sign, and if its corrected width agreed with these two the rise would be established against the one objection the present pair of recipes cannot answer, that both share a blind spot. Two recipes agreeing is evidence; two recipes with opposite biases agreeing would be much stronger evidence.

The width is a sample width. A width of 30 among 120 values says what the order looks like in a sample of that size; the width of all of day four is far larger and is not estimated here. What is compared is like with like — 120 against 120 — and the comparison is the finding.

And triples are the only step beyond pairs. Chains and wholly confused sets of four, or the full distribution of chain lengths, would say more about the shape of the order, and they are not measured. The width is the one measure here that looks at arbitrarily large sets, and it is the one that moves most.

The convention the order is taken under

The values are short game values in canonical form, and a value is born by day n if it has a form whose options are all born by day n − 1 — the construction how old a value is sets out. Two values compare when one is at least the other for Left: G≥HG \ge H when Left, moving second in G−HG - H, wins. They are confused when neither holds. Days two and three are every value born by then, 22 and 1,474 of them. The recipes are the earlier essay’s: {L∣R}\{L \mid R\} with one or two, or one to four, day-three values a side, chosen at random with a seeded generator, canonicalised and kept when born on day four exactly. The honest day-three yardstick is twenty random samples of 120 of the day’s values.

The surprise: the floor was the pair statistic’s

The pair share looked like the most natural summary of how orderly a set of values is, and at day four it seemed to say something about the order itself: that games stop becoming less comparable as they get older. What the other measures show is that the order goes on spreading at day four exactly as it did at day three. Chains thin, wholly confused triples thicken, and the largest confused sets grow by half. The pair share stands still because of how it is built, not because the order does. Read against independence, the chains thin only as fast as the pairs require; the wholly confused triples thicken faster than that; and the largest confused sets grow by half, which no statistic of single pairs could have shown.

How a statistic can level while the thing it summarises changes is worth one sentence. The pair share is an average over pairs, and it can be held constant by a shift in how comparable pairs are arranged — fewer long chains and more small confused clusters, with the same number of comparable pairs between them. The width and the confused triples see the arrangement; the pair share cannot. Comparison is a search is where deciding a single comparison is priced, and the reason a statistic of single comparisons was the first thing anyone measured.

Still open: where the width goes

The width of a sample grows by half from day three to day four. The next question is whether it keeps growing at that rate, and it needs a day-five sample, which cannot be built the same way: a recipe choosing day-four values needs a pool of day-four values, and the only pool available is itself a built sample, so two biases would compound and the calibration would have no day four to be measured against. What can be measured without that is the rate at a fixed size — the width of samples of 30, 60, 120 and 240 at days three and four — which says whether the width grows with the day as a power of the sample size or as a fixed proportion of it. If the proportion itself rises from day to day, the order is becoming wider in a way no floor on pairs will ever show; if it settles, the pair share was a slower indicator of the same thing. The number tree is where the birthdays these days are counted in come from.

Part 7 of 7

One argument about Comparison. The parts either side of it:

The objects named here

The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.

BirthdayCalibrationCanonical formComparisonConfusedCounterexampleEnumerationExhaustive searchPartial orderSampling