Sums and comparison

A floor, and not a decline

Comparability fell eighteen points from day two to day three and the next day cannot be enumerated. It can be built — and the construction's bias measured one day lower, where the truth is known. Corrected, day four comes to 60.6 per cent against day three's 59.7: the fall was a one-day event.

Assumes: How rare it is to be bigger · Confused is not the same as unknown

How rare it is to be bigger counted the order. Among the 22 values born by day two, 179 of 231 pairs can be compared; one day later the share falls to 60 per cent, and the share of values comparable with nought falls from 13 in 22 to 106 in 369. It closed on the obvious next question and on why it could not be answered:

The rung above is the trend: whether comparability has a limit above zero, which needs a day this evaluator cannot enumerate.

Day four cannot be enumerated. It can be built, and a built sample is worth exactly what its calibration is worth.

The fall stops. Comparability on days two, three and four, the last built and corrected.
Fig. 1 The sequence with day four built and corrected by a bias measured one day lower. 77.5, 59.7, 60.6 — a fall of nearly eighteen points and then a rise of eight tenths of one. The figure refuses to draw unless the two enumerated days reproduce the rung below’s measurement, since a page extending a curve must first sit on it.

The fall stops. Comparability has a floor near three pairs in five, not a decline toward nothing.

Why the day cannot be listed

The construction runs by days: every value born by day nn has all its options born by day n1n-1, and a value born by day n+1n+1 is any pair of subsets of them. Day one has four values, day two twenty-two, day three 1,474 — and day four is the number of pairs of subsets of 1,474 things, which is not a number anybody enumerates.

So the rung below’s method — list a day, take every fourth value, compare every pair — cannot reach day four, and the honest response is either to stop or to sample. Stopping is what the rung below did, and it was right to: an uncalibrated sample would have been worse than no third point at all.

The fall that raised the question. Comparability on day two and day three, as it was first measured.
Fig. 2 The two days the rung below could enumerate. Twenty-two values and 231 pairs at day two, 369 sampled values and 67,896 pairs at day three. A fall of 17.8 points in a single day, which is what made the trend worth asking about and what makes a third point worth having.

Building a day, and distrusting it

A day-four value can be made rather than found: choose one or two day-three values for each side, form the game, and keep it when its canonical form is genuinely four days deep.

That produces day-four values and it does not produce a fair sample of them. It favours narrow forms — one or two options a side, where a real day-four value may have dozens — and narrowness is exactly the sort of thing that could move comparability.

A sample that has to be calibrated. The three pools: an evenly sampled day three, a built day three, and a built day four.
Fig. 3 The three pools. An evenly sampled day three, taken every fourth value from the enumeration as the rung below did; a day three built the same way day four has to be; and day four itself, built. The middle row exists only to be compared with the first.

So the sample is not trusted. It is calibrated, by running the identical construction one day lower — where the honest answer is known — and reading off how far it misses.

The bias, measured. The construction's effect on both measures at day three, where the honest answer is known.
Fig. 4 The construction run at day three, where the truth is 59.7 per cent. It gives 65.5, so it adds 5.7 points to pairwise comparability. On the other measure it adds 36 points, against a true value of 29 — a bias larger than the quantity it is a bias in.

On pairwise comparability the construction adds 5.7 points. That is a real bias and it is small enough to subtract, and subtracting it is the whole of the correction: day four’s built sample gives 66.3 per cent, which corrects to 60.6.

The direction of that bias is worth noting because it is the direction that matters. The construction overstates comparability, so an uncorrected reading of the built day-four sample — 66.3 per cent — would have said comparability was rising, which is a more dramatic and entirely spurious finding. The correction turns a rise of 6.6 points into one of 0.8, and the difference between those two sentences is the whole value of running the calibration.

Against day three’s 59.7. The fall was a one-day event.

Why the correction is a subtraction and not a ratio is worth one line, because the choice is doing work. The bias could have been applied multiplicatively — the construction gives 1.095 times the honest figure at day three, so day four’s 66.3 would correct to 60.5 — and it lands in the same place, 60.5 against 60.6. The two corrections agree to a tenth of a point, which is the one piece of evidence here that the correction is not an artefact of how it was applied.

That agreement is a small thing and it is the sort of check that costs nothing and occasionally saves a result. Where a bias is small relative to the quantity the two corrections coincide; where it is large they diverge wildly, which is exactly what happens on the other measure below.

The half that cannot be rescued

The rung below measured two things and this method reaches one of them.

The half this cannot reach. Comparability with zero, and why the sampling method cannot be corrected for it.
Fig. 5 The other measure. Comparability with nought falls from 59 per cent to 29 across the two enumerated days, and the construction adds 36 points to it at day three — larger than the quantity. Subtracting a correction bigger than the number being corrected is not a correction, so no day-four figure is quoted for it.

The share of values comparable with nought falls faster — 59 per cent to 29 — and the construction inflates it by 36 points at day three, against a true value of 29. A correction larger than the quantity is not a correction, and there is no honest day-four figure for it here.

The reason is legible. A value is comparable with nought when one player wins it whoever moves, and narrow forms are far more likely to be decided that way — a value with one option a side is very often a straightforward win. So the construction’s preference for narrow forms bites hardest on exactly the measure that asks about being decided, and hardly at all on the measure that asks about two values relative to each other.

That asymmetry is worth more than either number. It says the built sample is usable for questions about the relation between two values and unusable for questions about a value’s own outcome, and it says so on evidence rather than on a hunch.

What a floor at three fifths means

A floor, not a decline. The rung above's question with the answer the built and calibrated sample supports.
Fig. 6 The rung above’s question with the answer the sample supports. The fall is a one-day event; comparability appears to level near three pairs in five; and the figure a reader should carry is the method rather than the number, since a built sample is worth what its calibration is worth.

Three pairs in five comparable is not a small number, and it is worth saying what it does and does not mean.

It does not mean the order is nearly total. Two values in five cannot be put in order at all, which is confused is not the same as unknown’s subject and is a great deal of confusion — an order in which forty per cent of pairs are incomparable is not an order anybody would call one. And the two in five are not a fringe of exotic values: at day three they are drawn from the same 369 as the three in five, by the same construction, and nothing about a value announces which side of the line it falls on until the difference game has been played out.

It does mean the confusion is not swallowing everything. A reader looking at 77 per cent falling to 60 would reasonably guess the share was heading for nothing, and that the values born deep into the construction are almost all mutually incomparable. On this evidence they are not: the share levels almost immediately, and whatever governs it is set by day three rather than accumulating.

And it makes the day-two figure the outlier rather than the start of a trend. Day two is 77.5 per cent because day two is tiny — 22 values, of which seven are numbers, and numbers are totally ordered among themselves.

The obvious reading is that the fall is the numbers being diluted: 32 per cent of day two are numbers and 1 per cent of the day-three sample are, so a population going from a third totally-ordered to almost none should lose comparability without anything else changing. That reading is wrong, and it is worth checking rather than assuming. Comparability among the non-numbers alone falls from 76.2 per cent at day two to 59.8 at day three — nearly the whole of the drop, inside the population the dilution argument says should be unaffected.

So the numbers account for about a point and a half of the seventeen-point fall, and the rest is the non-numbers becoming harder to compare with each other. That is a fact about what a day of the construction does to the values it produces, and it is what makes the levelling at day four surprising: whatever was making non-numbers incomparable stops making them more so.

What a built sample is worth

The method here is the transferable part of the page and it is worth stating on its own, because this site will meet the same obstacle again.

A population that cannot be enumerated can often be constructed, and a construction always biases. The temptation is to construct, measure, and report — and the report is then a measurement of the construction rather than of the population, with no way to tell from inside it which.

The way out is that the construction can usually be run somewhere it can be checked. Every construction that reaches day four also reaches day three, and day three is listable. So the bias is not a matter of judgement: run the same code one step lower, compare with the truth, and read the number off.

And the calibration decides which questions the sample can answer, which is more useful than the correction itself. Here it says the sample is good for the relation between two values — a 5.7-point bias on a 60-point quantity — and useless for a value’s own outcome, where the bias is larger than the quantity. Without the calibration both numbers would have been reported with the same confidence, and one of them would have been wrong by more than its own size.

That is a general procedure and this anchor is not the only place on this site that needs it. Every ladder here has a rung that ends which needs a day this evaluator cannot enumerate, and most of them stop there.

What the levelling would have to mean

A quantity that falls sharply once and then stops is a different object from one that decays, and it is worth saying what each would have implied.

A decay toward nought would say that comparability is a property of simple values — that two values picked at random from deep in the construction are almost certainly incomparable, and that the partial order, though a genuine order, is essentially empty at scale. That is the reading the first two points invite, and it would make the simplest game above both’s finding — that day-two values have least upper bounds inside day two — an artefact of a small population rather than a structural fact.

A floor says the opposite: that a fixed and substantial share of pairs are ordered however deep one goes, and that whatever makes two values comparable is available at every scale. Three pairs in five is enough that an argument about the order is an argument about most of the population.

The sample supports the second and cannot say why. What would say why is the kind of pair that stays comparable — whether the comparable pairs at day four are pairs sharing a component, or pairs of very different temperatures, or something structural — and that is a classification this page did not attempt on a sample of 120 values. It is the natural companion to the robustness check named below and a good deal more interesting.

One thing the levelling is not is evidence that the order becomes total. Two values in five incomparable is not a small residue, and it is the same residue at day four as at day three. Comparing positions is where the four relations were set out, and the fourth one is not going away.

What the solver computed, and how

Three populations. Day two is every one of its 22 values. Day three is every fourth value of the enumeration, 369 of them, which is the rung below’s sample and reproduces its figure exactly — 59.7 per cent against the 60 it reported.

The two built pools are made the same way: choose one or two values from the previous day for each side, form the game, reduce to canonical form, and keep it when its depth is exactly the day wanted and it has not been seen before. Both are 120 values, from one seed, so the day-three and day-four built pools differ in nothing but the pool they draw from.

Every pair in every population is compared by the ordinary test — is one at least the other, decided by playing the difference game — and a pair is counted comparable when either direction holds. Comparability with nought is the same test against the empty game.

The bias is then the built day-three figure less the honest one, and the corrected day-four figure is the built one less that bias. Nothing is fitted; the correction is one subtraction of one measured quantity.

Three things are asserted rather than reported. The two enumerated days must reproduce the rung below’s measurement, or the two pages disagree about one population. The construction’s bias on pairwise comparability must be small enough to subtract. And its bias on comparability with nought must exceed the day-three value itself, since the page declines to quote a corrected figure there and the refusal has to be earned.

Where the model stops

One construction, one seed, 120 values. The day-four figure rests on a single sample from a single method, and the correction rests on that method behaving the same way at day three as at day four — which is an assumption, not a measurement. It is the assumption every calibration makes and it is worth naming: if the construction’s bias grows with the day, the correction is too small and the true day-four figure is lower.

And the construction is one of many. Choosing one or two options a side is the cheapest way to reach day four and not the only one; a method allowing up to four options a side would produce wider forms, a different bias, and a different corrected figure. Whether the corrected answers agree across methods is the obvious robustness check and it has not been run.

Two points do not make a floor. 59.7 and 60.6 is a level, and a level over one step is consistent with a slow decline, with a genuine floor, and with a shallow rise. What the sample rules out is the reading a reader would take from the first two points — a steep continuing fall — and it does not distinguish among the others. A fifth day would separate them and is built the same way, from the day-four sample rather than from an enumeration, so its calibration would have to be a calibration of a calibration; nothing here says how far that can be pushed.

Normal play throughout, and comparison is a normal-play notion: GHG \geq H means Left wins GHG - H moving second, which is comparison is a search’s subject and is decided by the recursion this convention makes well-founded.

And the figures cannot show the sample. Six tables of shares describe three populations of values, and the object — 120 built day-four forms, most of them a dozen characters wide — is a list rather than a picture. What could be drawn is one of them beside a day-three value it cannot be compared with, which is confused is not the same as unknown’s figure at one day deeper.

Where the ladder goes next

The comparison anchor has five rungs: comparing positions, comparison as a search, the fourth answer as a fact about the pair, how often each answer comes up, and now whether the last of those has a limit.

The rung above is the second construction. This page’s day-four figure rests on one way of building a day-four value, and the honest test of a calibrated sample is whether a differently-biased construction, calibrated the same way, lands in the same place. Allowing up to four options a side gives wider forms and a bias that can be measured at day three exactly as this one was; if the two corrected day-four figures agree the floor is real, and if they disagree the calibration is method-specific and the whole approach is worth less than this page claims. It is the same code with one parameter changed, and it is the first thing anybody should ask of a number arrived at this way.

Two neighbours are worth the trip. How rare it is to be bigger is where the two enumerated days were counted, and it is the page this one extends by one point and one method. And how old a value is is where the birthday is established as the measure the construction proceeds by, which is what makes day four a well-defined population that nobody can list.

Part 5 of 6

One argument about Comparison. The parts either side of it:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.

BirthdayCanonical formComparisonConfusionDominanceEnumerationEqualityNormal playOutcome classPartial order