Twenty draws and a second recipe
Assumes: A floor, and not a decline · How rare it is to be bigger
A floor, and not a decline answered a question that cannot be answered by counting. The share of pairs of values that the game order can compare falls from 77.5 per cent on day two to about 60 on day three, and day four has too many values to list. So a sample of day-four values was built — one or two day-three values a side, kept when the result is genuinely four days deep — its bias was measured by building day three the same way and comparing with the truth, and the day-four figure was corrected by that bias. The answer was 60.6 per cent, against day three’s 59.7. The fall had stopped.
That essay also named the test it had not run. The honest test of a calibrated sample is whether a differently-biased construction, calibrated the same way, lands in the same place. This essay runs it — and a second test that essay did not name, which turns out to matter more: building the same sample again with a different seed.
What a calibration corrects, and what it does not
A built sample is biased by its recipe. Choosing one or two options a side at random and keeping the result favours forms with few options, and forms with few options are more often comparable, so the built sample overstates comparability. The calibration measures how much: build day three the same way, measure its comparability, and subtract the true day-three figure. Whatever the recipe adds at day three, it is assumed to add at day four too, and the day-four figure is corrected by that amount.
Two things can go wrong with that, and they are different.
The assumption can fail: the recipe’s bias at day four could differ from its bias at day three, because day-four forms are built from a larger pool and the recipe’s preference for narrow forms bites differently. The test for that is a second recipe with a different bias. If the correction is sound, two recipes that overstate by different amounts should still land in the same place once corrected; if it is not, they will disagree, and by roughly the difference in their biases.
The numbers can be noisy: a sample of 120 values is a sample, and so is the day-three figure the calibration subtracts. A correction is a difference of two noisy numbers and inherits the noise of both. The test for that is to repeat the whole procedure with different random choices and look at the spread. The earlier essay did this once, reported one number, and read a gap of 0.9 points between it and day three as evidence of a floor.
The number the correction is made against
The calibration’s reference point is the one number nobody had re-examined.
The earlier figure of 59.7 came from every fourth day-three value — one quarter of the day, starting at the first. Taking every fourth value starting at the second, third or fourth instead gives 58.5, 62.3 and 59.8. So the honest day-three figure that the whole calibration rests on was itself uncertain by nearly four points depending on which quarter was used, and the particular quarter chosen sat below the middle of that range.
Every pair of the whole day can be compared — 1,085,601 pairs, which takes seconds rather than the hours day four would — and the answer is 60.2 per cent. That is the number used as the calibration target from here on. It moves the reference half a point from where it was, and it removes one source of noise entirely, since the whole day has no sampling error.
Two recipes with different biases
The second recipe allows one to four options a side rather than one or two. Wider forms are built more often, and wider forms are less often comparable, so its bias should be smaller.
The two biases are well separated. The narrow recipe — the original — overstates day three by 7.2 points on average across twenty seeds; the wide recipe by 1.0. Individual seeds scatter by one and a half to two points around those means, but the two clouds barely overlap. That is exactly what the test needs: two instruments that are wrong by different amounts, so that agreement after correction means something.
The wide recipe is not simply the better instrument. Its bias is smaller, so it needs a smaller correction, but its day-four sample is drawn with a different mix of widths and there is no guarantee that a small bias at day three stays small at day four. What makes the comparison informative is not that either recipe is unbiased but that their biases differ.
Corrected, seed by seed
Now the day-four figure, built and corrected twenty times with each recipe.
The first thing the rows show is how wide the spread is. The narrow recipe’s corrected day-four figure runs from about 52.0 to 64.5 across twenty seeds; the wide recipe’s from about 52.0 to 64.1. A single seed can land anywhere in a band twelve points wide. The earlier essay’s 60.6 was one draw from that band, and the gap of 0.9 points between it and day three was a gap between one draw and one quarter of a day.
The spread is not surprising once it is seen. A sample of 120 values has about seven thousand pairs, and the share of them that are comparable is estimated with a standard error of a few tenths of a point — but the correction subtracts a second sample’s figure, and the two samples’ errors are partly shared and partly independent. Worse, the values in a built sample are not independent: a recipe that happens to draw several variants of one hot switch will fill a sample with near-duplicates whose pairwise comparisons all go the same way. So the effective sample is smaller than 120, and the noise larger than a naive calculation suggests.
How many draws a number needs
The spread says what a single draw was worth, and it also says what twenty are worth. Each recipe’s corrected figure varies from seed to seed with a standard deviation of about three and a half points. The mean of twenty such figures has a standard error of about , a little under 0.8 of a point. Reporting a day-four figure to the nearest tenth of a point would need a standard error of a few hundredths, which at this spread means thousands of seeds; reporting it to the nearest half point needs about fifty.
So the question the earlier essay answered — has comparability stopped falling? — was answerable with the samples it had, and the answer it gave was right. What it could not support was the decimal. A figure of 60.6 against 59.7 reads as “slightly above”, and on one draw each it meant nothing more precise than “within a few points”. The fault is not in the method, which is sound and is now shown to be sound by a second recipe; it is in reporting the output of a noisy method as if it were exact, which is a mistake the method itself gives no warning of.
Every comparison in the sample is exact. Each pair of values is compared by playing out their difference — comparison is a search, and here the search is run to the end — so there is no noise at all in whether one particular pair is comparable. The noise is entirely in which pairs were drawn. That is worth keeping in mind whenever an exactly computed quantity is summarised over a sample: exact parts do not make an exact whole.
The two recipes agree
The spread makes single draws useless. The means of twenty draws are not.
Corrected, the narrow recipe puts day four at 58.5 per cent and the wide recipe at 59.6, each the mean of twenty seeds. The difference, 1.1 points, is about one standard error of the difference between two means of twenty noisy numbers. So the recipes agree: two instruments with biases six points apart land in the same place once their biases are removed, which is the test the earlier essay asked for, and it passes.
What they agree on is a figure a point or so below day three’s 60.2, not above it as the single draw suggested. That gap is smaller than the uncertainty in either mean, and a gap that small is not a measurement of anything. The honest summary is: day four’s comparability is the same as day three’s to within the precision these samples allow, which is about a point and a half.
That is still the finding the earlier essay reported, and with better support. The fall from day two to day three was eighteen points. The change from day three to day four, measured two ways with twenty seeds each, is indistinguishable from nothing. A decline continuing at even a third of the first day’s rate would have shown up as six points, far outside the error bars. So the fall was a one-day event, the floor holds, and the decimal attached to it was noise.
The sequence, redrawn
With the error bars in place the three points of the curve can be set side by side again.
Day two is enumerated and exact at 77.5 per cent. Day three is now enumerated too, at 60.2, rather than sampled at 59.7. Day four is built and corrected, at 58.5 by one recipe and 59.6 by the other, each within about three quarters of a point of its own mean. Drawn with those error bars, the curve drops eighteen points in one day and then runs flat within its uncertainty. The drawing above keeps the earlier numbers so that the correction can be read against them; the dots at the head of this essay are what replaces its third point.
A reader who wants a single number for day four should take something like 59 per cent, give or take a point and a half — the average of two recipes that agree, with an allowance for the one error they might share. A reader who wants to know whether day four is below day three should conclude that nothing measured here can tell. Both statements are more modest than 60.6, and both are true.
What a floor at three fifths would mean
It is worth pausing on why the answer is interesting rather than merely settled. On day two, three pairs in four are comparable; on day three, three in five; and on day four, three in five again. How rare it is to be bigger had found the first two numbers and taken them to mean that comparability erodes as values get older — that deep enough in the construction, almost every pair of values is confused rather than ordered.
A floor says otherwise. It says the proportion of comparable pairs settles, and that the ordering of values keeps a fixed share of its structure however far the construction goes. Nothing here proves that — two built days after an enumerated one are three points on a curve — but the specific worry that the order dissolves has now been tested twice and has failed to appear both times. That the partial order stays three-fifths ordered is a claim about the whole universe of short games, and it is a surprising one to find resting on a comparison between two sampling recipes.
The error the two recipes could share
Agreement between two instruments rules out the errors in which they differ, so it is worth being specific about what they have in common.
Both recipes build a day-four value by choosing day-three values uniformly at random from the list of 1,474 and putting them on the two sides of a new form. Neither ever chooses a day-three value more often because it is simpler or hotter or more common as an option of other values. Real day-four values do not have that property: a day-four canonical form’s options are whatever day-three values survive domination and reversal, and those are not a uniform draw from day three. Some day-three values are likely to appear as options of day-four values far more often than others — plausibly the numbers and the small switches, which are the hardest to dominate away — though that has not been counted here.
If comparability depends on that — if day-four values whose options are “popular” day-three values are more or less often comparable than values built from arbitrary ones — both recipes would miss it by the same amount, and their agreement would be agreement on the wrong number. The calibration cannot catch it either, because at day three the options are day-two values and there are only twenty-two of them, so every recipe samples them nearly uniformly by necessity.
A third recipe would probe exactly this. Weight each day-three value by how often it appears as an option in the day-three canonical forms themselves — a measurable popularity — and build day four from the weighted draw. If the corrected figure moves, the uniform recipes shared a bias; if it lands with the other two, the floor has survived a third instrument built to break it. That is the obvious next measurement of this kind, and it needs nothing new beyond a weight.
The convention named
Comparability here means the game order — when Left, moving second, wins under normal play, which is the definition comparing positions starts from — and a pair is comparable when either or . The values are canonical forms, so each value is counted once however many forms it has. “Day four” means values whose canonical form is exactly four days deep, and the samples are built, not drawn uniformly: there is no known way to draw a uniform day-four value without listing the day.
What the dots cannot show
The error bars here are spreads over seeds and not proofs of anything. They say how much the procedure varies, not how far its average is from the truth, and if both recipes shared a bias that the day-three calibration does not see — some property of day-four forms that neither recipe builds — both would be wrong together and agree anyway. Agreement between two instruments rules out the errors in which they differ. It cannot rule out an error they share.
The comparison with zero, which the earlier essay found beyond rescue, is not revisited. The narrow recipe’s bias on that quantity was larger than the quantity itself, and nothing about a second seed changes that.
Still open: whether the floor is the same at day five
The next point on the curve is day five, and it cannot be built the same way: a recipe choosing day-four values needs a pool of day-four values to choose from, and the only pool available is itself a built sample. Building from a built pool compounds two biases, and the calibration trick — measure the bias one day lower, where the truth is known — has no day four to be calibrated against.
What can be done is to change the question rather than the day. The floor is a statement about pairs; what a floor at three fifths means for triples, or for the largest sets of mutually confused values, can be measured on the enumerated days and on the built one with the same two recipes and the same twenty seeds. If those quantities also stop moving at day four, the floor is a property of the order and not of one statistic; if they keep moving while pair comparability stands still, the pair statistic was levelling for a reason that does not generalise. How old a value is is where the birthday measure these days are counted in comes from.
Part 6 of 6
One argument about Comparison. The parts either side of it:
The objects named here
The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.
BirthdayCanonical formComparisonConfusionDay threeDominanceEnumerationNormal playPartial orderSampling
- A factor, and not an overhead canonical form, comparison, dominance, enumeration, normal play
- A threshold is a detection limit comparison, dominance, enumeration, partial order, sampling
- An option nobody would take canonical form, comparison, confusion, normal play, partial order
- Every chance but a certainty birthday, comparison, day three, enumeration, partial order
- One of four questions birthday, canonical form, comparison, enumeration, partial order
- The easy case was not the reason canonical form, comparison, dominance, enumeration, normal play