A threshold is a detection limit
Assumes: The threshold was a fact about the census · The margin a count needs
The threshold was a fact about the census found the mobility rule’s 72 failures at a margin of two sitting entirely on the largest board the census covered, and then found the rule failing at a margin of three on the first board bigger. It closed on the two points that gave:
The rung above is the threshold as a function of the board. Three at fifteen squares and four at eighteen are two points, and a rule of thumb whose bound grows with the board is a different object from one with a constant bound — it would mean the count degrades as the game gets big. A third point would say whether the growth is real, and the cheapest one available is a 4 × 4 at four squares covered.
The third point is a 3, and so are most of the ten after it, and the question turns out to have presupposed something false.
Nothing about the board orders them
Thirteen board-and-depth sweeps give thresholds of 1, 2, 3 and 4. Hold any property of the board fixed and ask whether the threshold stops varying — a grouping test, which is what settles whether a quantity’s stated inputs are its real ones — and every grouping is split:
- by area: twelve squares gives 1 and 2, eighteen gives 3 and 4;
- by columns: six columns gives all four values;
- by rows: three rows gives 2, 3 and 4;
- by the longer side: the same split as the columns;
- by how deep into the game the position is: four squares covered gives all four values.
The clearest counterexample is the last row of the sweep against the one above it. A 3 × 6 with four squares covered has a threshold of four; the same board with six covered has a threshold of three. One board, two answers, and the board is not what changed.
That already disposes of the question as it was asked. There is no function from boards to thresholds to be growing or not growing, because two sweeps of one board disagree.
The rate falls, and the threshold is where it stops being visible
What is actually going on is visible the moment the failures are quoted as rates rather than as a boundary.
On every board in the sweep the rule fails at a margin of one somewhere between 3 and 17 per cent of the time; at a margin of two, between nought and 7 per cent; at a margin of three, once, at 0.12 per cent. Averaged over the boards where both counts are non-zero, the rate falls by a factor of about seven for each extra unit of margin.
A rate that falls like that never reaches zero. It reaches the point where the number of pairs available at that margin is too small for the expected count to exceed one — and that is what a threshold records. The 3 × 6’s threshold of four is not a statement that the rule never fails at a margin of four. It is a statement that among the 452 pairs at that margin, none happened to.
The 3 × 6’s threshold of four rests on four failures among 3,340 pairs. Draw fewer of the same pairs and the chance of drawing none of the four is a product of four terms — computed rather than simulated, because the population is finite and known. Half the pairs miss all four 6 per cent of the time. A quarter miss them 32 per cent of the time. A tenth of the sweep reports a threshold of three two times in three, on the identical board.
So the number is not stable under looking at less of the same thing, which is the defining property of a detection limit and the opposite of the defining property of a bound. Which option the reduction keeps is the one place on this anchor where a claim about options is exact rather than statistical, and the contrast is instructive: an exact claim is about every pair, and a rate is about a population.
What the quiet boards are evidence of
The other half of the argument is what a threshold of three means when it is reported.
Take the 3 × 6’s measured rate at a margin of three — four failures in 3,340, so about 0.12 per cent — and ask each quiet board how surprising its silence would be at that rate. The 3 × 5 at four squares covered has 656 pairs at that margin and no failure, which would happen about 46 per cent of the time by chance alone. The 4 × 4 at four covered has 1,548 and would be silent 16 per cent of the time. The 2 × 8 has 640 and 46 per cent.
None of those is evidence that the board is better behaved. Each is a small sweep of a rare event, and a small sweep of a rare event is silent about half the time whatever the truth. The rung below’s three at fifteen squares was never a measurement of fifteen squares; it was a measurement of 656 pairs.
One board’s silence is not unsurprising, and it is the interesting one. The 3 × 6 covered six squares has 9,128 pairs at a margin of three and no failure among them, which at the 3 × 6’s own rate four squares covered would happen about once in fifty thousand times. That board really does behave differently from itself.
The one thing that moves
Five boards in the sweep are measured at two depths, and on every one of them both failure rates fall as the board fills up.
A 3 × 5 goes from 17.1 to 11.0 per cent at a margin of one and from 6.67 to 1.18 at a margin of two. A 3 × 6 goes from 16.1 to 10.7 and from 3.24 to 1.09. A 3 × 4 from 8.6 to 2.7 and from nought to nought. A 2 × 6 from 11.5 to nought. A 4 × 4 from 16.0 to 14.7 and from 1.89 to 1.51.
That is the only quantity in this sweep that moves the rule’s reliability in one direction, and it moves it in the direction a player would want: the mobility rule gets better as the game goes on. Three rules and a tie-break measured the same rule against two others on a played board rather than on a swept one, and its residue — the decisions no rule reaches — is the same population seen through a strategy instead of through a census. The rung below’s worry — that the count degrades as the game gets big — is not what the data says. What degrades the count is room, and a board fills up as it is played.
The mechanism is the wall, which the rung below identified. The count is wrong when the option leaving fewer replies is the one that cuts the board in two, and cutting the board in two needs a long uninterrupted run for a domino to stand across. Half the difference in odd runs is where the same quantity — how many uninterrupted runs a region still holds — turns out to price a Domineering position exactly, so the two readings are about one property of a board seen from two sides. A board with four squares covered still has those runs; the same board with six covered has fewer, so there are fewer positions where a wall can be built and fewer where the count can be fooled.
Why the question felt answerable
It is worth asking why does the threshold grow with the board? looked like a question with an answer, because the mistake is a common one and this ladder has now made it twice.
A threshold is reported as a single number, and a single number attached to a board invites being plotted against a property of the board. Three at fifteen squares and four at eighteen is two points on a graph, and two points on a graph suggest a line. Nothing in the reporting says that the four was four events and the three was none, or that the two boards were looked at with sweeps differing by a factor of four in size.
The general form is that a statistic which discards its own population invites comparisons its population cannot support. That is exactly what link density read as a site average is a warning about in a different setting, and it is what the rung below found one level up: a threshold that was a fact about which boards the census covered rather than about boards.
The fix is the same both times and it is not a better statistic. It is carrying the population along with the number. A rate of 0.12 per cent measured on 3,340 pairs and a rate of nought measured on 656 are two facts that can be compared; thresholds of four and three are two facts that cannot.
There is a second reason the question felt answerable, and it is about the shape of the subject rather than about statistics. Combinatorial game theory is full of quantities that genuinely are functions of a board — a Grundy sequence’s period, the temperature of an empty rectangle, whether a pairing strategy exists — and every one of them is computed rather than sampled. A ladder that has spent five rungs on computed quantities acquires the habit of expecting the next one to behave the same way. The mobility rule’s failures are not a computed quantity; they are a residue, and a residue has a rate rather than a value.
What the ladder should say instead
The word threshold should go, and the reason is not pedantry.
A threshold invites a reading it cannot support. The rule is safe at a margin of four is what a threshold of four sounds like, and what the sweep licenses is no failure at a margin of four was seen in 452 pairs. A player told the first would trust a comparison the data has nothing to say about.
A rate with its population carries the same information and cannot be misread that way. Sixteen per cent at a margin of one, three per cent at two, 0.12 per cent at three, nothing in 452 pairs at four is longer, is honest at every margin, does not move when the sweep grows, and lets a reader do the one calculation a threshold hides — how many comparisons they are about to make, and therefore how often the rule will let them down.
This is the same correction the ladder has now made three times at three levels. The margin a count needs turned a rule into a bound; the weight that blunts the count found the bound’s proposed explanation doing nothing; the rung below found the bound to be a fact about which boards were covered. Each step made the object smaller and more honest, and this one takes it to a rate, which is the smallest honest object there is.
When a measurement is about the instrument
The finding here is that a number the ladder had been quoting as a property of the game is a property of the sweep, and it is worth setting out the signature, because it is recognisable in advance.
A property of the game is stable under changes to the measurement. Vary the board, the depth, the sample, the order of enumeration, and a real threshold stays where it is.
A property of the instrument moves with the instrument and not with the subject. That is what shows here: no property of a board orders the thresholds, the same board at two depths gives two of them, and a tenth of the sweep that produced one number produces a different one. What does move, on every board measured twice, is the depth — which is a parameter of the sweep rather than a feature of any position.
A quantity that tracks a parameter of the measurement is a detection limit. It is the point below which the sweep cannot see the effect, and reporting it as a threshold of the game is reporting the resolution of the apparatus as a fact about what was being looked at.
The diagnostic is cheap and it is the one this rung runs. Measure the same subject twice with the instrument set differently. If the number moves, it belongs to the instrument; if it does not, it belongs to the subject. Two rungs of this ladder quoted a threshold without running that check, and the check costs one extra sweep at a different depth — which is less than either of those rungs cost.
What this does not say
The rule still works, and works well. At a margin of one — the hardest case, where the two options differ by a single reply — it is right between 83 and 97 per cent of the time on every board here. Nothing on this page weakens that; what it weakens is the claim that some margin makes it certain.
Thirteen sweeps is not many. The largest is a 3 × 6 at six squares covered, which is 18,551 positions and 117,656 pairs, and boards past eighteen squares are out of reach because every pair needs two comparisons of full Domineering games. A 4 × 5 was in the sweep and was removed: it did not finish.
The decay is measured over two margins. About sevenfold per unit of margin is the ratio of the margin-two rate to the margin-one rate, averaged over the seven boards where both are non-zero. There is one board in the sweep with a non-zero margin-three rate, so the claim that the decay continues is an extrapolation from a single point and is used only to price the silences.
And the subsample argument is about the pairs, not about play. Drawing a random tenth of the margin-three pairs is not the same as sweeping a smaller board, which would draw a structured subset. What it shows is that the threshold is unstable under a smaller look; it does not show that any particular smaller board would have missed the failures for that reason.
The convention, named
Normal play throughout: Left places vertical dominoes, Right horizontal ones, and a player who cannot place loses.
The mobility rule is the rule of thumb two rungs below: of two Left options, the one leaving the opponent fewer replies is the better one. A pair of options is comparable when one of the two games is at least the other; incomparable pairs and equal ones are left out of every count here, because the rule says nothing about them.
The margin of a pair is the difference between the two options’ reply counts. A failure is a comparable pair with a non-zero margin where the option leaving fewer replies is the worse one.
The threshold is one more than the largest margin at which a failure was seen — so it is the smallest margin at which the rule has never been observed to fail, on that sweep. It is that definition, and not the rule, that this page is about.
A board is swept at a depth: every occupancy of exactly that many covered squares, whether or not it is reachable by play. That is the rung below’s population and is kept for comparability, and it is a superset of the positions a game produces.
Where the ladder goes next
The dominance anchor has seven rungs: how much a list of options can lose, the margin a count needs, which option a reduction keeps, the weight that blunts the count, the reduction that always shrinks, the threshold as a fact about the census, and now the threshold as a detection limit.
The rung above is the rate as a function of the depth. Every board here shows the failure rate falling as the board fills, and each is two points; the shape of that fall is the thing worth having, because a rate that falls to nought by the endgame would mean the rule is exact where it matters most. Measuring it means sweeping one board at every depth rather than two, which on a 3 × 5 is eight sweeps and is affordable. That would turn the rule improves as the game goes on from a direction into a curve, and a curve is something a player can be given.
Two neighbours are worth the trip. The margin a count needs is where the threshold was invented, and reading it beside this page shows the whole life of a quantity — proposed as a bound, narrowed to a census, and dissolved into a rate. And a bound instead of an answer is the standard the site holds a rule of thumb to, and it is the form the mobility rule should have been given from the start.
Part 7 of 9
One argument about Dominance. The parts either side of it:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 9.
The objects named here
The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.
ApproximationBoundComparisonCounterexampleDominanceDomineeringEnumerationHeuristicMobilityPartial orderSamplingValue
- The moves a player can be talked out of approximation, bound, counterexample, domineering, enumeration, heuristic, value
- Two errors that cancel approximation, bound, domineering, enumeration, heuristic, sampling, value
- One domino every three cells approximation, bound, domineering, enumeration, heuristic, value
- The criterion that cannot exist approximation, bound, counterexample, enumeration, heuristic, value
- Where to stop building approximation, domineering, enumeration, heuristic, sampling, value
- A catalogue that builds itself approximation, domineering, enumeration, heuristic, mobility