The threshold was a fact about the census
Assumes: The margin a count needs · The weight that blunts the count
The margin a count needs asked what the oldest rule of thumb in board games is worth as a statement — leave the opponent as few moves as possible — and got a threshold. Over 57,879 pairs of Left options on eight Domineering boards, the option leaving the opponent fewer replies is the worse of the two 1,052 times at a margin of one and 72 times at a margin of two, and at a margin of three or more never. The weight that blunts the count then tried to explain those 72 by weighing each reply according to what it leaves behind, and every weighting made the rule worse while leaving all 72 exactly where they were.
That page closed by naming the only route left:
The rung above is the pairs themselves. Seventy-two pairs have now survived two accounts of why the count fails on them, and they are a small enough set to be looked at rather than summarised: what those boards have in common, whether the option with more replies is the one whose replies are interchangeable, and whether that is a property with a count attached.
They were looked at. The first thing they have in common is not a shape.
Where the seventy-two are
Every one of them is on the 3 × 5. The 2 × 6 has twelve squares and produces none; the 3 × 4 has twelve and produces none; the 3 × 3, the 2 × 5, the 2 × 4 and the two smallest produce none. The failures are not distributed over the census at all — they sit entirely on its largest board, which is also the only board in it with fifteen squares.
They sit at two depths, and only two. Thirty-six of them are at positions with two squares covered and thirty-six at positions with four; nothing deeper fails at a margin of two, and nothing shallower can, since an empty 3 × 5 is one position rather than a population. So the failures are early-game positions on the biggest board available, which is exactly the region of the census where the option lists are longest and the values furthest from anything a count could reach.
And they are fewer than they look. A rectangle that is not a square has four symmetries — the identity, the two reflections and the half-turn — and the census sweeps every occupancy independently, so a failing position is counted once for each image of itself. Under those four symmetries the seventy-two failing pairs are sixteen positions. A set that two rungs have described as small enough to read one at a time is small for a duller reason than either of them supposed: the census stops one board short of where the phenomenon lives.
The guess the rung below made
The proposal was interchangeability: perhaps the option with more replies has replies that are all much the same, so a count of them over-states what the opponent really has. That is a real quantity and it is cheap to compute — instead of counting an option’s replies, count the distinct values among them, which is the number of genuinely different positions the opponent can reach rather than the number of dominoes they can put down.
The count of distinct replies is a worse rule than the count of replies. It is wrong at every margin up to five, so a bound stated in it needs a margin of six — twice the plain count’s — and the extra work bought nothing, since computing the distinct values means evaluating every reply position rather than listing every reply.
Worse for the hypothesis, the direction is wrong. On 56 of the 72 failures it is the option leaving fewer replies that also leaves fewer distinct ones, and on sixteen it is the other way; on none of them are the two counts level. If the option with more replies were the one whose replies were interchangeable, its distinct count would be the smaller, and three times in four it is the larger. The guess was not merely unhelpful. It described the pairs backwards.
What they do have in common
Something else is true of them, and it is visible in the position rather than in the count.
On sixty of the seventy-two the option leaving fewer replies is the one that cuts the board into two pieces or three, and the option leaving more replies leaves the free squares in one piece. On the remaining twelve both options leave one piece. Not one failure has it the other way round — the count of replies has never yet been lowered by the option that keeps the board whole while its rival splits it.
The mechanism is one sentence long. Left plays vertically, so a Left domino occupies two squares of one column; Right plays horizontally, so every Right placement needs two adjacent squares of one row. A Left domino standing in the middle of a row-span removes every Right placement that would have straddled the column it stands in — several at a stroke — which is precisely what a wall does. The count sees the replies vanish and reads it as the opponent being held down. What has actually happened is that Left has spent a move in the middle of the board, where it takes room from Left as well.
Both values were computed by the recursion and neither is a number, so what is at stake in each is a genuine fight and not an accounting exercise. The middle move is worth and the edge move , and the edge move is greater in the game order — greater in every sum it could sit in, which is what the comparison means and is a much stronger statement than winning more often.
A description is not a prediction
Sixty out of seventy-two is a strong enough coincidence to want to be a rule, so the wall was scored the way any rule on this site is scored: over the whole population rather than over the cases that suggested it.
At a margin of one the count fails 5.69 per cent of the time when the fewer-replies option splits the board and 4.17 per cent when neither option splits it. That is a real difference and it is nothing like the difference the failing set implies: three quarters of the pairs at that margin have the fewer option splitting, so a condition satisfied by three quarters of everything will be satisfied by most of any subset picked out of it.
This is the trap the whole ladder has been walking into from the other direction. A property that every member of a set has is not a property that predicts membership, and the way to tell the two apart is to apply it to the pairs that did not fail. Sixty of the seventy-two split, and so do 7,326 pairs at the same margin that the count gets right.
The quantity that does predict it
One property does separate the failures from the rest, and it is the one the count exists in order not to compute.
At a margin of one the count is wrong 14.89 per cent of the time when the option leaving fewer replies is the hotter of the two, and 3.12 per cent when it is the colder — nearly five times as often. At a margin of two the same split is 3.65 per cent against 0.37 per cent, a factor of ten. Nothing else measured anywhere in this ladder separates the two populations by that much.
The reason is not mysterious once stated. A count of the opponent’s replies is a statement about how much room is left; temperature is a statement about how much is at stake in moving. A hot option is worth taking even when it leaves the opponent plenty of room, because the fight in it is worth more than the room is — and the count has no access to that, since it reads the board after the move and not the game inside it.
Which is exactly why the finding does not produce a better rule. Knowing which of two options is hotter means computing two thermographs, and the whole point of a mobility count is to decide between options without computing anything about them. The repair is available and it costs precisely what the heuristic was invented to save. That is the same shape of result as the moves a player can be talked out of, where a count that was wrong turned out to be repairable only by an interval that had to be computed — and it is worth reading the two together, because a heuristic priced against the thing it replaces is the honest way to state what it is for.
It also accounts for only 32 of the 72. Twenty-eight of the failures have the colder option leaving fewer replies and twelve have the two at the same temperature, so temperature predicts failure across the census and does not describe this particular set. The wall describes the set and does not predict. Neither is the account the rung below was asking for, and after three rungs the honest position is that the failures at a margin of two have no single account.
One board larger
If the seventy-two are a fact about the largest board in the census rather than about Domineering, there is an obvious test, and it has a definite answer: run the same census on a bigger board and see whether the threshold moves.
A 3 × 6 has eighteen squares, so sweeping every even occupancy of it is out of reach — but it does not need to be swept. Every failure on the 3 × 5 sits at two squares covered or at four, and the deeper of those two is the cheaper, so the sweep here is every 3 × 6 occupancy with four squares covered: 3,060 positions, the same pairs of Left options, the same three relations kept apart.
The count is wrong at a margin of three. Four times, out of 3,340 pairs at that margin, and all four with the fewer-replies option cutting the board — which makes them the same phenomenon as the sixty, one board and one margin further out.
So the threshold is not three. Three is what the count needs on Domineering boards of at most fifteen squares, and four is what it needs once eighteen are allowed, and nothing in either measurement suggests a value it settles at. The rule two rungs of this ladder have been sharpening is a statement about a population, and the population was the eight boards small enough for an exact evaluator to sweep exhaustively. That is a real result about a real census, and it is not a fact about the game.
What this does not say
It does not say the count is useless. At a margin of three the count is right on 1,584 of the 1,596 comparable pairs it speaks about, and the twelve exceptions are confused rather than reversed. A rule that names the better option correctly whenever it speaks with a gap of three is a good rule; what it is not is a rule with a threshold that can be quoted without naming the boards.
The margin-three failures are four. Four out of 3,340 is a rate of one in eight hundred, and a reader entitled to be sceptical of a threshold read off 1,596 pairs is entitled to be equally sceptical of one refuted by four. What the four establish is that the threshold is not a law; what they do not establish is where it goes next, and the honest reading of the 3 × 6 sweep is that it moved by one and might move again.
One depth on the larger board. The 3 × 6 sweep is at four squares covered, and the shallower depth — two covered, sixteen free — was swept separately and produced no failure at a margin of three at all. The two depths are not interchangeable, and the finding is stated at the depth that has it.
And Domineering only. Everything here is a count of the opponent’s dominoes on a rectangle, and the wall mechanism is a fact about a game where one player’s pieces are vertical and the other’s horizontal. In Clobber, where a move removes a piece rather than filling space, there is no wall to build and the whole account would have to be restated before it could be tested.
The convention, named
Normal play throughout: a player who cannot move loses.
An option is the board after one Left move, not the move itself, and a reply is a Right move from that board. The count read by the rule is the number of Right moves available on the board the option leaves.
Better and worse are the game order rather than the outcome: option is better than for Left when , which is a claim about every sum either of them could be part of, decided by comparison as a search. Pairs that are equal and pairs that are confused are counted in their own columns and never credited to the rule, since a rule scored on the pairs it cannot distinguish scores well by saying nothing.
Regions are connected pieces of free board, counted through edges and not corners, which is the connectivity Domineering actually decomposes along — the board falls apart is where that is established and priced.
Temperature is the height at which a thermograph’s walls meet, and a number is treated as colder than every hot position, which is what makes the comparison of two options well defined when one of them has stopped being a fight.
A failure is a comparable pair, at a margin of one or more, in which the option leaving the opponent fewer replies is the worse of the two. The margin is the difference in replies and never the difference in anything else, which is what makes this census and the two below it the same census.
Where the ladder goes next
The dominance anchor has six rungs: what deleting removes, what decides how much, what a survivor looks like on a board, how large a mobility gap has to be before it decides a comparison, what happens when the gap is measured more carefully, and now what the pairs it fails on turn out to be.
The rung above is the threshold as a function of the board. Three at fifteen squares and four at eighteen are two points, and a rule of thumb whose bound grows with the board is a different object from one with a constant bound — it would mean the count degrades as the game gets big, which is the opposite of what a rule of thumb is usually assumed to do. A third point would say whether the growth is real, and the cheapest one available is a 4 × 4 at four squares covered: sixteen squares, a different aspect ratio, and the same sweep.
Two neighbours are worth the trip. A rule that is never right and cannot be far wrong is where a heuristic is priced properly, with a bound it satisfies rather than a score it achieves, and it is the standard this ladder has been trying to reach for three rungs. And two errors that cancel is the other place on this site where a reading is tested by being added across a whole board rather than checked on the positions that suggested it, and it is the model for what the paragraph above is asking for.
Part 6 of 9
One argument about Dominance. The parts either side of it:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
The objects named here
The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.
ApproximationBoundComparisonCounterexampleDecompositionDominanceDomineeringEnumerationHeuristicMobilityNormal playPartizanTemperature
- Half the difference in odd runs approximation, bound, counterexample, decomposition, domineering, enumeration, heuristic, normal play
- Three rules and a tie-break approximation, counterexample, decomposition, domineering, enumeration, heuristic, mobility, normal play
- A heuristic that becomes a theorem approximation, counterexample, decomposition, domineering, enumeration, heuristic, mobility
- An effect that changes sign approximation, counterexample, decomposition, enumeration, heuristic, mobility, temperature
- How wrong a nearly-independent split is approximation, bound, counterexample, decomposition, domineering, enumeration, temperature
- One domino every three cells approximation, bound, decomposition, domineering, enumeration, heuristic, partizan