Values

The threshold was a fact about the census

Two rungs failed to account for the seventy-two pairs where a mobility count gets the direction of a comparison wrong, and the third looks at them one at a time. They are not a class of shapes. All seventy-two are on the largest board in the census, at two depths, and sixteen positions up to symmetry — and one board larger the count fails at a margin of three, which the ladder has been quoting as the point at which it never does.

Assumes: The margin a count needs · The weight that blunts the count

The margin a count needs asked what the oldest rule of thumb in board games is worth as a statement — leave the opponent as few moves as possible — and got a threshold. Over 57,879 pairs of Left options on eight Domineering boards, the option leaving the opponent fewer replies is the worse of the two 1,052 times at a margin of one and 72 times at a margin of two, and at a margin of three or more never. The weight that blunts the count then tried to explain those 72 by weighing each reply according to what it leaves behind, and every weighting made the rule worse while leaving all 72 exactly where they were.

That page closed by naming the only route left:

The rung above is the pairs themselves. Seventy-two pairs have now survived two accounts of why the count fails on them, and they are a small enough set to be looked at rather than summarised: what those boards have in common, whether the option with more replies is the one whose replies are interchangeable, and whether that is a property with a count attached.

They were looked at. The first thing they have in common is not a shape.

Every failure is on one board. The eight Domineering boards of the mobility census with the number of failing pairs on each. Seven of them contribute none; every failure at a margin of two is on the largest board, at two depths, and sixteen positions up to symmetry.
Fig. 1 Every failure at a margin of two, by board. Seven of the eight boards contribute none; the eighth is the largest in the census, and it contributes all seventy-two.

Where the seventy-two are

Every one of them is on the 3 × 5. The 2 × 6 has twelve squares and produces none; the 3 × 4 has twelve and produces none; the 3 × 3, the 2 × 5, the 2 × 4 and the two smallest produce none. The failures are not distributed over the census at all — they sit entirely on its largest board, which is also the only board in it with fifteen squares.

They sit at two depths, and only two. Thirty-six of them are at positions with two squares covered and thirty-six at positions with four; nothing deeper fails at a margin of two, and nothing shallower can, since an empty 3 × 5 is one position rather than a population. So the failures are early-game positions on the biggest board available, which is exactly the region of the census where the option lists are longest and the values furthest from anything a count could reach.

And they are fewer than they look. A rectangle that is not a square has four symmetries — the identity, the two reflections and the half-turn — and the census sweeps every occupancy independently, so a failing position is counted once for each image of itself. Under those four symmetries the seventy-two failing pairs are sixteen positions. A set that two rungs have described as small enough to read one at a time is small for a duller reason than either of them supposed: the census stops one board short of where the phenomenon lives.

The guess the rung below made

The proposal was interchangeability: perhaps the option with more replies has replies that are all much the same, so a count of them over-states what the opponent really has. That is a real quantity and it is cheap to compute — instead of counting an option’s replies, count the distinct values among them, which is the number of genuinely different positions the opponent can reach rather than the number of dominoes they can put down.

Counting the different replies instead. The same pairs sorted by the difference in how many distinct values their replies carry, rather than by how many replies there are. The rule that results is worse than the one it was meant to repair: its threshold is six against three.
Fig. 2 The same pairs sorted by the difference in distinct reply values rather than in replies. The rule this produces is worse than the one it was meant to repair: it is wrong at margins of one, two, three, four and five, so its threshold is six against the plain count’s three.

The count of distinct replies is a worse rule than the count of replies. It is wrong at every margin up to five, so a bound stated in it needs a margin of six — twice the plain count’s — and the extra work bought nothing, since computing the distinct values means evaluating every reply position rather than listing every reply.

Worse for the hypothesis, the direction is wrong. On 56 of the 72 failures it is the option leaving fewer replies that also leaves fewer distinct ones, and on sixteen it is the other way; on none of them are the two counts level. If the option with more replies were the one whose replies were interchangeable, its distinct count would be the smaller, and three times in four it is the larger. The guess was not merely unhelpful. It described the pairs backwards.

What they do have in common

Something else is true of them, and it is visible in the position rather than in the count.

What the seventy-two have in common. The failing pairs sorted by which of the two options breaks the board into more pieces. On sixty of the seventy-two it is the option leaving the opponent fewer replies, and on none of them is it the other one.
Fig. 3 The failing pairs sorted by which option breaks the board into more pieces. On sixty of the seventy-two it is the option leaving the opponent fewer replies, and on none of them is it the other one.

On sixty of the seventy-two the option leaving fewer replies is the one that cuts the board into two pieces or three, and the option leaving more replies leaves the free squares in one piece. On the remaining twelve both options leave one piece. Not one failure has it the other way round — the count of replies has never yet been lowered by the option that keeps the board whole while its rival splits it.

The mechanism is one sentence long. Left plays vertically, so a Left domino occupies two squares of one column; Right plays horizontally, so every Right placement needs two adjacent squares of one row. A Left domino standing in the middle of a row-span removes every Right placement that would have straddled the column it stands in — several at a stroke — which is precisely what a wall does. The count sees the replies vanish and reads it as the opponent being held down. What has actually happened is that Left has spent a move in the middle of the board, where it takes room from Left as well.

One of the sixteen positions. A three by five Domineering board with two of Left's moves drawn as the boards they leave. The middle move leaves Right two fewer replies and is the worse of the two; it is also the move that cuts the board in half.
Fig. 4 One of the sixteen positions, with two of Left’s moves drawn as the boards they leave. The middle move leaves Right two fewer replies and cuts the board in half; the edge move leaves Right seven replies and the board whole, and it is the better of the two.

Both values were computed by the recursion and neither is a number, so what is at stake in each is a genuine fight and not an accounting exercise. The middle move is worth {{3/21/2}{13}}\{\{3/2 \mid -1/2\} \mid \{-1 \mid -3\}\} and the edge move {{20}{1/22}}\{\{2 \mid 0\} \mid \{-1/2 \mid -2\}\}, and the edge move is greater in the game order — greater in every sum it could sit in, which is what the comparison means and is a much stronger statement than winning more often.

A description is not a prediction

Sixty out of seventy-two is a strong enough coincidence to want to be a rule, so the wall was scored the way any rule on this site is scored: over the whole population rather than over the cases that suggested it.

The wall describes and does not predict. The whole census split by whether the option leaving fewer replies breaks the board into more pieces. The failure rate is higher when it does, and only by a third — which is a description of the failing set rather than a rule for finding it.
Fig. 5 The census split by whether the option leaving fewer replies breaks the board into more pieces. At a margin of one the failure rate is 5.69 per cent when it does and 4.17 per cent when neither does, which is a raised rate rather than an explanation.

At a margin of one the count fails 5.69 per cent of the time when the fewer-replies option splits the board and 4.17 per cent when neither option splits it. That is a real difference and it is nothing like the difference the failing set implies: three quarters of the pairs at that margin have the fewer option splitting, so a condition satisfied by three quarters of everything will be satisfied by most of any subset picked out of it.

This is the trap the whole ladder has been walking into from the other direction. A property that every member of a set has is not a property that predicts membership, and the way to tell the two apart is to apply it to the pairs that did not fail. Sixty of the seventy-two split, and so do 7,326 pairs at the same margin that the count gets right.

The quantity that does predict it

One property does separate the failures from the rest, and it is the one the count exists in order not to compute.

The quantity the count cannot see. The census split by which of the two options is the hotter. The count fails nearly five times as often when the option leaving the opponent fewer replies is the hotter of the two, which is the best predictor of failure anywhere in this ladder and the most expensive.
Fig. 6 The census split by which of the two options is the hotter. At a margin of one the count fails 14.89 per cent of the time when the option leaving the opponent fewer replies is the hotter of the two, and 3.12 per cent when it is the colder.

At a margin of one the count is wrong 14.89 per cent of the time when the option leaving fewer replies is the hotter of the two, and 3.12 per cent when it is the colder — nearly five times as often. At a margin of two the same split is 3.65 per cent against 0.37 per cent, a factor of ten. Nothing else measured anywhere in this ladder separates the two populations by that much.

The reason is not mysterious once stated. A count of the opponent’s replies is a statement about how much room is left; temperature is a statement about how much is at stake in moving. A hot option is worth taking even when it leaves the opponent plenty of room, because the fight in it is worth more than the room is — and the count has no access to that, since it reads the board after the move and not the game inside it.

Which is exactly why the finding does not produce a better rule. Knowing which of two options is hotter means computing two thermographs, and the whole point of a mobility count is to decide between options without computing anything about them. The repair is available and it costs precisely what the heuristic was invented to save. That is the same shape of result as the moves a player can be talked out of, where a count that was wrong turned out to be repairable only by an interval that had to be computed — and it is worth reading the two together, because a heuristic priced against the thing it replaces is the honest way to state what it is for.

It also accounts for only 32 of the 72. Twenty-eight of the failures have the colder option leaving fewer replies and twelve have the two at the same temperature, so temperature predicts failure across the census and does not describe this particular set. The wall describes the set and does not predict. Neither is the account the rung below was asking for, and after three rungs the honest position is that the failures at a margin of two have no single account.

One board larger

If the seventy-two are a fact about the largest board in the census rather than about Domineering, there is an obvious test, and it has a definite answer: run the same census on a bigger board and see whether the threshold moves.

A 3 × 6 has eighteen squares, so sweeping every even occupancy of it is out of reach — but it does not need to be swept. Every failure on the 3 × 5 sits at two squares covered or at four, and the deeper of those two is the cheaper, so the sweep here is every 3 × 6 occupancy with four squares covered: 3,060 positions, the same pairs of Left options, the same three relations kept apart.

One board larger, and the threshold moves. The same census run on a three by six board with four squares covered. The count of the opponent's replies is wrong at a margin of three, which it never is on any smaller board — so the threshold of three is a property of the boards it was measured on.
Fig. 7 The same census on a three by six board at four squares covered. The count of the opponent’s replies is wrong four times at a margin of three, which it never is on any board of fifteen squares or fewer, and all four have the fewer-replies option splitting the board.

The count is wrong at a margin of three. Four times, out of 3,340 pairs at that margin, and all four with the fewer-replies option cutting the board — which makes them the same phenomenon as the sixty, one board and one margin further out.

So the threshold is not three. Three is what the count needs on Domineering boards of at most fifteen squares, and four is what it needs once eighteen are allowed, and nothing in either measurement suggests a value it settles at. The rule two rungs of this ladder have been sharpening is a statement about a population, and the population was the eight boards small enough for an exact evaluator to sweep exhaustively. That is a real result about a real census, and it is not a fact about the game.

What this does not say

It does not say the count is useless. At a margin of three the count is right on 1,584 of the 1,596 comparable pairs it speaks about, and the twelve exceptions are confused rather than reversed. A rule that names the better option correctly whenever it speaks with a gap of three is a good rule; what it is not is a rule with a threshold that can be quoted without naming the boards.

The margin-three failures are four. Four out of 3,340 is a rate of one in eight hundred, and a reader entitled to be sceptical of a threshold read off 1,596 pairs is entitled to be equally sceptical of one refuted by four. What the four establish is that the threshold is not a law; what they do not establish is where it goes next, and the honest reading of the 3 × 6 sweep is that it moved by one and might move again.

One depth on the larger board. The 3 × 6 sweep is at four squares covered, and the shallower depth — two covered, sixteen free — was swept separately and produced no failure at a margin of three at all. The two depths are not interchangeable, and the finding is stated at the depth that has it.

And Domineering only. Everything here is a count of the opponent’s dominoes on a rectangle, and the wall mechanism is a fact about a game where one player’s pieces are vertical and the other’s horizontal. In Clobber, where a move removes a piece rather than filling space, there is no wall to build and the whole account would have to be restated before it could be tested.

The convention, named

Normal play throughout: a player who cannot move loses.

An option is the board after one Left move, not the move itself, and a reply is a Right move from that board. The count read by the rule is the number of Right moves available on the board the option leaves.

Better and worse are the game order rather than the outcome: option AA is better than BB for Left when ABA \ge B, which is a claim about every sum either of them could be part of, decided by comparison as a search. Pairs that are equal and pairs that are confused are counted in their own columns and never credited to the rule, since a rule scored on the pairs it cannot distinguish scores well by saying nothing.

Regions are connected pieces of free board, counted through edges and not corners, which is the connectivity Domineering actually decomposes along — the board falls apart is where that is established and priced.

Temperature is the height at which a thermograph’s walls meet, and a number is treated as colder than every hot position, which is what makes the comparison of two options well defined when one of them has stopped being a fight.

A failure is a comparable pair, at a margin of one or more, in which the option leaving the opponent fewer replies is the worse of the two. The margin is the difference in replies and never the difference in anything else, which is what makes this census and the two below it the same census.

Where the ladder goes next

The dominance anchor has six rungs: what deleting removes, what decides how much, what a survivor looks like on a board, how large a mobility gap has to be before it decides a comparison, what happens when the gap is measured more carefully, and now what the pairs it fails on turn out to be.

The rung above is the threshold as a function of the board. Three at fifteen squares and four at eighteen are two points, and a rule of thumb whose bound grows with the board is a different object from one with a constant bound — it would mean the count degrades as the game gets big, which is the opposite of what a rule of thumb is usually assumed to do. A third point would say whether the growth is real, and the cheapest one available is a 4 × 4 at four squares covered: sixteen squares, a different aspect ratio, and the same sweep.

Two neighbours are worth the trip. A rule that is never right and cannot be far wrong is where a heuristic is priced properly, with a bound it satisfies rather than a score it achieves, and it is the standard this ladder has been trying to reach for three rungs. And two errors that cancel is the other place on this site where a reading is tested by being added across a whole board rather than checked on the positions that suggested it, and it is the model for what the paragraph above is asking for.

Part 6 of 9

One argument about Dominance. The parts either side of it:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.

ApproximationBoundComparisonCounterexampleDecompositionDominanceDomineeringEnumerationHeuristicMobilityNormal playPartizanTemperature