Values

The weight that blunts the count

The rung below found that counting the opponent's replies gets the direction of a comparison right once the gap reaches three, and proposed a repair: weigh each reply by whether it leaves the opponent anything. Weighing it makes the count worse. The threshold goes from three to four, the failures from 1,124 to 1,320, and all seventy-two of the pairs the repair was written for come through it unchanged.

Assumes: The margin a count needs · Which option the reduction keeps

The margin a count needs took the oldest rule of thumb in board games — leave the opponent as few moves as possible — and asked what it is worth as a statement. The answer was a threshold. Over 57,879 pairs of Left options on eight Domineering boards, the option leaving the opponent fewer replies is the worse of the two 1,052 times at a margin of one and 72 times at a margin of two, and at a margin of three or more it never is.

That page closed by naming the repair its own failures suggested:

The seventy-two failures at a margin of two all have the same shape — the option leaving more replies leaves them useless — and turning useless into a count is the piece of work: replies weighted by whether they in turn leave the opponent anything, which is one more ply of the same measurement. If a weighted count has a threshold of one, it would replace this page’s rule with a much better one, and if it has no threshold at all that is worth knowing too.

The weighted count has a threshold. It is four, which is worse than the plain count’s three, and every other way of weighing a reply is worse still.

Four ways to count a reply. The plain count of the opponent's replies against three weightings of it, each scored on the same pairs of Domineering options. Every weighting has a larger threshold than the plain count and gets more pairs wrong.
Fig. 1 Four ways of counting what an option leaves the opponent, scored on one population of pairs. The plain count is the rung below’s; the three below it each weigh a reply by something. Every weighting needs a wider margin than the count it was meant to sharpen, and gets more pairs wrong on the way there.

What a weighted reply is

A reply is live when the player who made it still has a move afterwards, and dead when it does not. The weighted count of an option is the number of live replies it leaves; the plain count is all of them.

Everything else is the rung below’s census unchanged — the same eight boards from 2 × 2 to 3 × 5, the same unordered pairs of Left options, the same three relations kept apart so that pairs of equal value are not counted as a success. Both counts are taken on one pass over one population, so the comparison between them is not a comparison of two experiments.

Two further weightings are computed beside the live one, because a single alternative that behaves a certain way is an anecdote:

  • deep counts each reply once for every move it leaves the replier, so a reply opening three further moves counts three;
  • damage counts a reply by how many of Left’s moves it removes rather than by what it leaves Right, which is the same idea pointed at the other player.

The four counts are four functions of the same option. What differs is only which replies they notice.

There is something to weigh

The first thing to check is that the quantity exists. A weighting that never separates two options from each other cannot improve anything or spoil it, and dead replies could easily have been a rarity.

How many replies lead nowhere. Replies after which the replier has no further move, counted per board. A fifth of all replies are dead, and on the smallest boards it is three fifths.
Fig. 2 Replies after which the replier has no move at all, counted per board. A fifth of all replies are dead, and on the smallest boards it is three fifths — so the weighting has plenty to work with, and it is largest exactly where the boards are smallest.

Of 98,944 replies, 20,996 are dead — 21 per cent. The share falls steeply with the size of the board, from 62 per cent on 2 × 4 to 20 per cent on 3 × 5, which is what one would expect: a large board has room left over after a domino goes down. On between a third and a half of the options every reply is dead, and those options have a weighted count of nought.

So the weighting is not a formality. It changes the count on a large fraction of the options and it collapses a great many of them to zero. Whether that is an improvement is the rest of this page.

The census, and the wider margin

The rung below’s table is the thing to hold beside it. It is the same population and the same three relations, counted without the weight.

The margin the count needs. Every pair of Left options sorted by how many more replies one leaves the opponent than the other. At a margin of three the option leaving fewer replies is never the worse one.
Fig. 3 The plain count, from the rung below. The option leaving the opponent fewer replies is the worse of the two 1,052 times at a margin of one and 72 at a margin of two, and never at three — 1,736 pairs reach that margin.
The live count, margin by margin. Every pair of Left options sorted by the difference in live replies — replies after which the replier still has a move. The option leaving fewer of them is the worse one at margins of one, two and three.
Fig. 4 The rung below’s table with dead replies not counted. The option leaving fewer live replies is the worse of the two at a margin of one, of two, and of three — so the threshold is four, and the pairs sitting at three or more have gone from 1,736 to 7,020.

The live count is wrong 1,052 times at a margin of one, 252 times at a margin of two and 16 times at a margin of three. Nothing fails at four or above, so the threshold is four.

Both halves of that are worse than what the plain count buys. The threshold is one higher, so a player needs a wider gap before the rule is safe to trust. And the failures are more numerous: 1,320 against 1,124, which is 17 per cent more pairs on which the rule points the wrong way.

The third number is the one that matters most to somebody who wants to use a rule. A threshold is only useful over the pairs that reach it, and the two counts spread their pairs differently: 1,736 pairs reach a margin of three under the plain count, against 7,020 under the live one. Weighting makes the margins look larger while making the guarantee arrive later, which is exactly the combination a rule of thumb should not have — more apparent separation, less real information.

The other two weightings are not close. Damage — counting a reply by the Left moves it removes — is wrong 4,004 times and needs a margin of ten. Deep needs twenty-three, which on these boards is very nearly no guarantee at all, since a margin of twenty-three arises on 244 pairs out of 57,879.

The seventy-two, followed through

The proposal was not a general hope. It was a diagnosis of a specific set: the 72 pairs at a margin of two where the plain count points the wrong way, whose shape the rung below described as the option leaving more replies leaves them useless. If that diagnosis is right, the weighted count should get those 72 right.

The pairs the weighting was built for. The pairs where the plain reply count points the wrong way, tabulated against what the live count says about the same pairs. The seventy-two at a margin of two are unchanged by the weighting.
Fig. 5 Every pair the plain count gets wrong, followed through the weighting proposed to repair it. The seventy-two at a margin of two keep a margin of two and stay wrong — every one of them. The weighting changes which pairs it is wrong about; it does not reduce them.

All 72 arrive at the far side unchanged. Their live margin is two, exactly as their plain margin was, and the option leaving fewer live replies is still the worse one in every case. The weighting does not touch them.

That is a stronger refutation than a worse score. A weighting could easily have fixed the 72 and broken more elsewhere, and then the argument would be about how to trade one against the other. Instead the extra replies in those pairs are not dead replies at all: whatever it is that makes the option with more replies the better one there, it is not that its replies lead nowhere.

The 1,052 failures at a margin of one break up differently: 972 keep a margin of one, 12 move to two, 16 move to three, and 52 collapse to a margin of nought. The 52 are the only pairs the weighting helps, and it helps them by abstaining — a margin of nought is a count with no opinion, not a count with the right one. That is worth about as much as a rule that declines to answer.

The count that speaks where the other was silent

If the weighting rescues 52 pairs by falling silent, it should pay for that somewhere, and the ledger says where.

Of the live count’s 1,320 failures, 248 are pairs the plain count declined to judge: 168 at a live margin of two and 80 at a margin of one, all of them pairs where the two options leave the opponent the same number of replies. The plain count has no opinion about such a pair. The weighted count invents a difference between them, and on those 248 the difference points the wrong way.

A difference that was not there. Two Domineering options leaving the opponent the same number of replies but different numbers of live ones. The weighted count separates them and points at the worse of the two.
Fig. 6 Two Left options on one 3 × 3 board. Both leave Right two replies, so the plain count says nothing about the pair. One of them has two live replies and the other none, so the weighted count prefers the second — which is worth −1/2 and is the worse of the two for Left.

The pair is worth reading in full, because it contains the whole mechanism.

The option worth −1/2 leaves Right two replies, and after either of them Right has no move. Both replies are dead, so its weighted count is nought — the best score available. The option worth 0 leaves Right two replies, each of which leaves Right another move, so its weighted count is two. The weighting therefore ranks the first above the second, and the first is the one that loses half a move.

Why a dead reply is not a wasted one

The mechanism is one line of the normal play convention, and once it is said the result stops being surprising.

A player who cannot move loses. So a reply that leaves the replier with nothing is not a reply that achieved nothing — it may be precisely the move that ends the game with the opponent to play. On the board above, one of Right’s two dead replies leaves Left with no move either, and it is therefore Right’s winning move. The weighting scores it as worth nothing.

That is the flaw, and it is not a matter of degree. Weighting by liveness assumes that a move’s value lies in what it opens up, which is a reasonable assumption in a game where the object is to accumulate something and a wrong one in a game where the object is to be the last to move. In normal play, the terminal move is the whole point. A count that discards exactly the moves that end games is a count that discards exactly the moves that win them.

The same reading explains the ranking of the other two weightings. Deep does more of what live does — it rewards a reply in proportion to how much it opens up — and it does worse, needing a margin of twenty-three. Damage at least measures something about the position rather than about what remains for the mover, and it is worse again in failures, because Left’s own mobility is not what dominates Left’s options; Right’s is.

The order of the four counts is therefore not a table of arbitrary results. It is one statement — the closer a count gets to rewarding replies that continue the game, the worse it predicts a comparison — measured four times.

What the rung below got right and wrong

Two things separate cleanly here.

The threshold was right. The plain count’s threshold of three is reproduced exactly by this census, on the same population, and it is asserted rather than reported: if it came out anywhere else, the two censuses would not be measuring the same thing and the comparison in this page would be worthless.

The diagnosis was wrong. The option leaving more replies leaves them useless was an inference from looking at the failures, and looking at failures is how a plausible mechanism gets attached to a set of counterexamples. The 72 pairs are still there after the useless replies have been discounted, so whatever their shape is, it is not that one.

There is a general moral in the pair, and this site has now met it twice. The bend is the condition found a proposed cause that turned out to be a one-way implication; this one finds a proposed cause that turns out not to be a cause at all. A section closing an essay is the most inviting place in the collection to state a mechanism, and it is the place where nothing checks one.

A repair that makes things worse is more informative than one that helps

The refinement fails in every direction at once, and that unanimity is worth more than a partial success would have been. It is worth saying why.

A repair that helped a little would leave the diagnosis open. Perhaps the weighting is right and the constant is wrong; perhaps it helps on some shapes and hurts on others; perhaps a third refinement on top would finish the job. Each of those is a further sweep, and none of them is ruled out.

A repair that moves the threshold the wrong way, raises the failure count, and leaves every one of the seventy-two target pairs unchanged has ruled itself out completely. There is no version of it to tune. And the last clause is the sharp one: the pairs it was written for come through it unchanged, which means the quantity it adds is constant across exactly the cases that were supposed to distinguish them.

That is a determination rather than a score. If weighting a reply by whether it leaves the opponent anything gives the same value on both members of a failing pair, then no rule using that weighting separates them, and the whole family of weighted counts is refuted rather than one member of it.

So the useful output of this rung is a narrowed search rather than a smaller error. What is left has to be a quantity that differs across those seventy-two pairs, and the rungs above go looking for one — and find, eventually, that the pairs are not a class at all, which is a conclusion only reachable once the obvious repairs have been eliminated rather than merely scored.

What this does not say

Four limits, and the first is the important one.

Weighting is not tested; four weightings are. Live, deep and damage are three functions somebody would actually write, and the space of possible weightings is not exhausted by them. What the page establishes is that the specific repair the rung below proposed fails, and that two natural neighbours of it fail worse. A weighting built the other way round — rewarding a reply that leaves the opponent nothing — is a different experiment.

Eight small boards. The largest is 3 × 5, which is fifteen squares, and the dead-reply share is still falling steeply at that size. A census on 4 × 5 would have a smaller share of dead replies and might well show a smaller effect; it would not show the effect reversing, because the 72 pairs would still not be about dead replies.

Domineering only. The count is a count of the opponent’s dominoes, and the rule of thumb is about that. In a game where a move does not remove a fixed amount of room — Clobber, say, where a move removes a piece and can open lines for either player — the whole family of counts would have to be redefined before any of this could be asked.

And a threshold is a bound on the sign, not on the size. That was true of the plain count and is true of every weighting here. None of these counts tells a player how much better one option is; they tell a player which way the comparison runs, once the gap is wide enough. The quantity that would answer the other question is the value, and computing it is the thing the rule of thumb exists to avoid.

The convention, named

Normal play throughout: a player who cannot move loses, and it is the clause the whole result turns on.

An option is the board after one Left move, not the move itself. A reply is a Right move from that board. Live and dead describe a reply by what the replier has afterwards, which is the reading the rung below’s sentence asks for — whether they in turn leave the opponent anything, where the opponent of the option-holder is the one making the replies.

Better and worse are the game order, decided by comparison and not by outcome: option AA is better than BB for Left when ABA \ge B, which is a statement about every sum either of them might sit in. Pairs that are equal and pairs that are confused are counted in their own columns, because a rule that is credited with the pairs it cannot distinguish scores well by doing nothing.

Where the ladder goes next

The dominance anchor has five rungs: what deleting removes, what decides how much, what a survivor looks like on a board, how large a mobility gap has to be before it decides a comparison, and now what happens when the gap is measured more carefully.

The rung above is the pairs themselves. Seventy-two pairs have now survived two accounts of why the count fails on them, and they are a small enough set to be looked at rather than summarised: what those boards have in common, whether the option with more replies is the one whose replies are interchangeable, and whether that is a property with a count attached. This page rules out one hypothesis about them and does not offer another, which is the honest state of it.

Two neighbours are worth the trip. A rule that is never right and cannot be far wrong is where a heuristic is priced properly — with a bound it satisfies rather than a score it achieves — and it is the standard this whole ladder is trying to reach. And the moves a player can be talked out of is the other place on this site where a count of dominoes is corrected and where the correction turned out to be an interval; reading the two together shows what a repair looks like when it works and when it does not.

Part 5 of 9

One argument about Dominance. The parts either side of it:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.

ApproximationBoundComparisonConfusedCounterexampleDominanceDomineeringEnumerationHeuristicMobilityNormal playPartizan