Values

The price of taking the maximum

The seventy-two decisions where a Domineering strategy's rules name the wrong placement are never wrong by more than one reply, and a third of them are the second rule's fault rather than the mobility count's. The repair that follows — keep every placement within one reply of the best — retains a best placement every time and costs thirty decisions for every one it saves.

Assumes: Seventy-two of them were not silence · Three rules and a tie-break

Seventy-two of them were not silence split the 181 unanswered decisions in a stored Domineering strategy into two failures that had been counted as one: 109 where the rules say nothing and fall through to a tie-break, and 72 where the rules narrow the candidates to a single placement and the placement is worse than one they discarded. It closed on the number that would say what to do about them:

The question is whether the count is wrong on them by one reply or by more, because a count wrong by one is a tie the rule broke badly and a count wrong by three is the rule being outside its own bound.

It is never wrong by more than one. That is the hoped-for answer, it makes the obvious repair available, and pricing the repair is what this page mostly is — because the repair costs thirty decisions for every one it saves.

Wrong by one, or by nothing. The seventy-two decisions the rules get wrong, by how many replies the named placement misses a best one.
Fig. 1 The seventy-two decisions the rules get wrong, by how many replies the named placement misses a best one. Nothing is wrong by two.

The margin

For each of the 72, take the number of replies the named placement leaves the opponent, and the fewest a best placement leaves. The difference is how far outside its own bound the count is.

Fifty are wrong by one reply. Twenty-two are wrong by nought.

Nothing at all is wrong by two, and that settles the rung below’s question in the direction it wanted. The mobility count is never making a large mistake: on every one of the 72 there is a best placement leaving at most one more reply than the placement the rules chose. The rule is not out of its depth on these positions; it is finishing a photo-finish and getting it wrong.

Two things about that measurement need saying before it is leaned on.

A best placement is one that preserves the outcome, not one that maximises anything. A decision here is a position and a player to move, the candidates are that player’s placements, and a placement is best when playing it keeps the win the position holds. So there can be several best placements and usually are; the rules fail when everything they keep is outside that set.

And the reply count is counted after the placement, not before. Leaves the opponent fewer replies means the count of the opponent’s placements in the position the move produces, which is what makes it a property of the move rather than of the position. That is the quantity the whole ladder has been using since how many moves are worth making, and it is the one being scored here.

The twenty-two are somebody else’s fault

A margin of nought is a strange thing to find in a list of the mobility count’s errors, and it is worth reading slowly.

If a best placement leaves exactly as many replies as the named placement, then the mobility count did not choose between them. It scored them the same, kept them both, and handed both to the next rule. So on those decisions the rules were narrowed to one by connectivity — the second rule, leave the board in fewer pieces — and connectivity is what discarded the best placement.

Whose error it is. The seventy-two split by which of the two rules narrowed the candidates to one.
Fig. 2 The seventy-two split by which of the two rules narrowed the candidates to one. Twenty-four of them are the second rule’s doing.

Twenty-four of the 72 were narrowed by connectivity rather than by mobility. The rung below’s sentence — the mobility count narrowing to one placement and naming the wrong one — is right about forty-eight of them and wrong about the rest, and the correction matters for the repair: a fix inside the mobility count cannot reach a decision the mobility count did not make.

The two figures differ by two, and the difference is instructive rather than an inconsistency. Twenty-two decisions have a margin of nought, and those are exactly the ones where mobility kept a best placement and connectivity threw it away. The other two of the twenty-four are decisions where mobility had already discarded every best placement and connectivity then chose among the survivors — so mobility was wrong first and connectivity merely failed to rescue it.

And connectivity is worth having anyway

Assigning a third of the residue to connectivity would be an argument for dropping it, if the rest of connectivity’s ledger were not so strongly the other way.

What the second rule buys and costs. The two rules scored separately and together over every decision.
Fig. 3 The two rules scored separately and together over every decision a stored strategy has to make.

Mobility alone answers 2,528 of the 3,308 decisions, is wrong on 48, and is silent on 732. Adding connectivity answers 3,034, is wrong on 72, and is silent on 202.

So connectivity converts 530 silences into answers and creates 24 new errors. A rule that buys five hundred and thirty and costs twenty-four is not a rule to drop; it is a rule whose cost has now been measured, which is a different thing from a rule that was thought to be free.

This is the shape of nearly every result on this anchor. Three rules and a tie-break built the list by adding rules that each answered more than they broke, and the list’s 94.5 per cent is the accumulation of those trades. What was missing was the second column: each rule’s own error count, separately, rather than the list’s residue as one lump.

The repair, and what it costs

Wrong by at most one reply has an immediate consequence. If a best placement is always within one reply of the mobility maximum, then a rule that keeps every placement within one reply of the best always retains a best one. The 72 would stop being errors: on every one of them, a best placement would still be among the candidates when the tie-break arrives.

That is exactly the repair the rung below imagined, and it is available. It is also a disaster.

The repair, priced. Mobility relaxed from a maximum to a band, scored over every decision.
Fig. 4 Mobility relaxed from a maximum to a band, scored over every decision. Answering falls from 3,034 to 704.

A rule answers a decision when everything it keeps is a best placement — the standard the whole anchor uses, and the only one under which a player following the rule cannot go wrong. Widening the band retains best placements and retains everything else too, so a decision the strict rule settled by keeping exactly one good placement now keeps two or three of which one is bad.

Answering falls from 3,034 to 704. A band of two replies falls to 336. The repair saves at most 72 decisions and loses 2,330.

Nothing put after it helps. Six tie-breaks applied after a relaxed mobility, none of which recovers the strict rule's score.
Fig. 5 Six tie-breaks applied after the relaxed mobility. The best of them answers 1,160, against 3,034 for taking the maximum outright.

Nor is the loss something a better tie-break recovers. Putting each of six rules after the relaxation — four reading the value of what a placement leaves, one reading the shape, one reading nothing — gets the count back to 1,160 at best. That is a third of what taking the maximum answers. The loss is the band itself, not the absence of a rule after it.

Why a targeted repair is not available either

The obvious answer to a repair that fires too often is to make it fire less often: relax the band only where it is needed.

What a repair would have to touch. How many decisions a repair firing on singletons would reach, against how many it would correct.
Fig. 6 How many decisions a repair firing on singletons would reach, against how many it would correct.

The trouble is that where it is needed is not something the rule can see. The 72 are decisions where the rules name one placement and the placement is wrong; the rules also name one placement, correctly, on 2,658 decisions. From inside the rule the two are indistinguishable — a single surviving candidate, with a reply count and a piece count and nothing else.

So a repair conditioned on the rules named exactly one placement fires on 2,730 decisions and improves 72 of them, one in thirty-eight. Every other firing takes a correct singleton and widens it into a set that may contain a bad placement.

The target is not observable from inside the rule. That is a stronger objection than the arithmetic, because it does not depend on the arithmetic coming out any particular way: any repair that fires on a condition the errors share with the successes will damage the successes in proportion.

A rule cannot be graded on its residue alone

There is a general lesson in the connectivity finding and it is worth separating from Domineering, because it applies to every list of rules this site has built.

A list of rules is normally scored by what it leaves over. That is the natural summary — 3,308 decisions, 3,034 answered, 274 left — and it is the summary the rung below inherited. It hides the fact that a residue is produced jointly: each rule in the list both closes decisions and opens them, and the list’s residue is a net figure with the two flows cancelled against each other.

Separating them changes what a rule looks like. Connectivity, scored on the residue, is responsible for a third of the errors and looks like the weak link. Scored on both flows it converts 530 silences and creates 24 errors, which is a ratio of twenty-two to one. Neither number is wrong and only the second is a description of the rule.

The same arithmetic is available for every rule on every list here, and it has not been done. How much a list can lose prices a list of options against the position it summarises; this is the same question asked of a list of rules, and the answer for the one rule measured here is that its gross contribution is twenty-two times its net cost.

What the seventy-two actually are

Put together, the three measurements say something about what kind of object the residue is.

The rules are not misjudging these positions. They are choosing between placements that differ by nought or one reply — the closest calls the population contains — and on 72 of the 2,730 such calls they choose the wrong one. That is a precision limit rather than a mistake, and it is the same shape as the finding on the neighbouring anchor: a threshold is a detection limit found the mobility rule’s threshold to be a statement about how fine a difference the sweep can resolve rather than a property of a board.

Two rules made of counts have a finest difference they can see, and below it they are guessing. The 72 are the guesses that came out badly, and their being bounded at one reply is the same statement as the rules being right whenever the difference is two.

Read that way, the residue is not a defect at all. It is the price of an argmax: a rule that keeps the single best-scoring candidate is sharp, and sharpness on close calls is exactly the property that makes a rule occasionally wrong and usually decisive. A heuristic that becomes a theorem measures the other end of the same trade, where the calls stop being close and the rule stops being wrong at all.

What the figures cannot show

Six tables, and not one of them shows a position. That is a real gap on a site whose whole method is that the picture carries the argument, and it is worth saying what the missing picture would be.

It would be one of the seventy-two, drawn twice: the region with the placement the rules name, and the region with the placement they discard, side by side, with the reply counts under each and the outcome under that. A reader could then see the thing the numbers only assert — that the two look equally sensible, that the discarded one leaves the board in more pieces, and that the one the rules prefer loses.

That figure is drawable and it is not here, for a reason worth admitting: a single pair would be an anecdote. The finding is about a distribution — that every one of the seventy-two is this close, and that the closeness is what makes them unfixable — and a distribution is a table. What a worked pair would add is the feel of a close call, and what it would risk is a reader generalising from one region to seventy-two. Which shapes are worth fighting over is where the regions themselves get drawn and is the right place to look at one.

What this does not settle

The 72 are not repaired. This page prices two repairs and rejects both; it does not offer a third. What would work, if anything does, is a rule that reads something the two counts do not — a property of the position that distinguishes a correct singleton from a wrong one — and the rung below already established that the value predicts on part of the silent residue. Whether it predicts on this one is untested, and it is the obvious next thing.

The standard is severe and it is the ladder’s. A rule answers only when every candidate it keeps is best. Under a laxer standard — the rule keeps at least one best placement — the relaxed band answers far more, and the number would mean something quite different: not a player following this cannot go wrong but a player following this has a good option available if they can find it. The second is not a stored strategy, because finding it is the work the strategy exists to avoid.

And the population is the strategy’s, not a game’s. All 3,308 decisions come from the compressed strategy for regions of at most eight cells, which is a census of what a player would have to remember rather than of what a player meets. What a strategy has to remember is where that population is built and what it is a census of.

The band is the only relaxation tried. Widening mobility by a fixed number of replies is the natural reading of a tie-break inside mobility, and it is not the only one: a rule could widen by a proportion, or widen only where the count is small, or widen only where the number of candidates is small. Each is a different experiment and none of them is here. What the measurement establishes is that the simplest relaxation is very expensive, and that the expense comes from a mechanism — retaining bad candidates — which every relaxation shares to some degree.

Normal play throughout, and the two players treated symmetrically: a decision is Left’s or Right’s and the rules are the same rules read through the appropriate orientation.

One number is worth carrying out of all of this, because it is the one a reader of the rung below would have guessed wrong. The rules, taken together, make 2,730 decisions in which they name exactly one placement. They are right on 2,658 of them. That is 97.4 per cent, on the hardest sub-population the strategy contains — the decisions where the counts had to separate placements differing by nought or one reply — and it is a much better number than the list’s headline 94.5 per cent, which averages the close calls together with the easy ones and with the silences.

Where the ladder goes next

The tempo anchor has six rungs: what a value leaves out, how many moves are worth making, what a strategy has to remember, three rules and a tie-break, what the residue is made of, and now how far wrong the wrong part of it is.

The rung above is the value on the seventy-two. The rung below found that a rule reading the value of what a placement leaves answers 118 of the 202 genuinely silent decisions where a rule reading only shape answers 93, and it never asked the same question of the 72 — because at the time the 72 were thought to be a mobility error rather than a close call. They are close calls, which is exactly the regime where a different kind of information is most likely to help. Fitting the same value panel to them, one rule per value class, is the same computation on a smaller population and would say whether the residue is close calls the counts cannot resolve or close calls nothing can.

Two neighbours are worth the trip. What a strategy has to remember is where the 3,308 comes from, and every fraction on this page is a fraction of it. And when a real board falls apart says which of these regions a played game actually produces, which decides whether seventy-two errors concentrated on close calls is a small problem or an invisible one.

Part 6 of 7

One argument about Tempo. The parts either side of it:

The objects named here

The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.

ApproximationCompressionCounterexampleDecompositionDomineeringEnumerationHeuristicMobilityPartial orderStrategyTempoValue