Values

The easy case was not the reason

The rung below found the mobility rule reaching a failure rate of exactly nought near the endgame and named what a proof would need: that a decomposed board's comparable options are ordered by reply count. That statement is false on all five boards, at margins up to two — and split positions go exact two squares of depth before whole ones, so decomposition is the easy case rather than the cause.

Assumes: A heuristic that becomes a theorem · The weight that blunts the count

A heuristic that becomes a theorem swept the mobility rule — of two Left options, the one leaving the opponent fewer replies is the better one — at every depth on five boards, and found its failure rate falling to exactly nought before the endgame. From two to four moves out, on every board in the sweep, the rule has no counterexample at all.

It then said, plainly, what it did not have:

The exactness is exact on the population and not proved. No failure in 2,872 pairs is a much stronger statement than a rate and it is still a census. What would make it a theorem is an argument that a decomposed board’s comparable options are ordered by reply count, and this page does not have one.

The argument is not available, and the reason is the best possible one: the statement is false.

The theorem a proof would have needed. The mobility rule's failures on decomposed positions against connected ones, across every board in the depth sweep.
Fig. 1 The proposed theorem tested on every board in the depth sweep. A decomposed position’s comparable options are misordered by the reply count 1,762 times in 70,296 pairs, on all five boards, at margins up to two. The figure refuses to draw unless the theorem fails on every board and unless the failures reach past a margin of one — a rule wrong only ever by a single reply would be a rule needing a tolerance, not a false one.

Decomposition is nevertheless worth having. The split positions fail at 2.5 per cent against the whole positions’ 7.2, which is a factor of nearly three. It is a real effect and it is not an exemption, and the difference between those two things is the whole of this page.

What was being proposed

The proposed mechanism is a good one and it is worth stating at its strongest before it is tested.

A decomposed position is a sum: two or more regions with no square in common and no move touching both. The board falls apart is where that is established and priced, and the reason it makes a rule about counting replies plausible is that a sum’s arithmetic is additive in a way a single region’s is not. Left’s reply count on a sum is the sum of the reply counts, and Left’s value on a sum is the sum of the values. Two quantities that both add ought to be easier to relate than two that do not.

More sharply: on a sum where Left moves inside one component, the other components are untouched. So comparing two Left options that both play in the same component is comparing two positions that agree everywhere else — which is exactly the situation in which a local statistic like a reply count has the best chance of deciding a global comparison.

That argument is not wrong about anything. It is a reason to expect the rule to do better on decomposed positions, and it does. It is not a reason to expect it to be exact, and the gap between those is where the mechanism fails.

One board, split two ways. The failure rate at each depth on a single board, with decomposed and connected positions counted separately.
Fig. 2 The four-by-four at every depth, with the pairs separated by whether the position had already fallen into more than one region. Both rates fall, the split rate falls faster, and neither reaches nought at the depths where most positions are split. At ten empty squares 71 per cent of positions have already decomposed and the split pairs still carry 796 misorderings.

What a comparable pair is, and why so few of them there are

The population all of this is a rate on deserves a sentence, because comparable pair is doing more work than it looks.

Two of Left’s options AA and BB are comparable when one of them is at least as good as the other in every context — ABA \geq B or BAB \geq A — which is decided by playing the difference game ABA - B and asking who wins moving second. Most pairs are not comparable. Two options that each do something the other does not are confused, and no rule of any kind can order them, because there is no order to get right; confused is not the same as unknown is where that fourth relation is established as a fact about the pair rather than a shrug.

So the sweep is not scoring the rule on every choice a player faces. It is scoring it on the subset where a correct answer exists, which is the only place a rule can be right or wrong. On a 4 × 4 with ten squares still empty that subset is 37,078 pairs out of a much larger number of option pairs, and the rule gets 1,870 of them backwards.

That framing matters for reading the rates. Seven per cent is not the rule is wrong seven times in a hundred moves; it is of the choices where one option is genuinely better, the rule picks the worse one seven times in a hundred. The confused pairs — where a player still has to choose — are outside the measurement entirely and the rule says something about them too, which nothing here scores. Seventy-two of them were not silence is the page that scores a mobility rule on that other question, and it is a different measurement with a different answer.

Where splitting makes it worse

If decomposition were what the exactness is made of, a split position could never be worse than a whole one. It is, at several depths.

Where splitting makes it worse. The depths at which decomposed positions carry a higher failure rate than connected ones.
Fig. 3 Every depth at which a split position is misordered more often than a whole one. They are at the shallow end, where a board has just come apart and the regions it came apart into are still large — a two-by-seven at ten empty squares has 71 per cent of its positions split and the split ones fail at four per cent against the whole ones’ three.

There are several such depths and they are all early. That is the shape to expect if decomposition is a proxy for something else rather than a cause: a board that has split at two squares covered has split into large pieces, and a large piece is exactly the object the rule is bad at. Splitting a twelve-square board into a seven and a five has not made either half easy; it has made two hard problems out of one.

So the rows in that figure are not noise. They are the mechanism’s own prediction failing in the direction it cannot afford — a sufficient condition that is sometimes worse than its negation is not a sufficient condition.

Which class sets the depth

The sharper measurement is not the rate but the depth: at what point does each class of position stop failing altogether?

Split positions go exact first. The depth at which decomposed and connected positions each stop failing, against the depth the board as a whole becomes exact.
Fig. 4 The depth at which each class stops carrying a failure, against the depth the board as a whole becomes exact. On the two largest boards the split positions are clean two squares of depth before the whole ones, and the board’s own exactness waits for the whole ones. So the class the mechanism nominated as the explanation is the class that was never the obstacle.

On the four-by-four, decomposed positions carry no failure from ten squares covered onward; connected positions carry eight failures at that depth and the board does not become exact until twelve. On the three-by-five the same thing happens at the same distance: split positions clean at eight, whole positions still failing, exactness at ten.

So the split positions are not what the exactness is waiting for. They arrive early and then wait; what sets each board’s depth is the last few connected positions, which are single regions of four to six squares with two comparable options and a reply count that gets them the wrong way round.

On the 2 × 6 and the 3 × 4 the two classes go clean together, and on the 2 × 7 the split positions are actually the last to come right — at four squares covered the split pairs carry 72 misorderings against the whole pairs’ 16. So the lead is not universal; it is the pattern on the two boards with enough room for the question to have an answer, and the smallest boards say nothing either way because they have almost no shallow depths to say it at.

That inverts the proposed mechanism completely. The rung below reached for decomposition because decomposition is where a local statistic ought to work, and it does work there — first, and by a factor of three. It is the other class, the one the mechanism had nothing to say about, that decides when the rule becomes a theorem.

A cause that is lower where the effect is

There is one more measurement, and it is the one that closes the question rather than merely weakening it.

The share is not highest where the rule is exact. The proportion of positions already decomposed, at its peak and at the depth the mobility rule becomes exact.
Fig. 5 The proportion of positions that have already split, at its peak and at the depth the rule becomes exact. On three of the five boards the exactness arrives at a depth where fewer positions are split than at depths where the rule was still failing — the four-by-four peaks at 87 per cent and becomes exact at 60.

The share of decomposed positions is not monotone in depth. It rises as the board fills and then falls again, because a board with only four empty squares left usually has them in one place. On the four-by-four the share peaks at 87 per cent with ten squares covered — where the rule still fails — and stands at 60 per cent at twelve, where it is exact.

A quantity that is lower where the effect appears than where it does not is not the cause of the effect. That is not a statistical argument requiring a population; it is a single comparison on one board, and it is decisive on its own.

What is left

What is left of the mechanism. The proposed explanation for the mobility rule's exactness, split into its claims and scored.
Fig. 6 The proposed mechanism split into the claims it makes. The helpful half survives and the explanatory half does not, which leaves the rung below’s exactness where it was — measured on a large population, at a per-board depth nothing predicts, and unproved.

The rung below’s finding is untouched: the mobility rule really does reach a failure rate of exactly nought, and the census behind that is large. What has been removed is the one candidate explanation it offered, and nothing has replaced it.

That leaves the anchor in a specific and slightly uncomfortable position, which is worth naming rather than smoothing over. The exactness is a fact about depth, the depth is different on every board, and the threshold was a fact about the census already established that no property of the board — squares, perimeter, aspect ratio, longest side — orders those depths. Now decomposition joins that list. The quantity that governs when this rule becomes exact remains unidentified, and the set of things it is not has grown by one.

What the sweep does license is a sentence about where the rule is bad, and it is more useful to a player than the failed theorem would have been. The rule is at its worst on a large single region with room on it: 7.2 per cent on whole positions against 2.5 on split ones, and worse still at shallow depths. That is the same conclusion the weight that blunts the count reached from the other direction, and it is the standing complaint about mobility counting on this site — the rule is unreliable exactly where a search is expensive, and reliable where a search is cheap.

Two things this does not disturb

A negative result about a mechanism is easy to over-read, so it is worth being explicit about what still stands.

The rule is still worth using. A failure rate of 7.2 per cent on whole positions is a rule that gets fourteen comparable choices in fifteen right, at the cost of counting the opponent’s placements — against a search over a game tree. A bound instead of an answer is this site’s account of how to hold a rule of thumb to a standard, and the standard is know exactly where it fails and how thoroughly it was searched, which this anchor now meets better than before rather than worse.

And decomposition is still worth having, for its own reasons. The board falls apart prices what a split buys a solver — the components are searched separately and the costs multiply out rather than up — and none of that is affected by a mobility rule’s behaviour. The factor of three found here is a bonus on top of a saving that was already the largest single effect on this side of the site.

What has gone is one sentence: splitting is why the rule becomes exact. It was a plausible sentence, it was offered as a conjecture rather than a claim, and it took one sweep already built to refute. That is the cheapest kind of progress there is and it is the reason the rung below wrote the sentence down instead of leaving the exactness unexplained and unexamined.

What the solver computed, and how

Five boards — a 2 × 6, a 2 × 7, a 3 × 4, a 3 × 5 and a 4 × 4 — at every even number of covered squares from none up to four short of full. At each depth every occupancy of exactly that many squares is enumerated by combination, and a position is kept when Left has at least two legal placements.

For each kept position the free squares are flood-filled to count connected regions, and the position is labelled split when there is more than one and whole when there is one. That label is taken from the position being moved from, not from what a move produces, because the proposed theorem is a statement about a board that has already fallen apart.

Every pair of Left’s options is then compared. The two resulting games are evaluated by the ordinary Domineering recursion and reduced to canonical form, and the pair is counted only when the two are genuinely comparable — one at least as good as the other, and not equal — and their reply counts differ. The margin is the difference in reply counts and a failure is a pair in which the option leaving Right fewer replies is the worse one. The counts are tallied separately for the split and whole labels and never pooled.

Three things are asserted rather than reported. The proposed theorem must fail on every board, since a page written about a false statement needs it to be false everywhere rather than on average. The worst decomposed failure must reach a margin of at least two, because a rule wrong only ever by a single reply would be a rule wanting a tolerance rather than a false one. And the decomposed rate must be below the connected rate overall, or the positive half of the page is empty.

Where the model stops

Five boards to sixteen squares, which is the same sweep the rung below ran and is the reason the two are comparable. Nothing here reaches a board large enough for a decomposition into three or more substantial regions to be common, and the mechanism’s best case is presumably a sum of many small parts rather than a sum of two large ones. So the honest form of the negative is: decomposition into two pieces, on boards of this size, is not what the exactness is made of.

Left’s options only. The rule is stated for the player to move and the sweep runs it for Left throughout. Domineering is not symmetric — Left places vertically and Right horizontally — but the boards include both a 2 × 7 and its transpose in effect, and nothing suggests the asymmetry matters here.

Normal play, and the comparison the whole page rests on is a normal-play comparison: ABA \geq B means Left wins ABA - B moving second, which is decided by a search over the difference game.

And the figures cannot show the object that would settle the question, which is a witness. Every table here counts misordered pairs; none exhibits a decomposed position with two comparable options where the one leaving fewer replies is worse, drawn, with both regions visible and both reply counts marked. Such a position exists — there are 1,762 of them — and looking at a few would say whether they have a shape in common. That is the next rung and it is a different kind of work: reading counterexamples rather than counting them.

Where the ladder goes next

The dominance anchor has nine rungs: how much a list of options can lose, the margin a count needs, which option a reduction keeps, the weight that blunts the count, the reduction that always shrinks, the threshold as a fact about the census, the threshold as a detection limit, the curve it sits on, and now what the curve is not made of.

The rung above is the counterexamples themselves. The 1,762 decomposed misorderings are the sharpest population this anchor has ever had — every one is a case where an additive statistic gets an additive comparison wrong, on a position where the mechanism says it should not — and not one of them has been looked at. The measurement is cheap because the sweep already finds them: for each, record how the free squares were divided, which component the two options played in, and whether the two options played in the same component or different ones. The mechanism’s argument only ever covered the same-component case, so if the counterexamples are concentrated in the different-component case, the mechanism was right about what it claimed and was applied too widely — which would rescue a weaker version of it and would be the best outcome available.

Two neighbours are worth the trip. The threshold was a fact about the census is where the list of board properties that fail to order this rule’s behaviour was started, and decomposition joins it here. And how often a board falls apart is where the share of decomposed positions was measured for its own sake, and it is the page whose curve explains why the share peaks and falls — which is what makes the comparison in this page’s fifth figure possible at all.

Part 9 of 9

One argument about Dominance. The parts either side of it:

The objects named here

The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.

Canonical formComparisonDecompositionDisjunctive sumDominanceDomineeringEnumerationExhaustive searchNormal playStrategy