The easy case was not the reason
Assumes: A heuristic that becomes a theorem · The weight that blunts the count
A heuristic that becomes a theorem swept the mobility rule — of two Left options, the one leaving the opponent fewer replies is the better one — at every depth on five boards, and found its failure rate falling to exactly nought before the endgame. From two to four moves out, on every board in the sweep, the rule has no counterexample at all.
It then said, plainly, what it did not have:
The exactness is exact on the population and not proved. No failure in 2,872 pairs is a much stronger statement than a rate and it is still a census. What would make it a theorem is an argument that a decomposed board’s comparable options are ordered by reply count, and this page does not have one.
The argument is not available, and the reason is the best possible one: the statement is false.
Decomposition is nevertheless worth having. The split positions fail at 2.5 per cent against the whole positions’ 7.2, which is a factor of nearly three. It is a real effect and it is not an exemption, and the difference between those two things is the whole of this page.
What was being proposed
The proposed mechanism is a good one and it is worth stating at its strongest before it is tested.
A decomposed position is a sum: two or more regions with no square in common and no move touching both. The board falls apart is where that is established and priced, and the reason it makes a rule about counting replies plausible is that a sum’s arithmetic is additive in a way a single region’s is not. Left’s reply count on a sum is the sum of the reply counts, and Left’s value on a sum is the sum of the values. Two quantities that both add ought to be easier to relate than two that do not.
More sharply: on a sum where Left moves inside one component, the other components are untouched. So comparing two Left options that both play in the same component is comparing two positions that agree everywhere else — which is exactly the situation in which a local statistic like a reply count has the best chance of deciding a global comparison.
That argument is not wrong about anything. It is a reason to expect the rule to do better on decomposed positions, and it does. It is not a reason to expect it to be exact, and the gap between those is where the mechanism fails.
What a comparable pair is, and why so few of them there are
The population all of this is a rate on deserves a sentence, because comparable pair is doing more work than it looks.
Two of Left’s options and are comparable when one of them is at least as good as the other in every context — or — which is decided by playing the difference game and asking who wins moving second. Most pairs are not comparable. Two options that each do something the other does not are confused, and no rule of any kind can order them, because there is no order to get right; confused is not the same as unknown is where that fourth relation is established as a fact about the pair rather than a shrug.
So the sweep is not scoring the rule on every choice a player faces. It is scoring it on the subset where a correct answer exists, which is the only place a rule can be right or wrong. On a 4 × 4 with ten squares still empty that subset is 37,078 pairs out of a much larger number of option pairs, and the rule gets 1,870 of them backwards.
That framing matters for reading the rates. Seven per cent is not the rule is wrong seven times in a hundred moves; it is of the choices where one option is genuinely better, the rule picks the worse one seven times in a hundred. The confused pairs — where a player still has to choose — are outside the measurement entirely and the rule says something about them too, which nothing here scores. Seventy-two of them were not silence is the page that scores a mobility rule on that other question, and it is a different measurement with a different answer.
Where splitting makes it worse
If decomposition were what the exactness is made of, a split position could never be worse than a whole one. It is, at several depths.
There are several such depths and they are all early. That is the shape to expect if decomposition is a proxy for something else rather than a cause: a board that has split at two squares covered has split into large pieces, and a large piece is exactly the object the rule is bad at. Splitting a twelve-square board into a seven and a five has not made either half easy; it has made two hard problems out of one.
So the rows in that figure are not noise. They are the mechanism’s own prediction failing in the direction it cannot afford — a sufficient condition that is sometimes worse than its negation is not a sufficient condition.
Which class sets the depth
The sharper measurement is not the rate but the depth: at what point does each class of position stop failing altogether?
On the four-by-four, decomposed positions carry no failure from ten squares covered onward; connected positions carry eight failures at that depth and the board does not become exact until twelve. On the three-by-five the same thing happens at the same distance: split positions clean at eight, whole positions still failing, exactness at ten.
So the split positions are not what the exactness is waiting for. They arrive early and then wait; what sets each board’s depth is the last few connected positions, which are single regions of four to six squares with two comparable options and a reply count that gets them the wrong way round.
On the 2 × 6 and the 3 × 4 the two classes go clean together, and on the 2 × 7 the split positions are actually the last to come right — at four squares covered the split pairs carry 72 misorderings against the whole pairs’ 16. So the lead is not universal; it is the pattern on the two boards with enough room for the question to have an answer, and the smallest boards say nothing either way because they have almost no shallow depths to say it at.
That inverts the proposed mechanism completely. The rung below reached for decomposition because decomposition is where a local statistic ought to work, and it does work there — first, and by a factor of three. It is the other class, the one the mechanism had nothing to say about, that decides when the rule becomes a theorem.
A cause that is lower where the effect is
There is one more measurement, and it is the one that closes the question rather than merely weakening it.
The share of decomposed positions is not monotone in depth. It rises as the board fills and then falls again, because a board with only four empty squares left usually has them in one place. On the four-by-four the share peaks at 87 per cent with ten squares covered — where the rule still fails — and stands at 60 per cent at twelve, where it is exact.
A quantity that is lower where the effect appears than where it does not is not the cause of the effect. That is not a statistical argument requiring a population; it is a single comparison on one board, and it is decisive on its own.
What is left
The rung below’s finding is untouched: the mobility rule really does reach a failure rate of exactly nought, and the census behind that is large. What has been removed is the one candidate explanation it offered, and nothing has replaced it.
That leaves the anchor in a specific and slightly uncomfortable position, which is worth naming rather than smoothing over. The exactness is a fact about depth, the depth is different on every board, and the threshold was a fact about the census already established that no property of the board — squares, perimeter, aspect ratio, longest side — orders those depths. Now decomposition joins that list. The quantity that governs when this rule becomes exact remains unidentified, and the set of things it is not has grown by one.
What the sweep does license is a sentence about where the rule is bad, and it is more useful to a player than the failed theorem would have been. The rule is at its worst on a large single region with room on it: 7.2 per cent on whole positions against 2.5 on split ones, and worse still at shallow depths. That is the same conclusion the weight that blunts the count reached from the other direction, and it is the standing complaint about mobility counting on this site — the rule is unreliable exactly where a search is expensive, and reliable where a search is cheap.
Two things this does not disturb
A negative result about a mechanism is easy to over-read, so it is worth being explicit about what still stands.
The rule is still worth using. A failure rate of 7.2 per cent on whole positions is a rule that gets fourteen comparable choices in fifteen right, at the cost of counting the opponent’s placements — against a search over a game tree. A bound instead of an answer is this site’s account of how to hold a rule of thumb to a standard, and the standard is know exactly where it fails and how thoroughly it was searched, which this anchor now meets better than before rather than worse.
And decomposition is still worth having, for its own reasons. The board falls apart prices what a split buys a solver — the components are searched separately and the costs multiply out rather than up — and none of that is affected by a mobility rule’s behaviour. The factor of three found here is a bonus on top of a saving that was already the largest single effect on this side of the site.
What has gone is one sentence: splitting is why the rule becomes exact. It was a plausible sentence, it was offered as a conjecture rather than a claim, and it took one sweep already built to refute. That is the cheapest kind of progress there is and it is the reason the rung below wrote the sentence down instead of leaving the exactness unexplained and unexamined.
What the solver computed, and how
Five boards — a 2 × 6, a 2 × 7, a 3 × 4, a 3 × 5 and a 4 × 4 — at every even number of covered squares from none up to four short of full. At each depth every occupancy of exactly that many squares is enumerated by combination, and a position is kept when Left has at least two legal placements.
For each kept position the free squares are flood-filled to count connected regions, and the position is labelled split when there is more than one and whole when there is one. That label is taken from the position being moved from, not from what a move produces, because the proposed theorem is a statement about a board that has already fallen apart.
Every pair of Left’s options is then compared. The two resulting games are evaluated by the ordinary Domineering recursion and reduced to canonical form, and the pair is counted only when the two are genuinely comparable — one at least as good as the other, and not equal — and their reply counts differ. The margin is the difference in reply counts and a failure is a pair in which the option leaving Right fewer replies is the worse one. The counts are tallied separately for the split and whole labels and never pooled.
Three things are asserted rather than reported. The proposed theorem must fail on every board, since a page written about a false statement needs it to be false everywhere rather than on average. The worst decomposed failure must reach a margin of at least two, because a rule wrong only ever by a single reply would be a rule wanting a tolerance rather than a false one. And the decomposed rate must be below the connected rate overall, or the positive half of the page is empty.
Where the model stops
Five boards to sixteen squares, which is the same sweep the rung below ran and is the reason the two are comparable. Nothing here reaches a board large enough for a decomposition into three or more substantial regions to be common, and the mechanism’s best case is presumably a sum of many small parts rather than a sum of two large ones. So the honest form of the negative is: decomposition into two pieces, on boards of this size, is not what the exactness is made of.
Left’s options only. The rule is stated for the player to move and the sweep runs it for Left throughout. Domineering is not symmetric — Left places vertically and Right horizontally — but the boards include both a 2 × 7 and its transpose in effect, and nothing suggests the asymmetry matters here.
Normal play, and the comparison the whole page rests on is a normal-play comparison: means Left wins moving second, which is decided by a search over the difference game.
And the figures cannot show the object that would settle the question, which is a witness. Every table here counts misordered pairs; none exhibits a decomposed position with two comparable options where the one leaving fewer replies is worse, drawn, with both regions visible and both reply counts marked. Such a position exists — there are 1,762 of them — and looking at a few would say whether they have a shape in common. That is the next rung and it is a different kind of work: reading counterexamples rather than counting them.
Where the ladder goes next
The dominance anchor has nine rungs: how much a list of options can lose, the margin a count needs, which option a reduction keeps, the weight that blunts the count, the reduction that always shrinks, the threshold as a fact about the census, the threshold as a detection limit, the curve it sits on, and now what the curve is not made of.
The rung above is the counterexamples themselves. The 1,762 decomposed misorderings are the sharpest population this anchor has ever had — every one is a case where an additive statistic gets an additive comparison wrong, on a position where the mechanism says it should not — and not one of them has been looked at. The measurement is cheap because the sweep already finds them: for each, record how the free squares were divided, which component the two options played in, and whether the two options played in the same component or different ones. The mechanism’s argument only ever covered the same-component case, so if the counterexamples are concentrated in the different-component case, the mechanism was right about what it claimed and was applied too widely — which would rescue a weaker version of it and would be the best outcome available.
Two neighbours are worth the trip. The threshold was a fact about the census is where the list of board properties that fail to order this rule’s behaviour was started, and decomposition joins it here. And how often a board falls apart is where the share of decomposed positions was measured for its own sake, and it is the page whose curve explains why the share peaks and falls — which is what makes the comparison in this page’s fifth figure possible at all.
Part 9 of 9
One argument about Dominance. The parts either side of it:
The objects named here
The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.
Canonical formComparisonDecompositionDisjunctive sumDominanceDomineeringEnumerationExhaustive searchNormal playStrategy
- One number, stated two ways canonical form, comparison, decomposition, disjunctive sum, domineering, enumeration, normal play
- A factor, and not an overhead canonical form, comparison, disjunctive sum, dominance, enumeration, normal play
- Close calls nothing resolves canonical form, decomposition, domineering, enumeration, normal play, strategy
- The table that changes its mind decomposition, disjunctive sum, domineering, enumeration, normal play, strategy
- Three rules and a tie-break canonical form, decomposition, domineering, enumeration, normal play, strategy
- What a strategy has to remember canonical form, decomposition, domineering, enumeration, exhaustive search, strategy