A heuristic that becomes a theorem
Assumes: A threshold is a detection limit · The margin a count needs
A threshold is a detection limit took the mobility rule — of two comparable options, the one leaving the opponent fewer replies is the better one — and found that the margin at which it stops failing is not a property of a board at all. The same 3 × 6 gives one answer four squares covered and another six squares covered, and what really moves the rule’s reliability is how far into the game the position is. It closed by naming the object that reading was missing:
Every board here shows the failure rate falling as the board fills, and each is two points; the shape of that fall is the thing worth having, because a rate that falls to nought by the endgame would mean the rule is exact where it matters most.
The rate falls to nought, it gets there before the endgame, and on the way it does something a falling rate usually does not.
The sweep
The measurement is the rung below’s, taken at eight depths instead of two. For a board and a number of covered squares, enumerate every occupancy of that many squares; keep the positions where Left has at least two moves; and for every pair of Left options that are comparable — one of the two games at least as good as the other — record the difference in the number of replies each leaves Right, and whether the option leaving fewer replies is the better one.
Two things about that population are worth restating, because they decide what the rate means.
Only comparable pairs count. Two Domineering options are very often incomparable — neither at least as good as the other in every sum — and a rule of thumb that picks between two incomparable options is not wrong, because there is nothing for it to be wrong about. Scoring the rule on those would flatter or damn it arbitrarily depending on which way the tie-break fell.
And the margin is the whole of the rule’s confidence. A pair where the two options differ by one reply is a near thing; a pair differing by four is a claim the rule is making loudly. So the rate is reported at each margin separately, and the threshold — the smallest margin at which no failure occurs at that depth — is the summary the rung below used.
On a 3 × 5 the margin-one failure rate runs , , , , as the board fills from two squares covered to ten. On a 4 × 4 it runs , , , , , .
Five boards, five curves, one direction. No board’s rule gets worse as the board fills, at any depth, at either margin.
One more decision is buried in the population and it is the one that makes the sweep affordable at eight depths rather than two. The occupancies are every set of that many covered squares, not the ones a game of Domineering could actually produce. That is a larger population than the reachable one and a much cheaper one to enumerate, and the rung below used it for the same reason. It also makes the finding conservative in a specific way: an unreachable occupancy is a harder test than a reachable one, because it can put covered squares in arrangements no pair of dominoes would leave, and every one of those arrangements is a chance for the rule to fail that a real game would never offer.
The fall accelerates
A rate that falls is not news; the rung below had that. What the curve adds is the shape, and the shape is not the one a decaying quantity usually has.
On the 3 × 5 the successive ratios are , , . On the 4 × 4 they are , , , . The rate is not decaying by a constant factor per move. The factor itself is decaying, and monotonically, on both boards long enough to show it.
The consequence for a player is worth stating plainly. If the fall were geometric, the rule would improve steadily and a player would be trading a little accuracy for a little effort at every stage of the game. It is not: the first two moves buy almost nothing, and almost all of the improvement arrives in the last third. So the rule is at its worst exactly where a player most wants a shortcut — an open board, many options, an expensive search — and at its best where the search is cheapest anyway.
That is a discouraging finding read one way and a useful one read another, and the second reading is the one this page ends on.
It reaches exactly nought
The last point of every curve is not a small number. It is nought, and the threshold is one.
A threshold of one says something much stronger than a low failure rate. It says that at this depth, over every pair of comparable Left options on every occupancy of the board, the option leaving Right fewer replies is the better option — with no exception anywhere. On the 4 × 4 that covers 2,872 pairs; on the 3 × 5, 302. It is a guarantee rather than a rate, and it is the first thing on this anchor that is one.
The depths are: a 2 × 6 from four squares covered, a 2 × 7 and a 3 × 4 from six, a 3 × 5 from ten, a 4 × 4 from twelve. In moves remaining that is between two and four. So for the last two to four moves of every board in the sweep, the mobility rule is not a heuristic; it is a theorem about that board at that depth.
A bound instead of an answer is the standard this site holds a rule of thumb to, and the standard is met here in the only way it can be met by a measurement: not by proving the rule, but by identifying exactly the region in which it has no counterexample and saying how thoroughly that region was searched.
And room is not the variable either
The obvious guess about those depths is that they are really one depth in disguise — that the rule becomes exact when the board has few enough empty squares, whatever the board.
It does not. A 2 × 6 becomes exact with eight squares empty and a 3 × 4 — the same twelve squares, differently shaped — with six. The empty count is not the ordering variable any more than the perimeter, the aspect ratio or the longest side were on the rung below.
That is the rung below’s negative result reappearing one level up, and it is worth noticing that it reappears rather than being resolved. The quantity that moves the failure rate is the depth, the depth is not a quantity about the board, and no property of the board so far tried orders the depths at which the boards become exact. What this page has is five curves and five per-board depths, and the honest form of the finding is a table rather than a formula.
Why the rule should improve at all
The measurement is the page and the mechanism is worth a paragraph, because it is not obvious that filling a board should help a rule about reply counts.
A Domineering position late in the game has fallen apart. The board falls apart is the standing fact that a partly-covered board is a sum of small independent regions, and when a real board falls apart measures how quickly that happens in play. Two options on a decomposed board differ in one region and agree everywhere else, so comparing them is comparing two small games rather than two whole boards — and small games are much more likely to be comparable at all, and much more likely to be ordered the way their move counts are.
There is a second mechanism and it pushes the same way. Deep in the game most regions are worth numbers or near-numbers, and a region worth a number is one where the count of moves each side has left is the value. Counting the moves each side has is where that identification is made precise. So as the board fills, the reply count stops being a proxy for the value and starts being the value, and a proxy that has become the thing it was a proxy for cannot be wrong.
Neither argument is a proof, and both predict what the curve shows. What neither predicts is the acceleration, which is the part of the shape this page found and cannot explain.
What the acceleration might be
Nothing on this page explains why the improvement should speed up rather than proceed at a constant rate, and it is worth writing down the candidate that fits, so that the next rung has something to refute.
The two mechanisms above are both threshold effects rather than gradual ones. A board does not decompose a little more with each domino; it decomposes when a covered square happens to cut a region in two, which is an event rather than a trend. And a region does not become a number gradually; it becomes one when its last fight is resolved. So the quantity actually improving is the proportion of pairs whose comparison has been reduced to a settled small region, and that proportion is a product of several such events rather than a sum — which is the ordinary way a fall gets faster than geometric.
If that is right, the acceleration should be visible in a much cheaper statistic than the failure rate: the share of positions at each depth whose board has already fallen into two or more regions. That share is computable without evaluating a single game, and if its curve has the same shape as the failure rate’s, the acceleration is decomposition and nothing else. If it does not, the acceleration is about the values rather than about the geometry, and the mechanism above is wrong.
Three regimes
The useful summary is not a rate and not a formula. It is that the mobility rule has three regimes on any one board, and the curve says where the boundaries are.
Early, with thirteen of fifteen squares empty, it fails once in six on the pairs it is least confident about, and its threshold is three — a player would need a margin of three replies before the rule was safe. In the middle it fails once in twenty-six and the threshold is two. At the end it does not fail, and the threshold is one.
A player who knows only the middle number has a rule of thumb. A player who knows the curve has three different instruments and knows which one is in their hand.
What the tables cannot show
Every figure here is a rate or a count, and the thing a reader would most want to see is a pair: two positions differing by one domino, one of them leaving the opponent fewer replies and being the worse option anyway. That is what a failure is, and it is drawable — the rung below draws several.
What is not drawable is the curve’s subject. A rate at a depth is an average over thousands of pairs, and the pairs are not alike: some are two options in the same region, some are two options in different regions of a decomposed board, some involve an option that splits the board and one that does not. How much a list can lose is where that heterogeneity was first measured on this anchor, and it is the reason a single rate per depth is a summary rather than a description. A figure that showed the distribution behind each rate would be a better figure and would need a different sweep — one that records, for each failing pair, which of those kinds it was.
So the honest reading of these tables is that they say how often and not which. The rung below’s contribution was to establish that the how-often is a function of depth; this page’s is the shape of that function; and the which is not on either page.
What this does not settle
Five boards, all small. The largest is sixteen squares and the sweeps are exhaustive over occupancies rather than over play, so nothing here is a statement about a 6 × 6 board or about positions a game actually reaches. When a real board falls apart is the measurement of the gap between a census of occupancies and the positions play produces, and it is a real gap: an occupancy sweep includes shapes no sequence of dominoes could leave behind.
The exactness is exact on the population and not proved. No failure in 2,872 pairs is a much stronger statement than a rate and it is still a census. What would make it a theorem is an argument that a decomposed board’s comparable options are ordered by reply count, and this page does not have one — it has the two mechanisms above, which are reasons to expect it rather than reasons to believe it.
The acceleration is measured on two boards. Three of the five are too short to have three decay factors, and the check names them rather than counting them as agreeing. Two boards agreeing is a pattern, not a law, and the boards that would test it are the ones the sweep cannot afford.
And the depths are even. Coverage runs in twos because a domino covers two squares, so the curve has a point every move rather than every square, and nothing here says what happens between two points. On a curve that accelerates, a finer grid would matter — the interesting stretch is exactly the one where the rate is dropping fastest.
And a threshold is not a strategy. Knowing that the mobility rule is exact from ten squares covered tells a player which of two comparable options to prefer; it does not tell them which options are comparable, and deciding that is a search rather than a look. What a value leaves out is the standing statement of that gap on this site, and it applies with full force here: the rule prices a comparison that somebody else has to have set up.
Normal play throughout, and Left’s options only. The rule is symmetric in the players by construction, and measuring both sides would double the population without adding a case.
Where the ladder goes next
The dominance anchor has eight rungs: how much a list of options can lose, the margin a count needs, which option a reduction keeps, the weight that blunts the count, the reduction that always shrinks, the threshold as a fact about the census, the threshold as a detection limit, and now the curve it sits on.
The rung above is the exactness itself. The last point of every curve is a threshold of one over thousands of pairs, and a claim that strong on a population that large is either a theorem waiting to be written or a census that has not gone far enough. The way to find out is cheap: take a board past the depth where it becomes exact and check whether every comparable pair on a decomposed board is ordered by its reply count — a statement about sums of small regions rather than about boards, which is a much smaller object than the sweep that produced it. If it holds, the rule stops being a heuristic on this site altogether and becomes a theorem with a hypothesis attached.
Two neighbours are worth the trip. The weight that blunts the count is where the reply count was refined and where refining it turned out to cost more than it bought, and it is worth reading beside a page where the unrefined count turns out to be exact somewhere. And seventy-two of them were not silence is the same rule scored on a different question, where its failures turned out to be a different kind of failure, and the two together are the site’s account of what a mobility count can and cannot decide.
Part 8 of 9
One argument about Dominance. The parts either side of it:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
The objects named here
The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.
ApproximationCounterexampleDecompositionDominated optionDomineeringEnumerationHeuristicIncomparableInvariantMobilityPartial orderValue
- An effect that changes sign approximation, counterexample, decomposition, enumeration, heuristic, invariant, mobility, value
- Half the difference in odd runs approximation, counterexample, decomposition, domineering, enumeration, heuristic, invariant, value
- Three rules and a tie-break approximation, counterexample, decomposition, domineering, enumeration, heuristic, mobility, value
- A catalogue that knows what it will meet approximation, decomposition, domineering, enumeration, heuristic, invariant, value
- One domino every three cells approximation, decomposition, domineering, enumeration, heuristic, invariant, value
- The ceiling was a plateau approximation, counterexample, decomposition, domineering, enumeration, invariant, value