Theme

The thread: Assertions that reject

Every claim here is given a test it could fail, and the tests that matter are the ones that have failed. These are the essays where a check refused something — a guess, a rival explanation, or the essay's own first draft.
What the folding costs to do. The same search over a 4 × 4 Domineering board run twice, once folding positions by symmetry and once not, with everything counted. The fold stores 3.75 times fewer entries and spends 17.5 times more elementary operations to decide where to put them. What it costs

What it costs to notice a repetition

Folding a 4 × 4 Domineering board by its symmetries takes the table from 5,700 entries to 1,522. It also spends 559,424 square-mappings to work out where each entry goes — seventeen and a half times the entire cost of not folding. The saving has a ceiling of four and the price has no ceiling at all, and knowing which currency each is paid in is the difference between an optimisation and a habit.

Every part on its own is worth nothing. The Grundy value of each chain and loop considered as a game by itself, and three real Dots and Boxes boards used to check the turn-by-turn walk against the win-or-lose solver already used here. Every component alone is a second-player win, which is exactly what makes the nim-sum useless. Out in the world

The parts are worth nothing and the sum is not

Every chain and every loop in Nimstring, taken alone, has Grundy value nought. So the Sprague–Grundy theorem predicts that every position built from them is worth nought — and ninety-six of the two hundred and seven positions checked here are not. The theorem is not being misapplied; it does not apply, because a capture keeps the turn. What replaces it is smaller and sharper: count the short chains, and one long component of any kind reverses the parity.

Every position, by how much is left. Sylver Coinage positions counted by genus — the number of integers still unnameable — with the share on which the player to move loses. The parity of the genus very nearly decides the game: odd rows run between a fifth and a half, even rows between nothing and a thirteenth. Out in the world

A parity with a first exception

Sort every Sylver Coinage position by how many numbers are still unnameable and the game very nearly falls to parity: odd rows are between a fifth and a half positions the mover loses, and the first three even rows hold none at all. The rule has a first counterexample at genus eight, where it is a single position out of sixty-seven, and eleven more at genus ten. It is a tendency wearing away from both ends rather than a law with exceptions.

How far a plain search gets. An exhaustive search of Sprouts and Brussels Sprouts, run on this site, with the number of positions each size costs. Sprouts settles at three spots and Brussels Sprouts at two crosses; the published results on Sprouts go to forty-seven. How it was found

What computing further has bought

Sprouts has been searched harder and longer than almost any game, and the period-six pattern has survived every extension. This site's own exhaustive search settles three spots; the published results reach forty-seven, and the gap is not a gap in hardware — the gentler of the two measured growth factors puts forty-seven spots at ten to the hundred and twenty-fifth positions. Beside it sits Brussels Sprouts, which has five million positions holding a choice and not one choice that changes who wins.

What the notation costs to write. Every game born by each day, written in the brace notation and measured. The expressions are all distinct, which is what the notation is for, and by day three the typical one is twenty-two characters and the longest is fifty. How it was found

Where the braces stop

The brace notation names every game exactly — 1,474 games born by day three, 1,474 different expressions, no two alike. It also gets long: the middle one is twenty-two characters and the abbreviations everybody actually writes cover one game in twenty-three. And it has two hard edges. A game with a cycle in it has no finite expression at all, and the equation the minus sign encodes — that a game and its negative cancel — is false under misère play on every one of those 1,474.

Which bit of the rule decides. Four properties of an octal rule table set against whether the game it describes settles into a period. Only one holds on every code that does not: whether a move may leave two non-empty heaps. It is necessary and not sufficient. How it was found

Three bits of rule

An octal code is three bits a digit. The Grundy sequence it determines costs anywhere from one bit to a hundred and thirty-six — a factor of two hundred and seventy-two across rules that differ by a single digit — or it cannot be written down at all. Of four properties of the rule table tested against that, exactly one holds on every code that never settles: whether a move may leave two non-empty heaps. It is necessary, it is not sufficient, and nine codes carry it and produce answers smaller than their own rules.

The order one day out. Whether the values born by day three still form a lattice. Twice as many pairs are incomparable as at day two, and every incomparable pair still has a least upper bound and a greatest lower bound — so the order becomes more tangled without becoming ragged. Values

One of four questions

Three rungs of this ladder rest on sweeps of day two — 22 values, 253 pairs. Day three is 1,474 values and over a million pairs, and only one of the four questions can be asked of it. The order can: twice as many pairs are incomparable and every one of 1,606 sampled still has a least upper bound and a greatest lower bound, none of them a value day two already had. The other three compare sums of day-three values, which are born on day six, and sixty of those exhausted an eight-gigabyte heap.

Two clauses, and what each is about. Four rulesets against the two clauses of the condition. A ruleset passes both or the one-number-per-component recipe fails on it, and the two clauses fail for different reasons: locality is about the state proposed, isolation is about the rule. Where it stops

Two clauses and a third question

A component can carry its own rule when two things hold: its moves are a function of what it carries, and a move in it leaves every other component alone. Two rulesets built to fail one clause each are both caught on a named witness. The four real games sort exactly — every one the recipe gets right fails no clause, every one it gets wrong fails one — and the two clauses still miss something, because Fibonacci Nim and a held pass fail the same clause and only one of them can be repaired.

The third point. Amazons regions on a 4 × 4 board holding one amazon of each colour, by the Chebyshev distance between them. A 3 × 3 board reaches distance two and this one reaches three, and the mean temperature rises across all three. Particular games

The third point on the curve

Four rungs below this one measure a shared Amazons region's temperature against how far apart its two amazons are, and every one does it on a 3 × 3 board, where the distance can only be 1 or 2. Two points make a direction, not a curve. A 4 × 4 board reaches distance 3 — and cannot be evaluated at all until the regions are cut to five free squares. Restricted that far, the answer is neither a sign that flips nor an oscillation: the rise continues and it is running out, the second step being 36 per cent of the first.

A criterion that does not lift. The NoGo separation criterion run on strips and on three-row boards. On strips it holds on nine boards and explains all nine independent ones. On 227 three-row boards it holds on none at all. Particular games

A wall that bends

On a NoGo strip, two empty stretches add when no group breathes into both — a wall of two stones of different colours does it, and the criterion explains nine of ninety-three boards and all nine that it covers. On a three-row board it explains none of 227, and not because it is less accurate. A wall across a board has to bend, a stone at the bend sees empty squares on both sides by itself, and every one of the 227 has a group breathing into both regions. The condition is unsatisfiable.

Where the two diagrams part. Pairs of hot day-two values with the true thermograph of their sum against the one made by adding the components' walls. Every pair parts, and every pair parts at the lower of the two temperatures. Temperature

The second bend is the boundary

Adding two thermographs wall by wall gives a diagram that is right at the mast and wrong below it. Over every pair of hot values born by day two, the added walls sit outside the true ones at every height — an outer envelope with the truth somewhere inside — and the two pictures separate at exactly the lower of the two temperatures, on all twenty-eight pairs. Above that height both components are still fights and the addition is exact; one sixteenth below it, every pair has parted.

Why a square holds fewer values than a line. The same number of squares laid out two ways, with the number of places a hop can start, the longest chain one can run, and the number of distinct values every arrangement of that shape produces. A hop needs three squares in a line, so a long row supplies more of them than a compact rectangle of the same area — and the value counts follow. The second dimension is not the way to reach the deeper values, which is the opposite of what the rung below expected. Out in the world

The second dimension is not the deep end

The rung below says a row of eight reaches every corner of the vocabulary and goes far into none of them, and that the narrowness is a fact about the board. So the obvious next move is a rectangle — and nine squares in a square hold twenty-five values where nine squares in a line hold fifty-eight. The geometry says why before any stone is placed.

The condition has to hold underneath, not on top. Pairs of coin rows sorted by where the incentive condition holds, with Milnor's bound checked on each pair. Rows that satisfy the condition at every subposition never break the bound. Rows that satisfy it only at the top break it on a counted fraction — and a reader who tested the row rather than the row's insides would have called those safe. The distinction is invisible from the position and decides whether the theorem applies to it. Out in the world

A hypothesis has to hold all the way down

Milnor's bound is proved by induction over the play, so the condition it needs has to hold at every position the play can reach. Checked on the row instead, ninety-two pairs pass the test and twenty-four of them break the bound. Checked at every subposition, twenty-eight pairs pass and none breaks it.

Two solutions to one set of equations. The winning condition written as a single predicate and solved twice: once as the least solution of its own equations and once as the greatest. The least says Left can force a win; the greatest says Left cannot be forced to lose, which admits the positions where Left can keep the game going for ever. On a graph with no cycle in it the two coincide and the equations determine an answer. Where they differ, the difference is exactly the set the backward propagation never reaches — so a draw is not a leftover of the algorithm, it is the equations failing to have one answer. How it was found

The gap between two answers

A draw is usually described as what the backward labelling never reached, which makes it sound like a shortfall of the algorithm. Written as one predicate the winning condition is an equation, the equation is monotone, and it has a least solution and a greatest one — and the set the two disagree about is exactly the drawn set, on every game checked.

What a certificate costs, in units of the one Guy and Smith wrote. Octal codes with the period of their Grundy sequence, the window a proof of that period needs, and the arithmetic each costs — counted as mex operations and exclusive-ors, which are the two things a person computing by hand actually performs. Everything is priced in units of the certificate for Dawson's chess, so the column reads as multiples of one hand computation rather than as a number of operations. Some codes cost tens of times as much, and some have no certificate at all. How it was found

What the arithmetic cost in 1956

The rung below ends by respecting a hand computation without pricing it. Priced in the operations a person actually performs, ·137's certificate is 7,919 of them — and the same sweep says ·47's is sixty-three times that, that a splitting move is what makes the cost quadratic, and that seventeen of sixty-four codes have no certificate at any price.

Three complete solutions, each asked about the others. Bouton's 1901 criterion for Nim, Wythoff's 1907 description of his own cold positions, Moore's 1910 rule for taking from several heaps, and the Grundy criterion that arrived thirty years later, each checked against the truth on every position of four games. Every one of the old criteria is exact about its own game and wrong about the others. The blanks matter more than the numbers: Wythoff's is a description of a pair and has no form for three heaps at all, and the Grundy criterion has no form for a game whose moves touch several heaps at once. How it was found

Three complete solutions in nine years

Bouton in 1901, Wythoff in 1907, Moore in 1910 — three airtight solutions of three games, all published before there was any theory of games at all. Asked about each other's games they all fail, and two of them fail by being wrong while one fails by having no form for the question. Only the last kind of failure decides anything.

How far a description of that kind could ever have gone. Subtraction games sorted by whether a Bouton-style column criterion describes their losing positions. His test reads the heap sizes in binary and counts the marks in each column, which works exactly when a heap's value is a function of its own bits — and that is true of a small minority of the family. Below it, the weaker readings: a criterion on the low bits, and a sequence that merely repeats. The method itself is available for every game and says nothing; what 1901 supplied was a set with a description shorter than the game. How it was found

A set with a short description

Bouton's argument is a closure argument about a set, and every impartial game has such a set — its own losing positions. So the method is complete and proves nothing. What made 1901 a theorem is that his set had a description shorter than the game, and swept over fifty-six subtraction games, exactly seven have one of his kind.

The misère sentence, asked of games it was not written for. Bouton's one-sentence solution of misère Nim put to four other impartial games and checked against a search on every position. It is exact on Nim, which is the game it is a theorem about, and wrong on all the others — and wrong in both directions, calling wins losses and losses wins, where the same paper's normal-play criterion errs only one way. The clause responsible is the one about heaps of size one, which is a statement about how many counters are left rather than about what a move can do with them. How it was found

The sentence that solved the other convention

Bouton's paper solves misère Nim too, in one line, and it is the only misère result in the subject that fits on one. Transplanted the way the normal criterion is, it fails differently — the normal one calls losses wins and never the reverse, and this one errs in both directions on every game tried, because the clause it adds is about counters rather than about moves.

All themes