Out in the world

Nothing to subtract with

Comparison is defined by contexts and computed by subtraction, and the equivalence between the two is a theorem about groups. A scoring game is not one — sixty-six of eighty-one coin rows do not cancel against their own negatives — and the difference test then fails on a row compared with itself, which every context accepts and nothing certifies.

Assumes: A hypothesis has to hold all the way down · Comparing positions

Three rungs below, the essay that opens this ladder names what a scoring game loses and puts it in one word:

The specific thing that goes when the score goes is the group structure. In normal play, positions form an additive group: every game has a negative, sums compose, and comparison is subtraction. That is what makes the whole apparatus work and it is a consequence of the last-move convention.

That is stated and it is not measured, and it is the sort of claim that ought to be. This rung measures it, and the number that comes back is larger than the sentence suggests.

The equation, and the strategy that proves it

Under normal play, G + (−G) = 0 is the most general fact in the subject. Its proof is a sentence: whatever one player does in one copy, the other answers with the mirror move in the other copy; the answerer therefore always has a reply; and the answerer never runs out first, so the second player wins, so the sum is worth nothing.

Every game has a negative is the general statement, and it is what makes comparison computable — G ≥ H is decided by asking one question about G − H, which is one search instead of a quantifier over every position in the world.

The mirroring strategy is still available in a scoring game. Nothing about it depends on the convention: the moves in −G are the moves in G with the players swapped, so a copy is always available and the answerer always has one.

It does not cancel, and the reason is one line. What is being added up is coins rather than moves. The mirror of a concession is another concession, and two mirror-image coins are not nothing; they are twice something.

A game beside its own mirror, and what is left over. Every coin row added to its own negative, played out exactly, with the resulting scores counted. Under the last-move convention every such sum is worth nothing, because the mirroring strategy guarantees the second player the last move. Here the same strategy is available and the score it produces is not nothing: the mirror of a coin conceded is another coin conceded. Gold is the sums that do come to nothing, which are a minority.
Fig. 1 Every row of four coins from {1, 2, 3} beside its own negative, played out exactly. Fifteen of eighty-one come to nothing. The other sixty-six do not, and the scores run from eight ahead to eight behind.

What a negative is, for a game with a score

The definition needs care, because the obvious reading is wrong and the wrong reading gives a wrong answer.

−G is the game with the players swapped. For a coin row both players have the same moves, so swapping them changes nothing about the moves — what it changes is who the coins are worth something to. So as a summand, the negative of a row is the same row with its coins negated: a coin worth v to Left in one copy is worth v to Right in the other.

That makes the sum an ordinary two-row position the same solver handles, with the second row’s coins carrying the opposite sign, and the number it returns is the measurement.

Fifteen of eighty-one rows cancel. Sixty-six do not, and the failures are not small: the worst is a row worth eight to whoever moves first, which is more than the row’s own total.

The fifteen that do cancel are worth a look, because they are not an accident. They are the rows whose two parity classes have equal sums — the rows where the strategy the rung three below is built on, taking every coin of one parity, gives both players the same haul. That strategy is what a mirroring argument reduces to here, and where it happens to balance, the sum comes to nothing.

The distribution, and what it says about the mechanism

The scores of G + (−G) over the eighty-one rows run from minus eight to plus eight in steps of two, and the shape of that distribution is worth reading rather than the extremes.

The steps are of two because the total of a row and its negative is nought, and a score is that total less twice one player’s haul — so every score has the same parity, which the figure asserts rather than assumes. The spread is symmetric, because the family is closed under negation and negating a row exchanges the two ends of the distribution.

And the mass is not at nought. Fifteen rows land there, which is under a fifth; the rest are spread across the other eight values, with the largest coins producing the largest departures. A reader expecting “mostly cancels, sometimes not” has the picture backwards.

That distribution is the mechanism showing through. If the failure were an occasional accident of a badly-shaped row, the mass would pile up at nought with a tail. It piles up nowhere, because the mirroring strategy is not nearly neutralising the opponent — it is systematically paying twice, and how much it pays depends on the coins rather than on anything subtle.

A game beside its own mirror, and what is left over. Every coin row added to its own negative, played out exactly, with the resulting scores counted. Under the last-move convention every such sum is worth nothing, because the mirroring strategy guarantees the second player the last move. Here the same strategy is available and the score it produces is not nothing: the mirror of a coin conceded is another coin conceded. Gold is the sums that do come to nothing, which are a minority.
Fig. 2 The same measurement over a different set of coins, which moves every number and no proportion. Larger coins, larger departures, the same shape — which is what says the finding is about the strategy rather than about the family.

The consequence, and it is worse than “no inverses”

Losing inverses is usually described as losing a convenience. It is not, because the convenience is the whole of how comparison is computed.

The definition of G ≥ H is contextual: Left does at least as well with G as with H, whatever else is on the board. That is a quantifier over every position there is, and nobody can evaluate it directly.

The reason anybody can compute a comparison at all is that in a group the quantifier collapses. G ≥ H if and only if G − H ≥ 0, and ≥ 0 is a single test on a single game. Comparison is a search and the search is finite because the group law made it one.

With no inverses that collapse is unavailable, and the honest thing is to check both tests rather than assume the collapse survives in weakened form.

Comparison stops being a subtraction. Every ordered pair of coin rows compared two ways: by asking whether one does at least as well as the other in each of a family of contexts, and by asking whether their difference is not worse than nothing. Under the last-move convention the two tests are equivalent, which is what makes comparison computable. Here they are not, and the sharpest case is a row against itself — the contexts all say yes and the difference test says no, because a game beside its own negative does not come to nothing.
Fig. 3 Every ordered pair of three-coin rows, compared both ways: by asking whether one does at least as well in each of a family of contexts, and by asking whether their difference is not worse than nothing. Five hundred and seventy-three of seven hundred and twenty-nine agree. The rest do not, and the direction they fail in is the interesting one.

The sharpest case is a game against itself

Of the seven hundred and twenty-nine ordered pairs, five are pairs the difference test accepts and the contexts refuse — the test is unsound on those, claiming an inequality that does not hold.

A hundred and fifty-one are the other way: the contexts accept and the difference test refuses. The test is incomplete there.

And among the hundred and fifty-one are the twenty-seven pairs of a row with itself.

That is the sentence to carry. The difference test cannot certify a game against itself. Every context accepts it, trivially and by definition — a row does exactly as well as itself in every context there is — and the test says no, because G + (−G) is not nothing and the test asks whether it is.

There is no repair to that by tuning a threshold. The test is asking the wrong question: it asks whether a particular sum is not worse than nothing, and in a group that question is equivalent to the one anybody wanted, and here it is not equivalent to anything.

A game where the last move decides nothing. Rows of coins taken from either end, with the exact score for each side moving first. Under the normal-play convention this family is settled entirely by the parity of the row — nobody is ever without a move until the coins run out — so normal-play theory returns the same answer for every row and it is not the answer anybody wants. The scoring answer depends on nothing but the numbers.
Fig. 4 The row that fails hardest, and its negative. Played together, the first player takes eight more than the second. Under the last-move convention this position is worth nothing and is a second-player win; here the same mirroring strategy is available and produces a rout.

Why the mirroring strategy stops working, exactly

It is worth being precise about which step fails, because the strategy is not merely weakened — it inverts, in the same way it inverts under misère play.

Under normal play, the answerer’s guarantee is that a reply always exists. That is a statement about existence, it is what the mirror supplies, and running out of replies is the only way to lose. So the guarantee is the win.

Under scoring, having a reply guarantees nothing about the score. The answerer replies to every move and the score is the sum of what both players took, and the mirror ensures the two haul the same shape of coins from opposite copies rather than the same value of coins. A player who takes a three in the first copy is answered by a player taking a minus-three in the second, which is a three for the first player again.

So the copier is not neutralising the opponent; the copier is paying twice. That is why the sums come out at eight rather than at nought, and why the failures are largest on the rows with the largest coins in them.

Misère play breaks the same equation by a third route: the mirror strategy still guarantees a reply, and under misère having a reply is what a player is trying to avoid, so the strategy that guaranteed the win now guarantees the loss. Three conventions, one strategy, three different fates — and the strategy is the same in all three.

Where the five unsound pairs come from

A hundred and fifty-one refusals are explained by the identity failing. Five acceptances that should have been refusals are a separate defect and are worth accounting for, because a sound-in-one-direction test is a much more useful object than a test that errs both ways.

The five are pairs where the difference G + (−H) happens to come out non-negative from both sides while some context separates the two rows. That is possible because the difference test asks about one position and the contextual test asks about several, and a single position can be silent about a distinction that a well-chosen companion exposes.

That is the same failure the outcome table has one field over. Two positions with identical outcomes can sit inside sums with different outcomes, which is the reason the outcome table was abandoned as a notation; here it is the difference test that is being too coarse, for the same reason and one level up.

What makes the count small is that the difference is already a fairly sensitive probe — it is a whole game rather than a label — so most distinctions survive into it. Five of seven hundred and twenty-nine is the residue.

What survives, and it is not nothing

A reader could take from this that comparison is simply unavailable in a scoring game, and that is too strong.

The contextual relation is still a relation. It is reflexive — a row does as well as itself in every context — and it is transitive, since the inequalities compose context by context. So there is a partial order on scoring games; what is missing is a way to compute it. That is the same object equality in every company is about, and the same object the normal-play theory reaches by subtraction rather than by quantifying.

And the difference test is not useless, it is one-sided. It refuses far more than it should and accepts almost nothing it should not: five unsound pairs against a hundred and fifty-one incomplete ones. A test that errs overwhelmingly in one direction is a sufficient condition, and a sufficient condition that can be computed is worth having even when the necessary one cannot.

That is the same shape the transplanted criterion has one field over, where Bouton’s condition on a subtraction game is sound and incomplete, and the direction of the error says what it may still be used for. Here the direction is the opposite one and the reading is the same: the difference test’s acceptances are almost always right, so a positive answer is nearly a proof and a negative answer is nearly nothing.

Comparison stops being a subtraction. Every ordered pair of coin rows compared two ways: by asking whether one does at least as well as the other in each of a family of contexts, and by asking whether their difference is not worse than nothing. Under the last-move convention the two tests are equivalent, which is what makes comparison computable. Here they are not, and the sharpest case is a row against itself — the contexts all say yes and the difference test says no, because a game beside its own negative does not come to nothing.
Fig. 5 The same comparison over different coins. The proportions move and the shape does not: the test accepts a handful it should not, refuses a great many it should, and refuses every row against itself. A test whose failures are that lopsided is a test with a use, and the use is not the one it was written for.

What a class restriction would have to buy

The rung below is about restricting the class of scoring games until a theorem applies, and this rung says what the most valuable restriction would be.

A class in which G + (−G) = 0 holds would give back the collapse, and with it the whole computational apparatus — comparison by subtraction, equality by comparison, and a canonical form to reduce to. That is the prize, and it is why the modern literature on scoring games keeps proposing classes rather than proving theorems about all of them.

What the count above says is how much such a class would have to exclude. Sixty-six of eighty-one four-coin rows fail the equation, so a class closed under the operation and containing the ones that cancel is at most a fifth of the family — and the fifteen that cancel are picked out by a condition on their parity sums, which is not a condition anybody would have proposed for its own sake.

That is the honest state of it. The restriction that buys the arithmetic is severe, it is not the same restriction as the one that buys the mean-value bound, and no single class has been found that buys both. A universe closed under the operations it is asked about is what the normal-play theory gets for free and what a scoring theory has to construct.

What the picture cannot show

The contexts are a stated family and not every context. The contextual test here asks about six small rows, which is a sample of the quantifier rather than the quantifier. A pair the sample accepts might be refused by some context outside it, so the “contextual” column is an upper bound on the true relation, and the count of incomplete pairs is therefore a lower bound. The unsound count is a lower bound too, for the same reason and in the same direction.

And the negation is a modelling decision. Negating the coins is the reading that makes the mirroring strategy available and the sum solvable by the same machinery, and it is the reading under which the rung three below’s claim about group structure is a claim about something. A different notion of what −G should mean for a scoring game would give different numbers, and there is no theorem here to say which is right — that is what having no group means.

A game where the last move decides nothing. Rows of coins taken from either end, with the exact score for each side moving first. Under the normal-play convention this family is settled entirely by the parity of the row — nobody is ever without a move until the coins run out — so normal-play theory returns the same answer for every row and it is not the answer anybody wants. The scoring answer depends on nothing but the numbers.
Fig. 6 Three of the fifteen rows that do cancel. Each has its odd-numbered squares summing to the same total as its even-numbered ones, which is exactly the condition under which the parity strategy hands both players the same haul — so the mirroring argument balances, and the sum comes to nothing after all.

Nothing here is about Go. Go is a scoring game and a Go player compares positions constantly, by counting; what a counting comparison is doing is asking about a number, which is a total order and has no difficulty at all. The difficulty measured here is about comparing games, which is a question about how they behave in sums, and a Go player never asks it.

The convention, named

Scoring, with no pass: the row empties and the higher total wins. The pass is absent deliberately, because a pass would make the sums easy in a way that hides the point — a player who may pass in a copy simply declines the coins they do not want, and the mirroring strategy becomes a strategy for keeping the score at nought rather than a strategy for doubling it.

That is worth checking rather than asserting and it is not checked here, which is a shortfall this rung records rather than repairs. What can be said from the argument alone is that a pass cannot restore the equation in general: passing keeps a player from being hurt, and the sum G + (−G) is a position where somebody profits rather than one where somebody is hurt.

The convention every other field uses is normal play, where the equation is a theorem with a one-sentence proof. Everything on this page is the measurement of what that sentence was worth.

The surprise: the failure is at the identity

The expected shape of “no inverses” is that some comparisons become hard. A few pairs would be undecidable by the cheap test, the cheap test would work on most of them, and a reader would learn to be careful at the edges.

What the counts show is that the cheap test fails at the easiest possible case. A row compared with itself is the comparison nobody needs a method for, every context accepts it by definition, and the difference test refuses all twenty-seven of them.

That is worth dwelling on because of what it says about where a broken structure breaks. The equation G + (−G) = 0 is not one fact among many; it is the fact that makes the difference test mean anything, so when it fails, the test does not degrade — it stops being a test of what it was a test of. Comparing a game with itself is where that is most visible, because it is where the true answer is known in advance and the method still gets it wrong.

The general shape: a computational shortcut is only as good as the identity it rests on, and the way to find out whether a shortcut survives a change of setting is to run it on the case whose answer is not in doubt. If it fails there it has not been weakened, it has been detached — and no amount of care at the edges will help, because the edges were never the problem.

Where the ladder goes next

scoring now has four rungs, and between them they price what a game with a number at the end costs: the last-move theory returns a constant, a pass fixes one hypothesis and removes the convention, the hypothesis has to hold hereditarily, and comparison stops being a subtraction.

The rung above is the one the field has been working on since 2010 and this ladder has only pointed at. Which restriction buys the most? The class with hereditary incentive gets the mean-value bound and not the arithmetic; a class closed under cancellation would get the arithmetic and is very small; and nothing found so far gets both. That is a rung about comparing the candidate classes on the same family and reporting what each one covers — and it is a live question rather than a historical one.

Part 4 of 8

One argument about Scoring. The parts either side of it:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.

AdditivityComparisonContextDifference gameEqualityExhaustive searchGroupNegationNormal playPartial orderScoring game