Out in the world

The restriction that buys the most

Four candidate classes of scoring game, scored on the same two families and the same three questions. The class everyone expects to be tiny — the rows that cancel against their own negatives — is empty on rows of three and the widest restriction on rows of four, where it holds fifteen rows against the hereditary class's twelve and gets all 225 of its comparisons right against 108 of 144. The trade everyone expected does not exist.

Assumes: Nothing to subtract with · A hypothesis has to hold all the way down

Four measurements here each establish one thing a scoring game lacks. Counting at the end changes everything finds the last-move theory returning one answer for every position. What a pass is worth to a theory finds a pass repairing Milnor’s hypothesis at the cost of the convention everything else is built on. A hypothesis has to hold all the way down finds the hypothesis needing to hold hereditarily. Nothing to subtract with finds comparison ceasing to be a subtraction.

Each is a measurement about the whole family, and between them they suggest a shape. A class small enough to have inverses would get the arithmetic and cover almost nothing; a class defined by the incentive condition covers a good deal and gets the bound and not the arithmetic; and nothing gets both. That is the trade this page was set up to price.

Priced on four classes, two families and three questions, the trade is not there. The class with inverses is the largest of the restrictions on one family and it buys all three, and on the other it is empty.

Four classes, three questions, one table

Four restrictions, and what each one buys. Four candidate classes of scoring game — every row, the incentive condition at the top, the same condition at every subposition, and the rows that cancel against their own negatives — scored on two families of coin rows for the mean-value bound, for comparison by subtraction, and for cancellation.
Fig. 1 Four restrictions scored on the same two families for the three things a scoring game is missing: the mean-value bound on pairs from the class, comparison by subtraction on ordered pairs from the class, and cancellation against a member’s own negative.

The classes are the candidates those four measurements produce.

Every row is the baseline and the thing to beat. The incentive condition at the top is Milnor’s hypothesis as a reader would check it, on the position in front of them. The same condition all the way down is what an induction over the play can actually use. The rows that cancel against their own negatives is the smallest thing that could be called a group, and it is the class the arithmetic would need.

The questions are what each class is being asked to buy. Does the mean-value bound hold on every pair drawn from it? Does comparison by subtraction agree with comparison in every context, on every ordered pair from it? Does every member cancel against its own negative?

Asking the questions of pairs from the class rather than from the family is what makes this a comparison of classes rather than four readings of one sweep. A property that holds of the family restricted to a class is a property the class has; a property that holds of the family is a property nobody had to restrict for.

On rows of three, the expected answer

The family of rows of three coins drawn from −2, 1 and 3 is where a hypothesis has to hold all the way down does its work, and on it everything behaves as expected.

The bound holds on 96 of the 378 pairs — a quarter — so a reader with no restriction at all has a mean-value theory that is wrong three times in four. Restricting to the fifteen rows satisfying the incentive condition at the top takes the pairs from 378 to 120 and the bound still fails on 24 of them, which is that essay’s finding and is why the hereditary reading exists. Restricting to the seven rows satisfying it everywhere takes the pairs to 28 and the bound holds on all of them.

So on this family the hereditary class buys exactly what it is advertised to buy: one property, at the price of keeping seven rows of twenty-seven.

The price is worth stating in the currency a reader would pay it in. Seven rows of twenty-seven is a quarter of the family, and the pairs fall from 378 to 28 — from three quarters of the positions a theory might be asked about to a fourteenth of them. A theorem covering a fourteenth of the cases is not useless, but it is a long way from the normal-play situation, where the mean value and the temperature are defined for everything without a hypothesis at all. Worth nothing, and worth fighting for develops those two quantities in their own setting, and the contrast is the whole reason a scoring theory is hard to state: the same quantities, defined the same way, needing a hypothesis that throws away nine tenths of the family before they can be used.

And the hypothesis is not one a reader can check by looking. Fifteen rows pass it at the top and seven pass it everywhere, so eight rows of twenty-seven look safe and are not — which is the same finding restated as a count of rows rather than of pairs.

It buys nothing else. Comparison by subtraction gets 24 of its 49 ordered pairs right — barely half — and not one of the seven rows cancels against its own negative. The cancelling class on this family is not small. It is empty.

On rows of four, it inverts

Coverage against arithmetic, and no trade between them. Each restriction by how much of its family it covers and how many of the three properties it buys. The class with inverses is empty on rows of three and the largest restriction on rows of four, where it buys everything — so coverage and arithmetic are not being traded against each other.
Fig. 2 The same restrictions read as a trade: how much of its family each covers against how many of the three properties it buys. The class with inverses has no members on one family and is the widest restriction on the other.

Add a coin. The family of rows of four from the same three values has 81 rows, and every expectation the smaller family produced is wrong on it.

The bound holds on all 3,321 pairs, with no restriction at all. Milnor’s condition holds at the top of every one of the 81 rows, so the restriction that took an essay to distinguish from its hereditary version separates nothing here — the class of rows satisfying it is the family, row for row and pair for pair. A reader who arrived at this family first would have concluded that the mean-value theory needs no hypothesis.

And fifteen rows cancel against their own negatives. That is more than the twelve satisfying the incentive condition hereditarily. The class that was supposed to be the small expensive one is the larger of the two restrictions on this family.

What it buys is everything. The bound holds on all 120 of its pairs. Every row in it cancels, by construction. And comparison by subtraction agrees with comparison in every context on all 225 of its ordered pairs, where the hereditary class manages 108 of 144 and the unrestricted family 6,037 of 6,561.

So the answer to which restriction buys the most is the one with inverses, and it is not paying for the privilege in coverage. On rows of four it covers more than its rival and buys three properties to its rival’s one.

A game where the last move decides nothing. Rows of coins taken from either end, with the exact score for each side moving first. Under the normal-play convention this family is settled entirely by the parity of the row — nobody is ever without a move until the coins run out — so normal-play theory returns the same answer for every row and it is not the answer anybody wants. The scoring answer depends on nothing but the numbers.
Fig. 3 Three rows of four, with the scores each produces. The first and third cancel against their own negatives and the second does not — a single coin’s difference, and the row leaves the class.

The fifteen are not a family a reader would have guessed at either. They include every constant row, the rows that repeat a pair, and the rows whose two halves mirror each other — and they exclude rows that differ from a member by one coin. Changing the last coin of −2, −2, 1, 1 from 1 to anything else takes the row out: −2, −2, −2, 1 added to its own negative scores 2 with Black moving first and −2 with White, which is a two-point advantage to whoever has the move on a position that is supposed to be nothing.

That is what a class defined by a test rather than by a construction looks like. The condition is checked on each row and the rows that pass it have no obvious shape in common, which is a reason to be careful about the word class and a reason the closure question at the end of this page is the one that matters.

Why coverage is not a property of the restriction

Empty on one family and the widest restriction on the next is a strange thing for a class to be, and it is worth being clear about what it is not.

It is not a small-sample artefact: the rows of three are all 27 of them and the rows of four all 81, and both counts are exhaustive over their values. It is not the values changing, because they do not. It is a parity. A row of three coins is taken by three moves, so one player takes two coins and the other takes one, and a row added to its own negative is a six-coin position in which the mover’s advantage does not vanish. A row of four gives each player two, and the cancellation the mirror strategy would produce under normal play has somewhere to land.

That is the sense in which the coverage of the cancelling class is not a fact about the class. It is a fact about the family — and the same is true of the incentive condition, which separates fifteen rows of twenty-seven at length three and nothing at all at length four.

A restriction is a subset of a family, and every measurement of what it covers is a measurement of both. The expectation — a class with inverses would be very small — reads as a statement about scoring games and is a statement about whichever family somebody happened to check. The board falls apart, and the arithmetic changes is the same hazard one level down: a property that looks like the object’s turns out to be the sample’s.

The two ways a difference test is wrong

The two ways a difference test can be wrong. Comparison by subtraction set against comparison in every context, for each class and family, with the failures split into the test accepting a pair the contexts refuse and the test refusing one they accept. Restricting the class removes the first kind entirely and leaves the second.
Fig. 4 Comparison by subtraction against comparison in every context, on ordered pairs, with the failures split. The test accepting a pair the contexts refuse and the test refusing one they accept are different complaints, and the restrictions remove one and not the other.

The comparison column hides a distinction that decides how much any of this is worth.

A difference test can fail two ways. It can say yes where the contexts say no, which is unsound — a reader would replace a position by one that is worse in some context and lose a game they had won. Or it can say no where the contexts say yes, which is incomplete — a reader loses a simplification they were entitled to and nothing else.

On the unrestricted family of rows of three there are nine unsound pairs. Every restriction removes all of them: not one unsound pair survives the incentive condition, the hereditary condition or the cancelling condition, on either family.

What survives is incompleteness, everywhere. The hereditary class on rows of four refuses 36 of its 144 ordered pairs that the contexts accept; the unrestricted family of rows of four refuses 524 of 6,561. The cancelling class refuses none.

So the restrictions are worth more than the agreement counts suggest. A class whose difference test is merely incomplete is a class in which the test can be used — it proves what it claims and fails to prove some things that are true, which is what every sound-and-incomplete tool in mathematics does. Comparing positions sets out what the normal-play test is and why it is both; here the soundness is the thing a restriction buys first, and it is bought cheaply.

Comparison stops being a subtraction. Every ordered pair of coin rows compared two ways: by asking whether one does at least as well as the other in each of a family of contexts, and by asking whether their difference is not worse than nothing. Under the last-move convention the two tests are equivalent, which is what makes comparison computable. Here they are not, and the sharpest case is a row against itself — the contexts all say yes and the difference test says no, because a game beside its own negative does not come to nothing.
Fig. 5 The measurement the comparison column comes from: comparison by difference against comparison in every context, over a family and a list of contexts. The disagreements are counted in both directions because they are different faults.

The counts also say something about where a scoring theory would have to be careful, and it is not where a reader would expect. The unsound pairs are all in the smaller family, where the bound is also at its worst — so the family on which the theory is hardest to state is the family on which a careless comparison is actually dangerous, and the family where the theory is easy carries no unsound pairs at all. A reader who checked the difference test on rows of four and concluded it was sound would be right about rows of four and wrong about the game.

What a class has to be before it is a class

A game beside its own mirror, and what is left over. Every coin row added to its own negative, played out exactly, with the resulting scores counted. Under the last-move convention every such sum is worth nothing, because the mirroring strategy guarantees the second player the last move. Here the same strategy is available and the score it produces is not nothing: the mirror of a coin conceded is another coin conceded. Gold is the sums that do come to nothing, which are a minority.
Fig. 6 The measurement the cancelling class is built from: a row added to its own negative, and what it scores. Under normal play that sum is nought by the mirroring strategy; here it is a number that depends on the row.

One thing none of the four candidates has is closure, and it is the property that would turn any of them into a theory rather than a test.

A class worth restricting to has to be closed under the sum: if two positions are in it, their sum should be too, or the arithmetic it licenses cannot be applied twice. Nothing above checks that, and the classes are defined by tests on a single row — this row satisfies the incentive condition, this row cancels — with no reason for the property to survive addition.

That is not a small omission and it is why this page compares what the classes buy rather than announcing one of them as the answer. The class with inverses on rows of four buys all three properties on the pairs measured, and whether the sum of two of its members is again one of its members is a question about rows of eight, which is outside the family and outside the search.

The condition has to hold underneath, not on top. Pairs of coin rows sorted by where the incentive condition holds, with Milnor's bound checked on each pair. Rows that satisfy the condition at every subposition never break the bound. Rows that satisfy it only at the top break it on a counted fraction — and a reader who tested the row rather than the row's insides would have called those safe. The distinction is invisible from the position and decides whether the theorem applies to it.
Fig. 7 The earlier measurement, which this table is set against: pairs of rows of three sorted by where the incentive condition holds, with the bound checked on each. The middle band is the finding — fine at the top, broken underneath, and the bound goes with it.

What the comparison cannot say

Two families are not many. Rows of three and rows of four, from one set of three values, and the answer inverts between them. That is enough to establish that the coverage of a class is not a property of the class, and it is not enough to say what the classes look like on rows of five or six — where the parity argument above predicts the cancelling class reappears at every even length and vanishes at every odd one, and nothing here checks it.

The contexts are six. Comparison in every context means every context, and what is computed is comparison in six small ones. A pair the six accept might be separated by a seventh, so the contextual column is an upper bound on agreement and the unsound counts are lower bounds. The incompleteness counts are the reliable direction.

And closure is unchecked, as the section above says. Without it none of these is a class in the sense the theory would need, and the table is a comparison of tests rather than of theories.

The convention the classes are defined under

Every row is played to the end: the coins are taken from the ends until none is left, and the score is the difference between what the two players took. There is no last-move convention anywhere on this page, which is the point of the whole question — counting at the end changes everything is what happens when one is imposed.

The negative of a row is the row with every coin’s value negated, which is the same position with the players exchanged, and it is what a negative has to be for cancellation to mean anything. The mean and temperature of a row are half the sum and half the difference of its two scores, which is Milnor’s definition and not an invention of this page.

The incentive condition is that the score with Left moving first is at least the score with Right moving first — having the move is not a disadvantage — and it is checked either at the row or at every subposition the play can reach, which is the distinction a hypothesis has to hold all the way down draws.

The surprise: the trade was a description of one family

The expectation this page was written to test is a good one and it is stated in the right form: a stronger restriction covers less. Inverses are a strong thing to ask for, so the class with them should be small, and the arithmetic should be paid for in coverage.

Every part of that is true on rows of three, where the cancelling class is empty and the hereditary class is seven rows of twenty-seven. Every part of it is false on rows of four, where the cancelling class is fifteen rows and the hereditary class twelve, and the larger class is the stronger one.

What went wrong is that coverage was read as a property of the restriction. It is a property of the pair — restriction and family — and the two families here differ in a way that has nothing to do with scoring games at all. Four coins split evenly between two players and three do not.

That is worth carrying beyond scoring games, because the shape recurs wherever a class is proposed. Outcomes do not add is the same lesson about a summary: a quantity that describes a family perfectly can describe nothing about the objects in it. A class is a quantity of that kind. How much does this restriction cost is not a question with an answer until somebody says what it is restricting.

Still open: whether any of these is closed under addition

The measurement this page most wants and cannot make is closure.

For each of the four classes, the question is whether the sum of two members behaves like a member — and behaves like has to be made precise, because a sum of two coin rows is a two-row position rather than a row and so is not in the family at all. The honest version is a test rather than a membership: for two members of a class, does the two-row position satisfy the class’s own condition, read as a condition on positions instead of on rows?

The incentive condition has an obvious reading on a sum, since a sum has two scores like anything else. Cancellation has one too: does the sum of two cancelling positions cancel? If it does, the cancelling class on rows of four is a group in the only sense a scoring theory could use it, and everything on this page about what it buys becomes a statement about an object rather than about a list. If it does not, then what was measured here is a set of rows with a shared property and not a class, and the right next question is what the smallest closed class containing them looks like.

Part 5 of 8

One argument about Scoring. The parts either side of it:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.

BoundComparisonContextDifference gameExhaustive searchGroupIncentiveMean valuePartial orderScoring game