Four positions, sampled
Assumes: The names are not built out of the old ones · Four thousand nine hundred regions with no name
A loopy region — a small game whose move graph has a cycle, so that play can go on for ever — has no brace expression, because the brace form is a recursion and a cycle gives it nothing to stand on. What it has instead is two ordinary games, one for each way of settling a play that never ends: give every infinite play to Left and the region behaves like one game, its onside; give every infinite play to Right and it behaves like another, its offside. A loop is written with two names introduced the notation, and the question it raised is how many names the notation needs.
Two censuses answered that for small regions. Every loopy region of two positions has both sides written by ten names. Four thousand nine hundred regions with no name counted the regions of three positions — 110,934 of them — and found 4,931 with a side the two-position vocabulary does not reproduce. The names are not built out of the old ones showed those missing sides are not sums of old names, and that forty-eight names — thirty-five of the old and thirteen new — cover every side of every three-position region.
Ten, then forty-eight. The curve has two points and a shape cannot be read from two points.
Why four positions cannot be counted
A region of four positions is a move graph on four points. At each point Left may have any subset of the four points as moves, and so may Right: sixteen choices each, 256 per point, and — over four billion graphs. Most are not regions in the relevant sense, because not every point is reachable or no play repeats, but a census would still have to generate them all, and each surviving region needs its sides identified against twenty-two test games by solving every sum as a game with draws. The three-position census takes under half a minute; four positions are sixteen thousand times as many graphs, and the regions among them are larger and slower to identify, so the same method would run for weeks.
So the measurement the earlier essay proposed is a sample. Draw regions at random, identify both sides of each exactly as the three-position census identified them, and count how often the earlier names suffice. A sample cannot give the vocabulary’s size; it can give a floor on it and a measure of how often the old names fail. It can also do something a census cannot, which is to be repeated under a different rule for drawing, and see which of its answers move.
How the regions were drawn
A random move graph needs a rule for which moves are present, and the rule matters more than it looks.
The first sample includes each possible move with probability one half — every one of the possible moves is a coin toss. That gives graphs with about sixteen moves each: dense graphs, in which almost every position can move almost anywhere. Nearly every such graph is a region — 3,000 of 3,204 drawn — and most of them have plays that go on for ever: 2,485 of the 3,000 have some sum with a test game that is a draw.
Dense graphs are not typical regions, only typical coin tosses. In a dense graph both players usually have a way to keep the play going, and a region in which one player can always loop reduces to one of a few simple stoppers — the games on and off and their relatives, which are among the old names. So a dense sample flatters the old vocabulary. The second and third samples include each move with probability about a third and a quarter, giving about twelve and nine moves per graph, and it takes 4,090 and 6,284 draws to find 3,000 regions among them.
What the old names cover
The answer depends on the density, and the dependence is the first finding.
In the dense sample, 94.7 per cent of regions have both sides written by the thirty-five names the two-position regions use. Adding the thirteen three-position names brings that to 99.5 per cent, and fifteen regions of three thousand need something new. In the middle sample the old names cover 91.2 per cent, the forty-eight cover 97.9, and fifty-six regions need a new name. In the sparsest the figures are 87.0, 95.9, and ninety-four regions — about one in thirty-two.
For comparison, the three-position census found 4.4 per cent of regions with a side the vocabulary available before it — the two-position names — could not write. The sparsest four-position sample finds 4.1 per cent of regions with a side outside the vocabulary available before it, the forty-eight names, and 3.1 per cent with a side no earlier name of any kind reproduces. So at comparable sparseness, a fourth position does not make the earlier names fail more often than a third did. It makes them fail about as often, and some of the failures are exactly the ones the third position already named.
That is the second finding, and it is the one the earlier essay asked about directly. The thirteen names invented for three-position regions are not a peculiarity of three positions. Eleven of the thirteen reappear in the dense and middle samples, and ten in the sparsest. The names added at each size are reused at the next.
How fast new names arrive
Every sample met sides that no earlier name reproduces. Twelve distinct new sides in the dense sample, thirty-one in the middle one, thirty-five in the sparsest. The count that matters is not how many were met but whether the samples are close to meeting them all.
The three curves differ in the way that matters. The dense sample’s count climbs steadily and has not slowed by three thousand regions, so it has not seen most of what dense regions need; but it meets new sides only rarely, fifteen regions in three thousand. The middle sample climbs faster and is still climbing. The sparsest climbs fastest early — eighteen new sides in its first thousand regions, fifteen more in the second — and then almost stops, with two new sides in its last thousand. A curve that flattens is a sample running out of new things to find, and it suggests, without proving, that the sparse four-position regions need a few dozen new names and not hundreds.
The structure of the repeats says the same. In the sparsest sample, of the thirty-five new sides only twelve were met once; nine were met twice and the commonest twenty-one times. A population with a few very common new sides and a tail of rare ones is a population whose size a sample of this scale can bound but not pin down.
The third point
The vocabulary curve now has three points: ten, forty-eight, and at least eighty-three. The third is a floor — the forty-eight plus the thirty-five new sides the sparsest sample met — and the true figure is larger by whatever the samples missed, which the growth curves say is not much for sparse regions and an unknown amount for dense ones.
What the three points rule out is the alarming shape. The regions themselves grow enormously with the number of positions: two positions give 256 graphs, three give 262,144, four over four billion. If the names needed grew anything like that, the two-name notation would be useless at four positions, a list as long as the objects it names. It grows by a few dozen names a position. Every region still has two names; there are more names to choose from; and most of the new ones are needed rarely.
That is the sense in which the notation grows by accretion rather than collapse. Each size adds names for the regions the earlier names cannot write, and keeps using the earlier names for everything else. Nothing about four positions makes the three-position names obsolete — ten or eleven of the thirteen come back — and nothing suggests that four positions need names of a different kind.
Why sparse regions need more names
The density dependence is not an accident of the sampling, and it can be read off what a side is.
A side is decided by who wins each sum when infinite play goes to one player. In a dense region both players usually have a move that returns the play to where it was — a loop they can enter at will — and a region in which a player can loop whenever they like tends to be decided entirely by that fact. Such regions collapse onto a handful of names: on, which lets Left pass for ever, off, which does the same for Right, and small games added to them. A stopper and how to find one is the account of which loopy games stop and which do not, and in a dense four-position graph almost every position is the non-stopping kind that those few names describe.
In a sparse region the loops are fewer, and a player can only reach one by going through positions where the other player has choices. The winner of a sum then depends on the details: which loop is reachable from where, who can leave it, what is waiting at the exits. Those details are what distinguish one side from another, and there are more of them to distinguish when the graph is sparse enough that the loops do not dominate. So sparse regions need more names for the same reason that sparse graphs have more different shapes: the density that makes a graph a region also makes it uniform.
That also explains why a census at three positions and a sample at four agree at comparable sparseness. Three positions are small enough that most regions are sparse whatever their edges, because there are only so many ways to connect three points. The four-position dense sample is a different population, and the sparse one is the population the three-position census resembles. One part that never ends is the other place this subject found that the density of a loopy game’s graph decides how much of it the simple outcomes can describe.
The earlier counts, for comparison
The four-position figures only mean something beside the three-position ones, which are worth having in view.
The three-position census is complete: every region, every side. Its figure of 4,931 unnamed regions in 110,934 is exact, and it is 4.4 per cent. The four-position figures are samples, biased by their drawing rule, and their shares move by a factor of six as the density changes. The honest comparison is between shapes, not numbers: at both sizes, a few per cent of regions need names the smaller vocabulary lacks, and the new names at each size are a small set relative to the regions.
The earlier essay’s negative result — the missing names are not sums of old ones — is the reason the four-position count matters. If new names were built from old ones, the vocabulary at four positions would be determined by the vocabulary at three, and the curve would be a question about arithmetic. Because they are not, each size can bring names nothing before it predicts, and only counting finds them.
What the notation was for
The two-name notation was never meant to be a catalogue. It came out of the treatment of games that go on for ever in Conway’s On Numbers and Games and the loopy chapters of Winning Ways, where the point of writing a region as onside & offside was that each side is an ordinary game with a value, and ordinary games can be added, compared and simplified. A loopy region was to be handled by handling two finite things. The notation was the argument is the general form of that idea: a way of writing things down is a claim about what matters in them.
A catalogue of names is what the notation becomes when it is used to identify regions rather than to reason about them, and that is what these censuses and samples do. The finding that the catalogue grows slowly is a finding about how well the original claim holds up: most loopy regions really are two ordinary games in disguise, and the ones that are not are rare enough to name one at a time. Eighty-odd names for four positions is a long list for a table and a short one for a theory, and the theory still works on every region the list covers.
The convention named
A region is a move graph on four positions with a marked start, every position reachable from it and some play repeating. Its onside is the game it behaves as when every infinite play is a win for Left, and its offside the game when every infinite play is a win for Right. A side is identified by its column of winners — Left moving first and Right moving first — against each of the twenty-two values born by day two, and two sides with the same column are given the same name. The names are the thirty-five two-position names the three-position regions used, and the thirteen invented for three positions; a side matching neither is new.
The test set is the day-two values, the smaller of the two the earlier census could use. The earlier census measured what that smaller test set gets wrong on two-position regions, and the same caution applies: two sides the day-two test set cannot tell apart might be separated by a larger one, so the new-side counts are themselves floors.
What the sample cannot show
The sample cannot give the size of the four-position vocabulary, only a floor under it and an impression of how fast it is approached. The three densities give three different floors — twelve, thirty-one and thirty-five new sides — and the true number is at least the largest, plus whatever sides the samples jointly missed, which the flattening of the sparse curve suggests is modest and the steady climb of the dense one says is not zero.
Nor does the sample say which regions need the new names, or whether the new sides have short descriptions of their own. It also cannot say anything about draws beyond the convention it uses: every side is identified with infinite play given wholly to one player, and an outcome with no value behind it is the reminder that a draw is a third result the two-name notation deliberately sets aside. The earlier essay found that the thirteen three-position names are not sums of old ones; whether the thirty-five four-position sides are sums of the forty-eight — or of each other — is not tested here, and it is the question that would decide whether the vocabulary is growing in a structured way or merely growing.
Still open: whether the new names have a pattern
The three-position names were invented and left as a list. With four positions there are now at least thirty-five more, and a list of eighty-three names is the point at which a notation either finds an organising principle or stops being a notation. The natural first test is the one that failed at three positions, repeated one size up: are the new four-position sides sums of two earlier names, or of an earlier name and a small game? If they are, the vocabulary at four positions is generated by the vocabulary at three; if they are not, as at three, then every size brings names that must be counted to be found, and the curve’s shape — accretion by a few dozen a position — is the most that can be said about it.
Part 9 of 9
One argument about Notation. The parts either side of it:
The objects named here
The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.
DrawExhaustive searchIdentificationLoopyNotationRetrograde analysisSamplingStopperVocabulary
- The first theorem, and the winner it declines to name draw, exhaustive search, loopy, retrograde analysis
- The one outcome that adds draw, exhaustive search, loopy, retrograde analysis
- What the play keeps coming back to draw, exhaustive search, loopy, retrograde analysis
- When never ending is a win draw, exhaustive search, loopy, retrograde analysis
- A ko is won somewhere else draw, loopy, retrograde analysis
- A position with no value, and the rule that gives it one draw, loopy, retrograde analysis