Temperature

How hot a real position is

Counted one value at a time, a tenth of the subject is hot. Counted one position at a time — every board this site has enumerated, all 11,397 of them — it is a twentieth, two thirds of the positions are worth numbers outright, and ten of the seventeen rulesets never produce a hot position at all.

Assumes: How hot a day gets · The values nobody's game produces

How hot a day gets measured the temperature of every value born by day three and found the collection overwhelmingly cold. It closed by naming the measurement it had not made:

The distribution here counts each value once; a distribution over positions — every board of a given size in a given ruleset, with the temperature of each — is a different object and the one a player would recognise.

Here it is, over the 11,397 positions the gamut census is built from: every Hackenbush string of at most seven edges, every Toads and Frogs strip of at most seven squares, every Clobber row, every small Domineering rectangle, and thirteen other rulesets besides. The reweighting is not a refinement. It moves every number on the page.

The same population counted twice. The temperature scale over the positions this site has enumerated, once with every position counted and once with every distinct value counted. The two disagree about how much of the subject is hot, about what the commonest hot temperature is, and about whether a number is the usual thing for a position to be worth.
Fig. 1 The same population counted twice: once with every position counted and once with every distinct value counted once. The two disagree about how much of the subject is hot, about the commonest hot temperature, and about whether being worth a number is the usual thing for a position to be.

Two thirds of everything is a number

Sixty-seven per cent of the positions are worth numbers. Twenty-eight per cent are worth a number plus something the temperature scale cannot see — the class at temperature nought, which below zero is the essay about. Four and a half per cent are hot.

Counted by value instead, the same population gives 39 per cent numbers, 51 per cent at nought and 11 per cent hot.

So weighting by position more than halves the hot share and nearly doubles the number share. The direction is asserted rather than reported — a run in which the hot share came out larger by position than by value would stop the build, because the sentence above is the page.

The reason is not subtle and it is worth stating plainly, because it is the whole mechanism. A value that a thousand positions are worth counts a thousand times here and once there, and the values with the most positions behind them are the cold ones. Every Hackenbush string of length seven is a different position and there are 254 of them, and every one is worth a number. Every Clobber row is all-small and sits at temperature nought, and there are 3,272 of them. The hot values, when they occur, are usually the value of one board or two.

The temperature scale, per position and per value. Two histograms over the same sweep of rulesets. The upper counts every position and the lower counts each distinct value once, and the shapes differ: the population of positions piles up at the numbers where the population of values spreads across the small temperatures.
Fig. 2 The two histograms side by side. The upper counts positions and piles up at the numbers; the lower counts values and spreads across the small temperatures, which is the shape the day-three census found.

The commonest hot temperature is not a half

Among the hot values of day three the mode is a half, and among the hot values in this sweep it is a half as well — thirty-one of them.

Among the hot positions it is one, with 144.

That is a reversal rather than a shift, and it comes from the same weighting with the opposite sign. The values at temperature ½ are many and thin: a great many distinct values happen to have that temperature and most of them are realised by one board apiece. The value {1 | −1} is one value with a temperature of one, and it is what a 2×2 Domineering board is worth, and a 3×3, and eleven of the 104 Domineering region shapes besides.

A histogram over values counts how many kinds of position there are. A histogram over positions counts how often a kind turns up. Nobody would confuse the two if they were about anything else, and the temperature scale is presented as a property of positions throughout, which is what makes the substitution so easy to make.

Ten rulesets that never get hot

The pooled figure is the wrong number to carry away, and the ruleset breakdown is why.

Which games make hot positions. Every ruleset in the sweep, with the share of its positions that are numbers, that sit at temperature nought, and that are hot. Ten of the seventeen never produce a hot position, and the three rulesets contributing most positions to the pooled census are three of the coldest.
Fig. 3 Every ruleset in the sweep, with the share of its positions that are numbers, that sit at nought, and that are hot. Ten of the seventeen have no hot position at all, and three of them contribute more than eight thousand of the 11,397.

Snort is hot on five positions in six. Toppling Dominoes is hot on seven in ten. Domineering is hot on a third. And then Hackenbush is hot on none of 254 positions, Cutcake and Maundy Cake on none of twenty-five each, Shove and Push on none of 728 each, Clobber on none of 3,272 and End-Nim on none of 340.

The ten cold rulesets are cold for reasons a reader can state in a line apiece, and the reasons are all about the rules rather than about any board:

  • Hackenbush strings, Cutcake, Maundy Cake, Shove and Push are games every position of which is a number, so they have no temperature at all. Nothing worth fighting over is the essay about the last of them and about how strange that is.
  • Clobber and green Hackenbush are all-small: a player has a move exactly when the opponent does, so no position can settle at anything but nought plus an infinitesimal. That class is the subject of its own essay and the reason Clobber is on this site at all.
  • End-Nim and Nim are impartial or nearly so, and an impartial position’s two walls are the same wall.

So hotness is a property of the ruleset far more than of the position. That is not what a scale defined position by position suggests, and it is what a sweep across rulesets makes obvious.

The two figures for the hot share say it in one line. Pooled over positions it is 4.6 per cent; averaged over the seventeen rulesets, one vote each, it is 14.1 per cent. The pooled number is mostly a fact about which games have small positions worth enumerating.

The three biggest contributors are three of the coldest

There is a sharper way to say what the pooled figure is measuring, and it is a fact about this sweep rather than about games.

Three rulesets supply 8,586 of the 11,397 positions: Clobber with 3,272, Toads and Frogs with 3,270 and Amazons on one line with 2,044. Clobber is all-small and produces no hot position ever. Toads and Frogs is hot on one position in eighteen. Amazons on one line is hot on one in fourteen.

What the pooled percentage is a percentage of. The seventeen rulesets of the sweep ordered by how many positions each contributes, with a running total and the share of each that is hot. The three largest contributors are 75.3% of the whole sweep and every one of them is colder than the ruleset average, so the pooled hot share of 4.6% is mostly a statement about which games have positions cheap to enumerate rather than about how hot a position is.
Fig. 4 The same seventeen rulesets ordered by how many positions each contributes rather than by how hot it is, with a running total down the third column. Three rulesets reach three quarters of the sweep, and all three are colder than the ruleset average — which the figure checks rather than asserts, and refuses to draw if a contributor at the top of the table ever comes out hotter than the average.

So three quarters of the sweep comes from three games that are between rarely and never hot, and they are the three not because they are typical but because their positions are strings — every arrangement of letters on a strip of seven squares, which is a large number cheaply enumerated. Snort, which is hot five times in six, contributes six positions, because Snort is played on graphs and this site has six of them.

The two orderings are worth setting side by side, because they answer different questions and only one of them is about games. Sorted by hot share, the table says which rules make fights; sorted by position count, it says what the pooled percentage is a percentage of. A pooled figure is a weighted mean whose weights nobody chose — they are the sizes of seventeen enumerations, and an enumeration is large when its positions are strings and small when they are graphs.

The pooled histogram is therefore a histogram of what is easy to enumerate. Correcting for that by giving each ruleset one vote is crude and it is at least a different crudeness, and it moves the hot share from 4.6 per cent to 14.1 — which is above the value-weighted figure rather than below it. Neither number is the answer to how hot is a position; what the two of them together establish is that the question has no answer independent of which games are being counted.

The values every game arrives at. The values produced by the most different rulesets, with a bar for the number of rulesets that produce each. The head of the list is the small integers and star; nothing complicated is shared by many games, and nothing shared by many games is complicated.
Fig. 5 The values the sweep produces most often, which is the same weighting applied to values rather than to temperatures. The commonest values are numbers and nimbers, and a hot value does not appear until well down the list — which is the position-weighted census arriving from the side of the values.

The hottest thing here is hotter than a day allows

Day three’s hottest value has temperature 2, and it is {2 | −2}, and it is unique on its day.

The hottest position in this sweep has temperature 4.

There is no contradiction and the reconciliation is the point. A day bounds the birthday of a value, and temperature and birthday are independent quantities: a Snort position on a six-vertex graph is worth a value born far later than day three, and a Domineering board with forty squares on it is routinely worth a value born on day two. Big is not the same as hot is the essay about that independence, and this is the same independence seen from the population rather than from a pair.

What follows is a small correction to how the day-by-day record reads. The record {n | −n} at temperature n on day n+1 is a fact about how far the integers get, and the integers get one further each day; a real game reaches a large temperature by having a large stake on the board rather than by being born late. Amazons on one line reaches 4 with seven squares.

The value census agrees about which games those are, which it need not have: the ten rulesets producing no hot position are the same ten producing no hot value, so the two sweeps identify the same cold games from opposite ends. A ruleset could in principle produce a hot value on some rare position and never in the enumerated range, and none of the seventeen does.

One thing the reweighting leaves alone, and it is the thing a reader would have expected to move.

The set of temperatures that occur is the same either way. Both censuses find the same fourteen values — the numbers at −1, nought, and twelve positive temperatures from an eighth to four — and neither finds a temperature the other misses. What changes is only how much weight sits on each. The agreement is asserted rather than observed: a sweep in which one count found a temperature the other did not would stop the build, because the sentence above would then be describing two populations rather than one.

That is a small point with a use. It means the reweighting is not reaching into a different part of the subject: the same values are being counted, in the same proportions within each ruleset, and the whole of the difference is the ruleset mix. Anything a reader concludes about which temperatures are possible is safe under either count. Anything concluded about which are common is not.

What a solver spends its time in

The practical corollary is about search rather than about theory, and it is the reason to have made the measurement at all.

A solver walking a game tree meets positions in proportion to how often they occur, which is what the position-weighted histogram counts. Two thirds of what it meets are numbers, where numbers avoid numbers says the right move is never there. Another quarter are worth a number plus an infinitesimal, where the temperature theory has nothing to say and the atomic weight has everything. One position in twenty-two is hot.

So the theory a reader meets first — thermographs, means, playing the hottest — applies to about five per cent of what a program actually looks at, and the theory that applies to the other ninety-five is the one presented as an appendix.

That is a statement about this sweep and not about games in general, and the honest form of it carries the caveat from the previous section: the sweep is heavy in the rulesets whose small positions are cheap to enumerate, and those are disproportionately the cold ones. A sweep weighted by how often a ruleset is played would look different and nobody has one.

The two shapes are already drawn together above, in the pair of histograms: the day-three census has its mode at a half and its tail at two, and the position-weighted version has its mode at a number and its tail at four. Same subject, same instrument, two populations — and every disagreement between them is a disagreement about weights rather than about temperatures.

The quarter at nought, which is the real story

If the page had to be reduced to one number it would not be the 4.6 per cent. It would be the 28.3.

More than a quarter of every position in the sweep is worth a number plus something the temperature scale is blind to. Those positions are not cold in any useful sense — a Clobber row is a genuine fight and whoever misplays it loses — and the scale gives every one of them the same reading, nought, which is the reading it gives a star. The scale has stopped resolving, and the resolution it has stopped at is where a quarter of the population lives.

That is a different complaint from most positions are numbers. A position worth a number is one the theory has completely described; a position at temperature nought is one the theory has described as far as a number and no further. Counted by value the two classes are 39 and 51 per cent; counted by position they are 67 and 28. Either way the second is large, and either way the instrument that reads it is the atomic weight rather than anything on this page.

What a weighted census is really weighting

Two counts of the same collection can disagree by a factor of two, and it is worth being explicit about which weighting each of them applies, because the choice is invisible in the word “count”.

Counting values weights every value equally, however many positions carry it. That is the right weighting for a question about the theory: how much of the value space is hot, how the construction distributes its objects, what a day contains. It is the wrong weighting for a question about a game, because it counts a value produced by one obscure position exactly as heavily as the value of nought.

Counting positions weights each value by how many positions realise it. That is the right weighting for a question about a solver — what a search spends its time in — and it is dominated by the values that are easy to realise, which are the numbers.

Neither is the weighting a player wants, and it is worth saying what that third one would be. A player meets positions in proportion to how often play reaches them, which is neither of the above: a game visits a small, structured subset of its positions and visits some of them constantly. The three weightings can disagree by an order of magnitude and each is correct for its own question.

That is why the same collection supports several honest headline numbers, and why a figure quoting one of them has to say which. The habit this site follows is to name the denominator in the caption — per value, per position, per play-out — because a percentage without one is a number that cannot be checked and can always be defended.

What the census does not say

Three limits, and the first is the one that most changes the reading.

The sweep is a sweep of small positions. Every ruleset here contributes what its enumeration reaches — seven squares of Toads and Frogs, six of Shove, nine small Domineering rectangles — and a ruleset whose positions get hot only when they get big is under-represented by construction. Domineering is the clearest case: nine boards, three of them hot, and the temperatures of larger boards are not in the figure because larger boards are not in the sweep.

Temperature nought is one bar and is not one thing. Twenty-eight per cent of the positions sit there, and they are worth a number plus something the scale cannot resolve. Which end of the interval is open is the essay about how various that class is, and it finds all four possible relations between a position and its own stop inside it.

And a position is not a board. Every count here is of a whole position of a ruleset. A real board in play has usually broken into pieces, and the temperature of a sum is not determined by the temperatures of its parts — so this census says what a component is likely to be worth and says nothing directly about what a board in play is likely to be.

The convention, named

Normal play. Every temperature is computed from the position’s own thermograph, as the height at which the two walls meet, and a number is given temperature −1 rather than nought throughout — the convention below zero sets out, which puts every number strictly below every position carrying a star.

The population is the one the gamut census uses, unchanged: seventeen rulesets, 11,397 positions, every value computed by the recursion and reduced to canonical form. Two positions are counted as one value when their canonical forms are identical, which is the only identification made.

Where the ladder goes next

The cold anchor has three rungs to here: the floor of the scale, how far the top of it reaches, and now what happens when the count is taken over positions instead of over values.

The four rungs above narrow the same measurement twice more and then reverse it. What a game actually produces goes from a catalogue of regions to the regions a game actually reaches: 53 per cent of Domineering shapes of at most eight squares are hot, and only sixteen per cent of the components eleven hundred random games produce. The figure holds at three board sizes, so it is a property of play — and it means every temperature census taken over a catalogue, this page’s included, overstates the heat by about a factor of three.

One fight makes a board a fight then tallies at the board rather than at the piece and gets 32 per cent, which is twice the piece figure and not ten times it: a Domineering board carries only 1.68 pieces, and the hot ones cluster on the same boards rather than spreading across them.

The reversal is the rung after. The obstacle was the catalogue removes the limit that made the early game unmeasurable — a twelve-square region evaluates in five milliseconds, an eighteen-square one in under a second, and what was expensive was cataloguing every shape rather than sweeping the positions a board actually reaches. Swept that way, a board is hot four times in five three moves in, and cools as it breaks up.

So the coldness this page reports is a fact about the endgame rather than about the game. A played board starts hot and ends cold, and a census over all positions is dominated by the ending, because that is where most positions live.

Eight squares and no hotter closes the anchor with a ceiling: no Domineering position is hotter than three halves, the attaining region has eight squares, five of them exist up to symmetry, and the ceiling holds at nine and ten where the obvious extrapolation predicted more.

Part 3 of 9

One argument about Cold. The parts either side of it:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 10.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.

All-smallCold gameCold positionDay threeDecompositionEnumerationExhaustive searchFreezing pointHot gameHot positionInfinitesimalMean valueNumberRule tableTemperatureThermograph