Temperature

A description, and not a detector

The rung below noticed that the count of shapes attaining the hottest temperature grew across a plateau and collapsed at the step, and proposed it as a way to read a plateau off a single size. The growth is exact — five plateaus, no exception — and the rule is impossible: five orbits precede a rise at seven squares and no rise at eight.

Assumes: The ceiling was a plateau · Eight squares, and no hotter

The ceiling was a plateau extended the region sweep to eleven squares and found the hottest Domineering region at 7/47/4, against the 3/23/2 that had held at eight, nine and ten. A bound turned out to be a plateau three sizes wide, and the correction cost a sweep of 34,053 shapes to make.

It closed by naming a diagnostic it had used to explain the result and never to produce one:

This page’s best diagnostic was one nobody used: the number of shapes attaining the maximum, which grew from eight to fourteen to forty-three across the plateau and fell to four at the step. If that is a general signal — a maximum attained by many shapes is about to be exceeded, and one attained by few has just been reached — then it reads a plateau off a single size, which is what the whole difficulty here is.

It is not a general signal. The half of it that is an observation is exact, without a single exception in the sweep; the half that would be a rule cannot be made into one at all.

The counts, beside what happened next. The hottest Domineering region of each size with the number of shapes attaining it, and whether the next size was hotter.
Fig. 1 Every connected region to eleven squares, with the hottest temperature of each size, the number of shapes reaching it, and what the next size did. The fourth column is the diagnostic and the last is what a rule built on it would have to predict. Reading the two columns together is the whole of the test, and it can be done by eye: fourteen attainers at nine squares precede no rise and ten at seven squares precede one.

What is being counted

The quantity has to be pinned down before a rule about it can be tested, because two of the three words in shapes attaining the maximum are doing work.

A region here is a connected set of squares, treated as a Domineering position on its own: Left places vertical dominoes and Right horizontal ones, and neither may use a square outside the region. Its temperature is what what is at stake defines — the height at which the two walls of its thermograph meet, which is half the gap between what the two players can achieve if each moves first. A region of temperature 7/47/4 is one where moving first is worth seven-quarters of a move, and below zero is where the bottom of that scale is set.

The maximum over a size is then the hottest such region among all shapes of that many squares, and the attaining count is how many shapes hit that number exactly. Both halves of that are exact: the values are dyadic rationals and the comparison is an equality, not a comparison against a tolerance. That matters here more than it usually does, because the entire question is how many shapes land on one number precisely, and a tolerance of any size would turn attaining into nearly attaining and produce a different count with no warning.

And the maximum over a size is a fact about the catalogue, not about a game. How hot a real position is measured the temperatures of regions that games actually produce and found them well below the hottest available, so nothing on this page is a claim about what a player will meet. The hottest eleven-square region is the hottest eleven-square region; whether any game reaches it is a separate question with a discouraging answer.

The observation, which is perfect

Take the claim that is about what a plateau looks like from outside it, and it holds everywhere.

Every plateau, and every step. The attaining count across each plateau of the maximum temperature, and what it does at the step that ends it.
Fig. 2 Each run of sizes sharing a maximum, with the attaining count across it. Every one of the five grows at every step — 1, 2, 3 at nought; 1, 3 at one; 4, 10 at five-fourths; 8, 14, 43 at three-halves — and the count falls at three of the four steps that end a plateau. The figure refuses to draw unless the growth is monotone on every plateau, since that is the claim it makes.

The attaining count grows across every plateau, at every step, five for five. The three-halves plateau the rung below described runs 8, 14, 43. The five-fourths plateau below it runs 4, 10. The plateau at one runs 1, 3. Even the trivial plateau at nought — sizes one, two and three, where no region is hot at all — runs 1, 2, 3. There is no exception and no near miss.

And the count falls at the step, three times in four. It goes 43 to 4 at eleven squares, 10 to 8 at eight, and 3 to 1 at four. The exception is the step to six squares, where it goes 3 to 4 — a rise, small and in the wrong direction, at the size where the counts are smallest and a difference of one is everything.

So as a description the diagnostic is very good. Anybody looking at the completed sequence and asked to say where the plateaus were could read them off the attaining counts alone, and would get five of five and three of four.

The rule, which is impossible

The rung below’s sentence was not a description. It said a maximum attained by many shapes is about to be exceeded, and one attained by few has just been reached — which is a rule taking one size’s count and predicting the next size’s maximum. To be usable it needs a cut: some number above which the count means a rise is coming.

No number is the cut. The pairs of sizes on which any threshold rule for the attaining count would give the wrong answer.
Fig. 3 Every pair on which a threshold rule has to be wrong — a size whose maximum then held, carrying at least as many attainers as a size whose maximum then rose. There are seven such pairs, and the worst is the sharpest object in the sweep: nine squares has fourteen attainers and holds, seven squares has ten and rises. Any cut low enough to fire at seven fires at nine, and any cut high enough to spare nine spares seven.

There is no such number. Nine squares has fourteen attainers and the maximum holds at ten. Seven squares has ten attainers and the maximum rises at eight. A cut low enough to catch seven catches nine; a cut high enough to exclude nine excludes seven. And nine and seven are not the only pair: five squares has three attainers and rises, while six has four and holds, and eight has eight and holds.

Seven pairs in a sweep of eleven sizes have the wrong order. That is not a threshold that needs tuning.

The natural repair is to count orbits rather than shapes, and it is the right instinct — a shape and its mirror image are the same region twice, so a count of drawings overstates how many genuinely different shapes attain the maximum. The rung below reported both, and the orbit counts across the three-halves plateau run 5, 7, 22 against the shape counts 8, 14, 43.

Counting orbits does no better. The same threshold test applied to the number of symmetry classes attaining the maximum rather than the number of shapes.
Fig. 4 The same test on the orbit count. It does no better and it fails harder, because the failure is exact rather than merely ordered: seven squares has five orbits and the maximum rises, eight squares has five orbits and it holds. The same value of the diagnostic sits before both outcomes, which is not a cut that needs finding.

It does no better, and the way it fails is worse. Seven squares has five orbits and rises; eight squares has five orbits and holds. The diagnostic takes the identical value before the two opposite outcomes. A threshold rule can be defeated by an inversion, which is a statement about ordering and might be repaired by a better statistic; a rule defeated by an exact tie is being told that the quantity does not carry the answer.

Why the count grows

The reason the observation is perfect and the rule is impossible turns out to be the same reason, and it is deflating.

The share goes the other way. The attaining count as a fraction of all shapes of that size, with the growth rate of each.
Fig. 5 The attaining counts as a share of the shapes there are. The catalogue roughly triples and a half per size — 730 shapes at eight squares, 2,542 at nine, 9,287 at ten — and the attaining set grows more slowly, so the share falls across every plateau it can be read on: 1.10 per cent, 0.55, 0.46 across the three-halves plateau. The count rises because the population rises.

The number of connected shapes of nn squares grows by a factor of about three and a half per square. Between eight squares and ten the catalogue goes from 730 to 9,287, and the attaining set goes from 8 to 43 — which is growth, and it is growth at a slower rate than the population. As a share of the shapes there are, the attaining set shrinks across the plateau: 1.10 per cent, then 0.55, then 0.46.

So the count grows across a plateau for the least interesting reason available. It is not that the bound is being approached and more and more shapes are crowding up against it; it is that there are more and more shapes, and a fixed-ish fraction of an exploding population is an exploding number. Normalising the growth away — dividing by the population — removes the signal entirely and replaces it with a gentle decline.

That also explains why the count falls so dramatically at the step. A new maximum has just been reached, so almost nothing attains it yet; the four eleven-square attainers are four shapes in 34,053, which is 0.012 per cent. The collapse at a step is real and it is a fact about a new record being new, not a fact about plateaus.

What this leaves the sweep with

A description, not a detector. The rung below's proposed signal split into its component claims, with what the sweep says about each.
Fig. 6 The rung below’s proposal split into the claims it actually contains, each scored. The two descriptive claims come out at five of five and three of four; the two predictive ones have no threshold at all. The gap between the columns is the difference between reading a plateau and detecting one, and it is the difference the rung below wanted to close.

The difficulty the diagnostic was meant to remove is still there in full. A sweep at eleven squares cannot tell whether it is on a plateau or at a step, and the only way to find out is to compute twelve — which is roughly four times eleven’s 34,053 shapes with a costlier evaluation on each, and is exactly the cost the rung below could not afford and named as its own recorded shortfall.

That is worth stating without softening, because a negative result about a diagnostic is easy to write as though it were nearly positive. It is not nearly positive. There is no cut, the orbit repair ties rather than inverts, and the growth the whole idea rested on is the catalogue’s growth rather than the temperature’s.

What survives is a description worth having. Given a completed sequence of maxima and attaining counts, the counts say where the plateaus were and where the steps were, and they say it more legibly than the maxima do — a plateau shows as a rising run and a step as a collapse. That is genuinely useful for reading a sweep somebody else has run, and it is not what the rung below hoped for.

The counting also cost nothing, which is the one comfort. Every number on this page was already in the sweep the rung below ran — the attaining counts were printed in its own figures, in a column headed shapes attaining it — so testing the diagnostic required no new evaluation of anything, and the answer arrived for the price of arranging numbers already on disk. A negative result at that price is a good trade even when it is a flat negative.

There is a second thing the sweep says and it is not about counting at all. The eleven-square attainers are all of one kind — lopsided fights, with stops of 33 and 1/2-1/2, a mean of 5/45/4 and a temperature of 7/47/4 — where three of the five eight-square orbits are symmetric fights worth {3/23/2}\{3/2 \mid -3/2\}. The record changed hands from a shape where both players have much to gain to one where the gap between the two players’ prospects is large. Whether that is the pattern at every step, or a fact about eleven squares, is a question about the kind of shape rather than the number of them, and it is a question the counting never had access to.

The twelfth size, and what it would cost

The reason a plateau detector was worth wanting is arithmetic rather than curiosity.

The catalogue of connected shapes grows by a factor of about three and a half per square: 730 at eight, 2,542 at nine, 9,287 at ten, 34,053 at eleven. Eleven squares took forty-six seconds of evaluation; twelve is roughly four times as many shapes, and the evaluation of each is dearer because the game tree is deeper. The obstacle was the catalogue is where that cost was first identified as the thing limiting this anchor, and it has been the limit at every rung since.

A detector would have bought the sweep the right to stop. If the counts at eleven had said this is a step, the next size will hold, the sweep could have declared 7/47/4 a plateau value and moved on. They cannot say it, so the only way to know whether 7/47/4 is itself a plateau is to compute twelve — which is precisely the shortfall the rung below recorded and precisely what this rung was supposed to remove.

So the practical answer to is 7/47/4 a plateau? is unchanged: nobody knows, and the four attainers at eleven squares are consistent with both answers. Four attainers is what a new record looks like on its first size whether or not it is about to be beaten, because a record’s first size always has few attainers — that is what makes it a record.

There is one thing the counts do license, and it is small. Any size whose maximum has just risen has a small attaining count, at every step in the sweep and by an argument rather than by measurement. So a size with a large attaining count is certainly not at a step, which is the contrapositive of half the rung below’s sentence and is the only implication that survives. It is worth almost nothing operationally — the largest size computed always has the fewest attainers when it has just risen, and that is the case one wants to decide.

A count that was really a population

This is not the first time on this site a count has looked like a signal and turned out to be a fact about how many objects there were, and the recurrence is worth naming because the two cases were caught in opposite ways.

What a game actually produces is the nearest one: the distribution of temperatures over a catalogue is not the distribution over a game, because the catalogue weights every shape equally and a game does not. The correction there was to reweight. Here there is nothing to reweight — the question is about the catalogue, and the catalogue’s growth is the whole of the effect.

The general shape is that a count over a growing population is almost never the quantity anybody means. What the rung below meant was how crowded is the region just below the maximum — a density — and it reached for a count because the count was in the table. Densities and counts agree when the population is fixed and come apart when it triples every step, and this population triples and a half.

The share figure above is the density, computed, and it declines steadily across the whole sweep with no feature at any plateau or step. So the quantity the argument wanted does not carry the signal either, and the count only appeared to because the population was doing the work.

What the solver computed, and how

Every connected polyomino of at most eleven squares up to translation — 34,053 of them, with 9,287 at ten squares and 730 at eight. Each is treated as a Domineering region: Left places vertical dominoes and Right horizontal ones, on that region’s squares alone.

Each region’s game value is computed by the ordinary recursion over placements, reduced to canonical form, and its temperature read off the thermograph as the height at which the two walls meet. The maximum over each size is taken exactly, in the dyadic rationals the values actually live in, so attaining the maximum is an exact equality rather than a comparison against a tolerance — which matters, because the whole question is how many shapes hit a number precisely.

The orbit count groups the attaining shapes under the eight symmetries of the square, so a shape and its mirror count once.

The plateau structure is then read mechanically: a run of consecutive sizes sharing a maximum is a plateau, the size after it is a step, and the attaining count is checked for monotonicity across each run. The threshold test is exhaustive rather than a fit — every size whose maximum then held is compared against every size whose maximum then rose, and any pair where the first has at least as many attainers as the second is an inversion no cut can survive. The same comparison is run on the orbit counts.

Two things are asserted rather than reported. The growth must be monotone across every plateau, or the description this page keeps is wrong. And there must be at least one inversion on both counts, or some threshold exists and the page is written about the opposite result. A sweep in which the counts were merely noisy would fail both assertions and support neither half of the page.

Where the model stops

Eleven squares, and the whole difficulty is that twelve is where the question is. The sweep has five plateaus and four steps in it, which is a small number of events on which to test a rule about plateaus and steps, and the two smallest plateaus involve counts of one, two, three and four — where a difference of one is the whole signal and nothing can be read reliably. The honest statement of the negative result rests on the two largest plateaus and the seven inversions, not on all five.

Domineering regions, and this ruleset only. Nothing here says whether the same diagnostic works on Clobber or Toads and Frogs, and the mechanism identified — that the count grows because the catalogue grows — is not specific to Domineering, so the negative result probably transfers and has not been checked.

Normal play throughout, and temperature is a normal-play quantity: the thermograph is built from the recursion that this convention makes well-founded.

And the figures cannot show what would actually settle the question, which is a twelfth size. Every table here is a table of counts over a completed sweep, and the claim that no threshold exists is a claim about eleven data points. A twelfth size with a large attaining count and no rise would strengthen it and a twelfth with a small one and a rise would strengthen it; a twelfth that behaved the other way would not rescue the rule, because the inversions already present cannot be removed by adding data.

Where the ladder goes next

The cold anchor has nine rungs: below zero, how hot a day gets, how hot a real position is, what a game actually produces, one fight makes a board a fight, the obstacle that was the catalogue, what the ceiling is a fact about, that it was not a ceiling, and now what the counting is worth.

The rung above is the kind rather than the count. The four eleven-square attainers are lopsided fights and the eight-square attainers are mostly symmetric ones, and that is a statement about what a hot region looks like rather than about how many there are — which is the sort of thing that could genuinely predict, because it is about the mechanism the temperature comes from. The measurement is available on the sweep already built: classify every attaining shape at every size as symmetric or lopsided by comparing its two stops, and ask whether a size whose attainers are all lopsided behaves differently from one whose attainers are mixed. If the record changes hands from symmetric to lopsided at every step, the sweep has a signal that survives being normalised, because it is not a count of anything.

Two neighbours are worth the trip. Eight squares, and no hotter is where the plateau was mistaken for a bound, and it is worth reading beside a page about the diagnostic that would have caught it — because the diagnostic would not have caught it. And how hot a real position is is where the temperatures of positions a game actually reaches were measured, which is the reminder that the hottest region of a size is a fact about the catalogue rather than about anything a player will meet.

Part 9 of 9

One argument about Cold. The parts either side of it:

The objects named here

The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.

CoolingDecompositionDomineeringEnumerationExhaustive searchMean valueNormal playStopsTemperatureThermograph