Temperature

Big is not the same as hot

A player sizes a move by how much changes hands when it is played, which is the number in every endgame book. The theory sizes it by temperature. On a plain switch the two agree exactly, so nothing shows; on a move with a follow-up they come apart, and the count gives up more than the guarantee the theory's rule carries.

Assumes: Playing the hottest · A rule that is never right and cannot be far wrong

Four essays on this site state the same rule: move in the hottest component. It is the practical content of the whole temperature field, and a bound instead of an answer is the guarantee that comes with it — playing that way rather than optimally costs at most the largest temperature on the board.

Players do not use that rule. They use a different one, they have used it for centuries, and it is the number printed in every endgame book.

The two are not variants of one idea. They measure different things, and on the positions a beginner meets they happen to give the same order — which is the worst possible arrangement, because it means the disagreement never shows up until it costs something.

Two ways of sizing a move

What changes hands. Look at where the position lands if one side takes the move and where it lands if the other does, and take the difference. Go players call this deiri counting. It is the natural quantity: it is how much the board changes.

What is at stake. The temperature, read off a thermograph. Half the swing, for a simple switch.

the switch {5 | 1}. Temperature runs up the page and value across it. Each wall is where a player is willing to move once a tax of that much is charged per move; above the temperature at which they meet, neither wants to move and the position is worth its mean value. The height of the meeting point is what is at stake.
Fig. 1 The switch {5 | 1}. Left moving takes it to 5 and Right to 1, so four points change hands and the temperature is two. On a position of this shape the two rules are the same number up to a factor of two, so they always rank a set of them identically and there is nothing to compare.

That factor of two is why the disagreement is invisible in the easy case, and it is why a pool of plain switches proves nothing.

How much changes hands, against how much is at stake. Two ways of choosing where to move, run against optimal play over every board from a pool of three components. Biggest-first takes the component where the most changes hands, which is the count in every endgame book; hottest-first takes the one with the highest temperature. They disagree on most of these boards, the count costs points more often, and — the difference that matters — the count sometimes loses more than the largest temperature on the board, which is the bound the theory's rule is guaranteed to keep.
Fig. 2 The experiment run on plain switches only: the two rules pick the same component every time, so both do equally well and the figure has nothing to say. Every comparison of the two rules that uses simple positions is this figure.

Where they come apart

They separate on a component with a follow-up — a position whose options are themselves contested.

Take {5 | {4 | 0}}. Left moving takes it to 5. Right moving takes it to {4 | 0}, whose mean is 2. So three points change hands, deiri counting says the move is worth 3, and the temperature is 1 — not a half, because the wall is held up by an option that is itself a fight.

Now compare with {6 | 0}: six points change hands and the temperature is 3.

The count ranks {6 | 0} first and so does the temperature, so far so good. But {10 | {9 | 1}} has five points changing hands and a temperature of 1, so the count ranks it second and the temperature ranks it last.

Two components are enough to make the disagreement, and running the experiment on exactly two is worth doing before running it on eight, because it shows how little is needed.

How much changes hands, against how much is at stake. Two ways of choosing where to move, run against optimal play over every board from a pool of three components. Biggest-first takes the component where the most changes hands, which is the count in every endgame book; hottest-first takes the one with the highest temperature. They disagree on most of these boards, the count costs points more often, and — the difference that matters — the count sometimes loses more than the largest temperature on the board, which is the bound the theory's rule is guaranteed to keep.
Fig. 3 The whole apparatus on a pool of two: {51}\{5 \mid 1\}, a plain switch of swing four and temperature two, and {10{91}}\{10 \mid \{9 \mid 1\}\}, a sente shape of swing five and temperature one. Four boards, eight lines of play, and the two rules already choose differently on four of them — the count prefers the sente component because five is more than four, and the temperature prefers the switch because two is more than one. The count loses something on one line of the eight, and that one line is also a line on which it loses more than the largest temperature on the board.

So the failure the rest of this essay measures over a hundred and twenty boards is available on four, which is worth knowing: it is not a rare interaction that a large sweep dredges up, it is what these two quantities do whenever they are pointed at a component with a follow-up.

A Go player has a name for this and it is sente: a move whose follow-up is nearly as large as the move itself, so the opponent must answer and the mover keeps the initiative. Deiri counting overstates a sente move, because it counts a swing the mover will not have to pay for.

The overstatement is systematic rather than random, and its direction is the damaging one. A rule that sometimes overrates and sometimes underrates a component would produce errors that partly cancel; this one always overrates the same class of component, so a player following it plays sente moves too early, every time, in every position containing one.

The experiment

Both rules are run as players against a solver, over every board of three components drawn from a pool built to separate them.

How much changes hands, against how much is at stake. Two ways of choosing where to move, run against optimal play over every board from a pool of three components. Biggest-first takes the component where the most changes hands, which is the count in every endgame book; hottest-first takes the one with the highest temperature. They disagree on most of these boards, the count costs points more often, and — the difference that matters — the count sometimes loses more than the largest temperature on the board, which is the bound the theory's rule is guaranteed to keep.
Fig. 4 A hundred and twenty boards, played out twice each — once with Left moving first and once with Right. The two rules pick different components on 104 of the 240 lines; the count is strictly worse on 45 against the temperature rule’s 33; and on six lines the count loses more than the largest temperature on the board, which is the bound the theory’s rule never breaks.

The last row is the one that matters and it is worth being exact about why.

Hottest-first has a guarantee: it never loses more than the largest temperature on the board. That is a theorem, it is checked here over every line, and it holds on all 240. Biggest-first has no such guarantee and breaks that bound six times, giving up as much as three points where the largest temperature is two.

So the difference between the two rules is not that one is slightly better on average. One of them comes with a proof and the other does not, and the proof is exactly what is lost by counting the wrong quantity.

A rule that is not optimal, and cannot be far wrong. Every line of play in the experiment as one point: the largest temperature on the board across, and what playing hottest-first cost against optimal play up. The diagonal is the bound the theory proves. Nothing gold sits above it; the magenta series is the same rule with its ordering reversed, and it sits above the line often — which is what makes the bound a claim rather than a description.
Fig. 5 The guarantee itself, measured on the pool the bound was written for: hottest-first never exceeds the largest temperature, and coolest-first — the same machinery with the ordering reversed — breaks it repeatedly. A bound that nothing can break is a bound about nothing, which is why the reversed rule is in the experiment.

What the six lines look like

Six lines out of 240 is a small number and it is worth saying what has to happen for one to occur, because the mechanism is the essay’s argument in miniature.

The count has to prefer a sente component — one with a large swing and a small temperature — over a gote component with a smaller swing and a larger temperature. Playing the sente move first is not merely suboptimal; it wastes the move, because the opponent answers, the position returns to something similar, and the large gote component is still sitting there for the opponent to take.

The result is that the counting player spends a turn achieving nothing and then loses the big component, which is a loss of about the gote component’s swing rather than of the difference between the two. That is how a rule that is only slightly wrong about the ordering can be badly wrong about the outcome, and it is why the loss exceeds the largest temperature.

Hottest-first cannot do this. It takes the gote component first by construction, and the guarantee follows: whatever it gives up by not searching, it gives up in the small change rather than in a whole exchange.

Six lines in two hundred and forty is a rate, and a rate depends on what the pool is made of. Build a pool of exactly the shape described above — one plain gote component against two sente ones — and the same failure stops being rare.

How much changes hands, against how much is at stake. Two ways of choosing where to move, run against optimal play over every board from a pool of three components. Biggest-first takes the component where the most changes hands, which is the count in every endgame book; hottest-first takes the one with the highest temperature. They disagree on most of these boards, the count costs points more often, and — the difference that matters — the count sometimes loses more than the largest temperature on the board, which is the bound the theory's rule is guaranteed to keep.
Fig. 6 Ten boards drawn from one gote component and two sente ones, twenty lines of play. The rules disagree on half of them, the count gives something up on three, and on all three of those it loses more than the largest temperature on the board. That is a breach rate of three in twenty against six in two hundred and forty, from a pool built out of the shape the mechanism above names — which is what it means for a mechanism to be identified rather than guessed at.

Why temperature is the right quantity

The reason is what a temperature is: not how much changes hands, but how much a player should be willing to pay to move here rather than somewhere else.

Those are different because of what happens after. A move with a large follow-up will be answered, so the mover does not keep the whole swing; a move with no follow-up is finished, so the mover keeps all of it. Temperature accounts for the answer and deiri counting does not.

There is a case that makes the point from the other side, and it is the one that stops this from being a story about a rule that is simply worse. Take a pool of nothing but sente components, all four with a temperature of one and swings running from three to five. The two rules then disagree constantly, because the swings differ and the temperatures do not — and the disagreement costs nothing at all.

How much changes hands, against how much is at stake. Two ways of choosing where to move, run against optimal play over every board from a pool of three components. Biggest-first takes the component where the most changes hands, which is the count in every endgame book; hottest-first takes the one with the highest temperature. They disagree on most of these boards, the count costs points more often, and — the difference that matters — the count sometimes loses more than the largest temperature on the board, which is the bound the theory's rule is guaranteed to keep.
Fig. 7 Twenty boards built from four components that all have a temperature of one. The two rules pick the same component on only twelve of the forty lines, so they are disagreeing more often than not — and neither of them gives up a single point anywhere, because when the stakes are equal every order is as good as every other. A disagreement is not an error. The count is wrong about these components’ sizes and the wrongness has nothing to bite on, which is why a comparison of the two rules has to be run on a pool where the temperatures differ.

Nor can the two components be added up and compared as a total, which is the other thing a reader tries. Means add and temperatures do not: a board holding {5{40}}\{5 \mid \{4 \mid 0\}\} and {60}\{6 \mid 0\} settles at seven, which is four plus three exactly, while its temperature comes out at three — the hotter of the two, not the sum of them and not their average. So there is no aggregate to rank by. The ordering is a comparison between components and has to stay one, which is the structural reason a rule is wanted here rather than a formula.

The two quantities, side by side

It is worth tabulating the pool once, because the numbers are what the argument is.

position swing temperature ratio
{6 | 0} 6 3 1
{5 | 1} 4 2 1
{5 | {4 | 0}} 3 1 1.5
{7 | {6 | 0}} 4 1 2
{10 | {9 | 1}} 5 1 2.5

The ratio column is swing divided by twice the temperature, so a plain switch scores 1 and anything else scores more. Everything with a follow-up scores above 1, and the further above, the more the count overstates the move.

A reader might reasonably ask why the count is not simply divided by the ratio. The answer is that the ratio is computed from the temperature, so there is nothing to be gained: knowing enough to correct the count means already knowing the answer. The correction is not an adjustment to deiri counting; it is a different measurement.

That is worth saying because it is the general shape of the relationship between a folk heuristic and a theory. The theory does not usually supply a patch for the heuristic. It supplies the quantity the heuristic was approximating, and once that quantity is available the heuristic has nothing left to do.

The count is not stupid

It is worth resisting the reading that players have been doing it wrong for centuries.

Deiri counting is the right number for a gote move — one that ends the exchange, where nobody answers, and the mover keeps everything. Every plain switch is gote, and plain switches are most positions. The count is exactly right there and half the temperature’s scale, which is a difference of units rather than of ranking.

What the count needs is a correction for sente, and Go players have one: a sente move is conventionally counted differently, by its follow-up rather than by its swing, and strong players apply the correction without deriving it. The correction is temperature. The theory’s contribution is not a new rule; it is the observation that one quantity handles both cases and that the case analysis is unnecessary.

That is also where the temperature’s authority comes from, and it is the fact that makes the two quantities incomparable rather than merely different. A swing is a measurement of one exchange: this much changes hands, once. A temperature is the error term in a long-run rate — pile up nn copies of a component and the pile stays within one temperature of nn times its mean, however large nn becomes. It was never a claim about a single move, which is precisely why the sente case does not disturb it: a quantity defined by what a hundred copies do cannot be embarrassed by what happens after one of them is played.

The count Go already has

Saying that players count the swing is only half of what players do, and the other half is close enough to temperature that the gap is worth measuring precisely rather than gestured at.

Alongside deiri counting there is miai counting, which asks not how much changes hands but how much changes hands per move spent. A plain switch takes one move from each side to resolve, so its two moves are worth half its swing each, and half the swing of a plain switch is exactly its temperature. On the whole class of positions where players use miai counting, the folk number and the theory’s number are the same number.

That is a stronger reconciliation than the earlier sections suggest, and it changes what the disagreement is about. It is not that players measure the wrong thing. It is that they measure the right thing with a divisor they have to choose, and choosing the divisor is where the case analysis enters.

For a sente move the divisor is not two. The convention is that a sente move costs its player nothing — the opponent answers, the initiative comes back — so the exchange is not two moves but zero net moves, and dividing a swing by zero is not a calculation. What Go practice does instead is count the follow-up: the size of the threat that makes the answer compulsory, which is the quantity that will actually be contested if the answer never comes.

So the practice is two procedures and a test to decide between them. Gote: halve the swing. Sente: measure the follow-up. And the test is whether the opponent must answer.

What the case analysis costs

The test is the expensive part, and it is expensive in a way that is invisible while it keeps working.

Whether the opponent must answer is not a property of the component. It depends on what else is on the board — sente is a fact about the rest of the board, and the same local shape is sente beside a quiet board and gote beside a hot one. So the divisor a player needs in order to size a move correctly is a function of the sizes of all the other moves, which are the quantities the sizing was supposed to produce.

That circularity is the real difference between the folk account and the theory, and it is sharper than the ranking errors the sweep counts. A player using the two-procedure method has to guess which procedure applies before they have the numbers that would tell them, and a wrong guess does not produce a slightly wrong size — it produces a size computed from the wrong quantity altogether, off by the whole follow-up.

Temperature has no such step. The thermograph is built from the component alone, the walls bend where the follow-ups take over, and the number read off the top is correct whether the ambient board turns out hot or cold. The case analysis has not been solved; it has been absorbed into the shape of the graph, and the crossing point where a move stops being urgent — the one the earlier figure draws — is exactly the place the folk method needs a decision and the theory needs nothing.

This is why one rule carries a bound and the other cannot. A guarantee has to hold on every board, and a procedure whose first step is a judgement about the board cannot be guaranteed without first bounding the errors in that judgement. There is no such bound available, because a component’s sente-ness can be flipped by a single distant move.

None of which makes the folk method a poor way to play. It makes it a method whose correctness is carried by the player rather than by the rule, which is the ordinary situation with expertise and the exact thing a theorem is for. The six lines where the count exceeds the bound are what that looks like when the judgement is removed and the procedure is run mechanically — which is the only way a rule can be measured at all.

The bound is the whole difference

Restating the finding as sharply as it goes:

Hottest-first loses at most the largest temperature on the board. Proved, and checked here on 240 lines with no exception.

Biggest-first loses more than that on six of the same 240, and there is no bound anybody can attach to it.

A heuristic with a bound and a heuristic without one are different kinds of object even when their average performance is similar. The first can be used inside a larger argument — a program that plays hottest-first can state what it might be giving up — and the second cannot.

That distinction is the one a bound instead of an answer is about, and it is the shape of every honest approximation in this subject: not “usually close”, which is a report about a sample, but “never worse by more than x”, which is a statement about every position including the ones nobody has tried. The whole reason to define temperature rather than to measure swings is that the second kind of statement can be made about it.

The distinction is not visible in an average and it is not visible in a typical position, which is the last thing this essay has to establish and the reason the pool was built rather than sampled.

The same finding in three other games

Three other essays here make a measurement of exactly this shape, and the pattern they share is worth naming once.

In Dots and Boxes the rule is take every box available, and it is exactly right on boards too small to hold a chain worth declining. In a pawn ending the rule is count the spare tempo moves, and it is exactly right when both sides have the same options in every file. In Amazons the rule is count territory, and it is exactly right when every region is settled. Here the rule is play the biggest move, and it is exactly right when every component is a plain switch.

Four folk rules, four exact domains, and four failures outside them. The domains are different and the reason is the same in every case: the rule computes a number and the position’s answer stops being a number — at the moment a chain becomes worth declining, at the moment the two players’ options differ, at the moment a region becomes contested, at the moment a component acquires a follow-up.

That is as close to a general principle as this field has produced, and it is worth stating as a test rather than as an observation. Given a rule of thumb about a game, ask: what is the class of positions on which the game’s value is a number? The rule will be right there and nowhere reliable.

What the picture cannot show

The pool here is eight abstract components and the boards are three at a time.

That is an artificial setting and it is chosen deliberately: a pool of real Go endgame positions would be more convincing and much less controlled, and the point of the experiment is to isolate the quantity under test. The pool is built from components with follow-ups precisely because a representative pool would be dominated by plain switches and would report agreement that was arranged in advance.

Running the same experiment on a pool chosen to be ordinary rather than to separate the rules says how much of the disagreement was arranged.

How much changes hands, against how much is at stake. Two ways of choosing where to move, run against optimal play over every board from a pool of three components. Biggest-first takes the component where the most changes hands, which is the count in every endgame book; hottest-first takes the one with the highest temperature. They disagree on most of these boards, the count costs points more often, and — the difference that matters — the count sometimes loses more than the largest temperature on the board, which is the bound the theory's rule is guaranteed to keep.
Fig. 8 The same machinery over a pool of ten in which six components are plain switches and four have follow-ups — a spread closer to what a board actually offers. Two hundred and twenty boards, four hundred and forty lines, and the two rules now choose the same component on 422 of them. The count gives something up on 23 lines against the temperature rule’s 17, and it breaks the bound on none. The failure this essay is about survives, and its rate collapses.

So what these figures establish is that the two rules can differ, by how much, and which of the two carries a guarantee. What they cannot establish is how often the difference arises in a real game, which depends on how many sente moves a real endgame holds and is a question about Go rather than about the theory. The pool of ten above is a better guess at that rate than the pool of eight and is still a guess: it is a pool of abstract components chosen by hand, and a Go board decides for itself how many of its regions have follow-ups.

The second thing not shown: the thermographic strategy, which is better than both rules here. Playing by temperature is a simplification of what the theory can actually recommend — the full account uses cooling and the orthodox accounting that the Go endgame essay is about — and comparing that with either rule would need a third player and a longer essay.

The convention, named

Everything here is normal play, and the components are all hot — positions both players want to move in.

That matters because the whole notion of “where to move” presupposes that moving is good. In a position with a zugzwang, moving is bad, the temperature is negative, and both rules above are answering the wrong question. What a value leaves out is the essay about the quantity that takes over when the temperature runs out, and it is the same quantity a Dots and Boxes endgame is fought over.

So the honest scope of this essay is: among hot components, temperature is the right size and swing is not. Once the components go cold the ordering stops mattering and the parity of the remaining moves takes over.

Where the ladder goes next

temperature reaches six rungs with this one, and the next is the one the last section named.

It is the thermographic move rule — cooling every component by the same amount and playing where the cooled position still has a move — which is what Berlekamp’s Go endgame accounting actually does and which is strictly better than playing the hottest. That would need a third player in the experiment above, and it is the natural next measurement rather than a new subject.

Part 6 of 8

One argument about Temperature. The parts either side of it:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 18.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.

Disjunctive sumGo endgameGreedy playHeuristicHot gameMean valueMove selectionSenteSwitchTemperatureThermograph