Big is not the same as hot
Assumes: Playing the hottest · A rule that is never right and cannot be far wrong
Four essays on this site state the same rule: move in the hottest component. It is the practical content of the whole temperature field, and a bound instead of an answer is the guarantee that comes with it — playing that way rather than optimally costs at most the largest temperature on the board.
Players do not use that rule. They use a different one, they have used it for centuries, and it is the number printed in every endgame book.
The two are not variants of one idea. They measure different things, and on the positions a beginner meets they happen to give the same order — which is the worst possible arrangement, because it means the disagreement never shows up until it costs something.
Two ways of sizing a move
What changes hands. Look at where the position lands if one side takes the move and where it lands if the other does, and take the difference. Go players call this deiri counting. It is the natural quantity: it is how much the board changes.
What is at stake. The temperature, read off a thermograph. Half the swing, for a simple switch.
That factor of two is why the disagreement is invisible in the easy case, and it is why a pool of plain switches proves nothing.
Where they come apart
They separate on a component with a follow-up — a position whose options are themselves contested.
Take {5 | {4 | 0}}. Left moving takes it to 5. Right moving takes it to {4 | 0}, whose mean is 2. So three points change hands, deiri counting says the move is worth 3, and the temperature is 1 — not a half, because the wall is held up by an option that is itself a fight.
Now compare with {6 | 0}: six points change hands and the temperature is 3.
The count ranks {6 | 0} first and so does the temperature, so far so good. But {10 | {9 | 1}} has five points changing hands and a temperature of 1, so the count ranks it second and the temperature ranks it last.
Two components are enough to make the disagreement, and running the experiment on exactly two is worth doing before running it on eight, because it shows how little is needed.
So the failure the rest of this essay measures over a hundred and twenty boards is available on four, which is worth knowing: it is not a rare interaction that a large sweep dredges up, it is what these two quantities do whenever they are pointed at a component with a follow-up.
A Go player has a name for this and it is sente: a move whose follow-up is nearly as large as the move itself, so the opponent must answer and the mover keeps the initiative. Deiri counting overstates a sente move, because it counts a swing the mover will not have to pay for.
The overstatement is systematic rather than random, and its direction is the damaging one. A rule that sometimes overrates and sometimes underrates a component would produce errors that partly cancel; this one always overrates the same class of component, so a player following it plays sente moves too early, every time, in every position containing one.
The experiment
Both rules are run as players against a solver, over every board of three components drawn from a pool built to separate them.
The last row is the one that matters and it is worth being exact about why.
Hottest-first has a guarantee: it never loses more than the largest temperature on the board. That is a theorem, it is checked here over every line, and it holds on all 240. Biggest-first has no such guarantee and breaks that bound six times, giving up as much as three points where the largest temperature is two.
So the difference between the two rules is not that one is slightly better on average. One of them comes with a proof and the other does not, and the proof is exactly what is lost by counting the wrong quantity.
What the six lines look like
Six lines out of 240 is a small number and it is worth saying what has to happen for one to occur, because the mechanism is the essay’s argument in miniature.
The count has to prefer a sente component — one with a large swing and a small temperature — over a gote component with a smaller swing and a larger temperature. Playing the sente move first is not merely suboptimal; it wastes the move, because the opponent answers, the position returns to something similar, and the large gote component is still sitting there for the opponent to take.
The result is that the counting player spends a turn achieving nothing and then loses the big component, which is a loss of about the gote component’s swing rather than of the difference between the two. That is how a rule that is only slightly wrong about the ordering can be badly wrong about the outcome, and it is why the loss exceeds the largest temperature.
Hottest-first cannot do this. It takes the gote component first by construction, and the guarantee follows: whatever it gives up by not searching, it gives up in the small change rather than in a whole exchange.
Six lines in two hundred and forty is a rate, and a rate depends on what the pool is made of. Build a pool of exactly the shape described above — one plain gote component against two sente ones — and the same failure stops being rare.
Why temperature is the right quantity
The reason is what a temperature is: not how much changes hands, but how much a player should be willing to pay to move here rather than somewhere else.
Those are different because of what happens after. A move with a large follow-up will be answered, so the mover does not keep the whole swing; a move with no follow-up is finished, so the mover keeps all of it. Temperature accounts for the answer and deiri counting does not.
There is a case that makes the point from the other side, and it is the one that stops this from being a story about a rule that is simply worse. Take a pool of nothing but sente components, all four with a temperature of one and swings running from three to five. The two rules then disagree constantly, because the swings differ and the temperatures do not — and the disagreement costs nothing at all.
Nor can the two components be added up and compared as a total, which is the other thing a reader tries. Means add and temperatures do not: a board holding and settles at seven, which is four plus three exactly, while its temperature comes out at three — the hotter of the two, not the sum of them and not their average. So there is no aggregate to rank by. The ordering is a comparison between components and has to stay one, which is the structural reason a rule is wanted here rather than a formula.
The two quantities, side by side
It is worth tabulating the pool once, because the numbers are what the argument is.
| position | swing | temperature | ratio |
|---|---|---|---|
| {6 | 0} | 6 | 3 | 1 |
| {5 | 1} | 4 | 2 | 1 |
| {5 | {4 | 0}} | 3 | 1 | 1.5 |
| {7 | {6 | 0}} | 4 | 1 | 2 |
| {10 | {9 | 1}} | 5 | 1 | 2.5 |
The ratio column is swing divided by twice the temperature, so a plain switch scores 1 and anything else scores more. Everything with a follow-up scores above 1, and the further above, the more the count overstates the move.
A reader might reasonably ask why the count is not simply divided by the ratio. The answer is that the ratio is computed from the temperature, so there is nothing to be gained: knowing enough to correct the count means already knowing the answer. The correction is not an adjustment to deiri counting; it is a different measurement.
That is worth saying because it is the general shape of the relationship between a folk heuristic and a theory. The theory does not usually supply a patch for the heuristic. It supplies the quantity the heuristic was approximating, and once that quantity is available the heuristic has nothing left to do.
The count is not stupid
It is worth resisting the reading that players have been doing it wrong for centuries.
Deiri counting is the right number for a gote move — one that ends the exchange, where nobody answers, and the mover keeps everything. Every plain switch is gote, and plain switches are most positions. The count is exactly right there and half the temperature’s scale, which is a difference of units rather than of ranking.
What the count needs is a correction for sente, and Go players have one: a sente move is conventionally counted differently, by its follow-up rather than by its swing, and strong players apply the correction without deriving it. The correction is temperature. The theory’s contribution is not a new rule; it is the observation that one quantity handles both cases and that the case analysis is unnecessary.
That is also where the temperature’s authority comes from, and it is the fact that makes the two quantities incomparable rather than merely different. A swing is a measurement of one exchange: this much changes hands, once. A temperature is the error term in a long-run rate — pile up copies of a component and the pile stays within one temperature of times its mean, however large becomes. It was never a claim about a single move, which is precisely why the sente case does not disturb it: a quantity defined by what a hundred copies do cannot be embarrassed by what happens after one of them is played.
The count Go already has
Saying that players count the swing is only half of what players do, and the other half is close enough to temperature that the gap is worth measuring precisely rather than gestured at.
Alongside deiri counting there is miai counting, which asks not how much changes hands but how much changes hands per move spent. A plain switch takes one move from each side to resolve, so its two moves are worth half its swing each, and half the swing of a plain switch is exactly its temperature. On the whole class of positions where players use miai counting, the folk number and the theory’s number are the same number.
That is a stronger reconciliation than the earlier sections suggest, and it changes what the disagreement is about. It is not that players measure the wrong thing. It is that they measure the right thing with a divisor they have to choose, and choosing the divisor is where the case analysis enters.
For a sente move the divisor is not two. The convention is that a sente move costs its player nothing — the opponent answers, the initiative comes back — so the exchange is not two moves but zero net moves, and dividing a swing by zero is not a calculation. What Go practice does instead is count the follow-up: the size of the threat that makes the answer compulsory, which is the quantity that will actually be contested if the answer never comes.
So the practice is two procedures and a test to decide between them. Gote: halve the swing. Sente: measure the follow-up. And the test is whether the opponent must answer.
What the case analysis costs
The test is the expensive part, and it is expensive in a way that is invisible while it keeps working.
Whether the opponent must answer is not a property of the component. It depends on what else is on the board — sente is a fact about the rest of the board, and the same local shape is sente beside a quiet board and gote beside a hot one. So the divisor a player needs in order to size a move correctly is a function of the sizes of all the other moves, which are the quantities the sizing was supposed to produce.
That circularity is the real difference between the folk account and the theory, and it is sharper than the ranking errors the sweep counts. A player using the two-procedure method has to guess which procedure applies before they have the numbers that would tell them, and a wrong guess does not produce a slightly wrong size — it produces a size computed from the wrong quantity altogether, off by the whole follow-up.
Temperature has no such step. The thermograph is built from the component alone, the walls bend where the follow-ups take over, and the number read off the top is correct whether the ambient board turns out hot or cold. The case analysis has not been solved; it has been absorbed into the shape of the graph, and the crossing point where a move stops being urgent — the one the earlier figure draws — is exactly the place the folk method needs a decision and the theory needs nothing.
This is why one rule carries a bound and the other cannot. A guarantee has to hold on every board, and a procedure whose first step is a judgement about the board cannot be guaranteed without first bounding the errors in that judgement. There is no such bound available, because a component’s sente-ness can be flipped by a single distant move.
None of which makes the folk method a poor way to play. It makes it a method whose correctness is carried by the player rather than by the rule, which is the ordinary situation with expertise and the exact thing a theorem is for. The six lines where the count exceeds the bound are what that looks like when the judgement is removed and the procedure is run mechanically — which is the only way a rule can be measured at all.
The bound is the whole difference
Restating the finding as sharply as it goes:
Hottest-first loses at most the largest temperature on the board. Proved, and checked here on 240 lines with no exception.
Biggest-first loses more than that on six of the same 240, and there is no bound anybody can attach to it.
A heuristic with a bound and a heuristic without one are different kinds of object even when their average performance is similar. The first can be used inside a larger argument — a program that plays hottest-first can state what it might be giving up — and the second cannot.
That distinction is the one a bound instead of an answer is about, and it is the shape of every honest approximation in this subject: not “usually close”, which is a report about a sample, but “never worse by more than x”, which is a statement about every position including the ones nobody has tried. The whole reason to define temperature rather than to measure swings is that the second kind of statement can be made about it.
The distinction is not visible in an average and it is not visible in a typical position, which is the last thing this essay has to establish and the reason the pool was built rather than sampled.
The same finding in three other games
Three other essays here make a measurement of exactly this shape, and the pattern they share is worth naming once.
In Dots and Boxes the rule is take every box available, and it is exactly right on boards too small to hold a chain worth declining. In a pawn ending the rule is count the spare tempo moves, and it is exactly right when both sides have the same options in every file. In Amazons the rule is count territory, and it is exactly right when every region is settled. Here the rule is play the biggest move, and it is exactly right when every component is a plain switch.
Four folk rules, four exact domains, and four failures outside them. The domains are different and the reason is the same in every case: the rule computes a number and the position’s answer stops being a number — at the moment a chain becomes worth declining, at the moment the two players’ options differ, at the moment a region becomes contested, at the moment a component acquires a follow-up.
That is as close to a general principle as this field has produced, and it is worth stating as a test rather than as an observation. Given a rule of thumb about a game, ask: what is the class of positions on which the game’s value is a number? The rule will be right there and nowhere reliable.
What the picture cannot show
The pool here is eight abstract components and the boards are three at a time.
That is an artificial setting and it is chosen deliberately: a pool of real Go endgame positions would be more convincing and much less controlled, and the point of the experiment is to isolate the quantity under test. The pool is built from components with follow-ups precisely because a representative pool would be dominated by plain switches and would report agreement that was arranged in advance.
Running the same experiment on a pool chosen to be ordinary rather than to separate the rules says how much of the disagreement was arranged.
So what these figures establish is that the two rules can differ, by how much, and which of the two carries a guarantee. What they cannot establish is how often the difference arises in a real game, which depends on how many sente moves a real endgame holds and is a question about Go rather than about the theory. The pool of ten above is a better guess at that rate than the pool of eight and is still a guess: it is a pool of abstract components chosen by hand, and a Go board decides for itself how many of its regions have follow-ups.
The second thing not shown: the thermographic strategy, which is better than both rules here. Playing by temperature is a simplification of what the theory can actually recommend — the full account uses cooling and the orthodox accounting that the Go endgame essay is about — and comparing that with either rule would need a third player and a longer essay.
The convention, named
Everything here is normal play, and the components are all hot — positions both players want to move in.
That matters because the whole notion of “where to move” presupposes that moving is good. In a position with a zugzwang, moving is bad, the temperature is negative, and both rules above are answering the wrong question. What a value leaves out is the essay about the quantity that takes over when the temperature runs out, and it is the same quantity a Dots and Boxes endgame is fought over.
So the honest scope of this essay is: among hot components, temperature is the right size and swing is not. Once the components go cold the ordering stops mattering and the parity of the remaining moves takes over.
Where the ladder goes next
temperature reaches six rungs with this one, and the next is the one the last section named.
It is the thermographic move rule — cooling every component by the same amount and playing where the cooled position still has a move — which is what Berlekamp’s Go endgame accounting actually does and which is strictly better than playing the hottest. That would need a third player in the experiment above, and it is the natural next measurement rather than a new subject.
Part 6 of 8
One argument about Temperature. The parts either side of it:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 18.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.
Disjunctive sumGo endgameGreedy playHeuristicHot gameMean valueMove selectionSenteSwitchTemperatureThermograph
- A rule with a guarantee disjunctive sum, greedy play, heuristic, hot game, mean value, move selection, sente, temperature, thermograph
- A rule with no promise at all disjunctive sum, greedy play, heuristic, hot game, mean value, move selection, sente, temperature
- An environment made of coupons disjunctive sum, go endgame, hot game, mean value, sente, switch, temperature, thermograph
- The answer that starts another fight go endgame, hot game, mean value, move selection, sente, switch, temperature, thermograph
- A pool built to punish greed disjunctive sum, greedy play, heuristic, mean value, move selection, sente, temperature
- What a move nobody makes is worth disjunctive sum, hot game, mean value, move selection, sente, temperature, thermograph