The first time it told somebody something
Assumes: Playing the hottest · The endgame, accounted for
A mathematical theory of games is easy to admire and hard to justify. Zermelo’s theorem is true of chess and says nothing about chess; Sprague and Grundy settle every impartial game and no impartial game anybody plays for money. Every theorem in it can be true, every value correct, and the whole thing still a description of situations nobody was confused about.
Temperature theory is where that stopped being a fair complaint. Applied to Go endgames it produces a move order that is provably right, and the order is not the one strong players’ experience offers — which is the specific thing a theory has to do to earn its keep: tell somebody who already knows the subject something they did not know.
What a Go endgame is, as an object
Late in a game of Go the board has broken into regions that do not interact — the same thing that happens to a Domineering board when it falls apart. A move in one has no effect on any other. That is a disjunctive sum, exactly, and it is the situation the whole theory was built for.
Each region is worth something, in the sense this site’s values are worth something. Some are worth a definite number of points to whoever takes them. Some are worth more to whoever moves there first — those are switches, and the amount at stake in one is its temperature.
The player’s question is not which region is biggest. It is which region to move in, and those are different questions.
The rule, and why it is not obvious
The rule is: move in the hottest region.
Stated flatly it sounds like common sense — take the biggest thing first. It is not common sense, for two reasons.
It is about temperature and not about value. A region worth ten points to whoever takes it, where the opponent taking it also gains ten, is hotter than a region worth twenty points that only one player can use. The quantity to compare is what is at stake in the exchange, not the size of the prize.
And the intuitive alternative is defensible. Taking the largest secure gain first, or answering the opponent’s last move locally, are both what experience suggests, and both are wrong in cases the theory identifies exactly.
The sharpest way to see that the two quantities are independent is a board on which one of them has been set to zero everywhere. Four regions, each worth exactly nothing on average — whatever Left can gain in one, Right can gain back — and yet all four are worth taking, in a definite order, and the order decides the game.
The quantity, defined
Temperature is defined by an operation rather than by a measurement, and the definition is worth having in front of the argument.
Cooling a position by means: at every stage, a player who moves must pay for the privilege, and the payment is subtracted from what they gain. Cool a switch enough and it stops being worth fighting over — both players would rather the other moved — and the position collapses to a number.
The temperature is the value of at which that happens. It is the tax that makes the position not worth contesting, and it is therefore exactly what the position is worth contesting for.
That construction is why the comparison across regions is meaningful. Two regions’ temperatures are on the same scale — both measured in the same units of payment — so “hotter” is a comparison of like with like rather than a judgement.
What is proved, and what is not
Here the essay has to be precise, because the strong form of the claim is false.
Hottest-first is not optimal. It is a greedy rule and greedy rules generally are not.
What is proved is a bound, of the kind an approximation with a guarantee always is: playing the hottest component never loses by more than the temperature of the hottest component. That is a guarantee of a specific size, established by an argument rather than by trial, and it is the kind of statement a heuristic never provides.
The site’s own check makes both halves visible. Over 440 measured lines, hottest-first stays inside the bound on all of them — and is strictly worse than optimal on 17, so the bound is doing work rather than being a restatement of optimality. Playing coolest-first breaks the bound on 54.
The account itself has the same character, and it fails in a way worth seeing rather than describing. On regions that are plain fights between two numbers it is not an approximation at all — the sum of the means plus the alternating stakes is the exact value, every time. Add one region whose own options are contested and the equality goes.
That is the exact shape of the theory’s honesty. The account is a bookkeeping device that assumes each region is spent when it is played, and a region with a follow-up is not; the error is bounded by the same quantity the move rule’s error is bounded by, which is why one number governs both. It is also why the bound is stated against the hottest region rather than against a total: means add and temperatures do not, so there is no total temperature for a guarantee to be denominated in.
The surprise: the theory says when to stop obeying it
The part of the analysis that reads least like a mathematician’s contribution is the one about when the rule expires.
As the temperature falls the regions stop being contested, and what is left is arithmetic: a set of regions each worth a fixed number to whoever takes it, with the moves alternating. At that point the ordering question is over and the accounting question begins, and mixing them up is the mistake the theory identifies.
Orthodox accounting is the name for doing that bookkeeping correctly, and its content is that a player’s total is the sum of the region values plus a correction for the tempo — for who has the move when the temperature reaches zero. Tempo is a quantity in the account rather than an intangible.
The tempo term is the alternating column, and the cleanest demonstration of it is a board with nothing to choose between: three identical regions, so that the ordering question is vacuous and only the count of them matters.
So the tempo is worth one temperature, it is the last term to survive as the board cools, and which player collects it is decided by a count that no ranking of regions can see. That is why a Go player counts the number of remaining gote moves as well as their sizes, and why a single move early can be worth a whole exchange late.
What the demonstration actually was
It is worth being specific about the external validation, because “the theory beat professionals” is the kind of claim that inflates.
The positions were constructed: endgames built so that the board genuinely decomposes into independent regions, with the couplings that would spoil the analysis removed. That is not cheating — it is stating the hypothesis of the theorem and then satisfying it — but it is not a claim that the theory plays Go.
Within those positions the result was unambiguous. The analysis produces a move order; the order differs from what strong players choose; and playing the analysis wins games that playing the intuition loses. Berlekamp’s demonstrations included giving professional players the stronger side and beating them.
What that establishes is narrow and real. It is not that the theory is a Go program. It is that on positions satisfying its hypothesis, it produces provably correct play in cases where a career of experience produces incorrect play — which is exactly the standard by which a theory is said to have told somebody something.
The number that made it a result rather than a curiosity
The claim that separates this from a demonstration of correctness is about size: how much the right order is worth.
The error from playing the wrong region first is bounded by the temperature — a proved bound, and this site checks it over 440 lines. Whether that bound is large is a question about positions rather than about theory, and in Go endgames it is often the whole margin: a point or two, in games decided by half a point.
So the practical content is that a quantity nobody was computing turns out to govern a decision everybody was making by feel, and the amount at stake is the amount that decides games.
The arithmetic of a real endgame is dominated by one or two large regions with a tail of small ones, and it is worth seeing the account on a board of that shape, because it is where the practical claim lives: the big region is not what the game turns on.
Where the model stops
The positions checked here are small enough to solve completely, and that is the whole basis of the numbers on this page. A real Go endgame is not, and the theory’s application there is an approximation whose error the bound above governs.
And Go’s rules do not decompose perfectly. The theory needs the regions to be genuinely independent, and in Go they are only nearly so: ko fights, seki and shared liberties couple regions that look separate. The published applications are careful about this and the coupling is where practical difficulty lives.
What the picture cannot show
A thermograph is a picture of a value as a function of a cooling parameter, and it shows the value beautifully and the move not at all.
Two positions with identical thermographs can require different moves, because the thermograph records what the position is worth and not which option achieves it. Everywhere this site draws one, the move — if the move matters — has to be stated separately, and this is the standing example of the rule that a value is not a strategy.
Why the gap was twenty years
A theory built in 1970 and applied in 1990 invites the question of what took so long, and there are three answers, all of them ordinary.
The theory had to be finished first. Temperature, cooling, thermographs and the mean value are not the foundational part of Conway’s work; they sit on top of values, canonical forms and the sum, and those came first.
The application needed a decomposition somebody trusted. Go’s endgame decomposes only approximately, and establishing which positions decompose exactly enough is a piece of Go knowledge rather than a piece of mathematics. That is Berlekamp and Wolfe’s contribution as much as the analysis is.
And the audience had to overlap. The people who understood the theory and the people who understood Go endgames were, for two decades, almost disjoint sets.
None of that is a criticism of anybody. It is what applying a theory to a subject usually costs, and the interesting thing is that the cost was paid at all — most of this subject’s machinery has never been applied to anything outside itself.
The object at the centre of it is one a Go player recognises instantly and had no name for as a mathematical thing: a region worth more to whoever moves in it first, with a definite amount at stake. Everything above is that object counted, ordered and summed. The vocabulary was already there — big, urgent, sente, gote — and what arrived was the arithmetic underneath it.
What a player takes away, and what they do not
The practical residue of the theory for somebody playing is short, and it is worth separating from the apparatus.
Compare what is at stake, not what is gained. The quantity to rank regions by is the temperature — what moving there is worth relative to moving elsewhere — and not the value to either player of the region itself.
Play the hottest first, and expect to be within one temperature of optimal. Not within one swing: for a plain switch the swing is twice the temperature, so the bound stated in swings is twice as loose as the one that was proved.
And stop ranking when the temperatures reach zero, at which point count rather than compare.
What a player does not take away is a method for computing the temperatures, which is the entire content of the theory and requires the position to be small enough to evaluate. The rule is usable with estimated temperatures, and then it is a heuristic with no bound — which is what strong players were already doing, and is why the theory’s contribution is the cases where estimation and computation disagree.
The two rankings, and where they part
That correction is worth more than a footnote, because ranking by the swing is not a mistake anybody makes carelessly. It is the thing Go practice has always done, it agrees with the theory on almost every position a reader will test it on, and it is wrong exactly where the theory’s contribution lies.
On a plain switch — a region worth to whoever takes it and to whoever is left with it, and nothing deeper — the swing is and the temperature is half of that. A constant factor cannot change an ordering, so ranking a board of plain switches by swing and ranking it by temperature give the same order, every time. Every region in the figures on this page is of that kind.
They part on a region with a follow-up, where the opponent’s reply is itself a fight. There the swing counts a change of position the mover will not be allowed to keep — the opponent answers, and the answer is nearly as large as the move — while the temperature accounts for the answer and comes out much smaller. The two quantities then rank the same board differently, and a player following the swing plays those regions far too early.
And this is where the guarantee is lost rather than merely loosened. Big is not the same as hot runs both rules as players over a pool built to separate them, and finds the swing-ranked player losing more than the largest temperature on the board — the bound this essay is about — on six lines of 240. Not a worse average: no bound at all, because the quantity being ranked by is not the quantity the error is measured in.
Why the ranking quantity has to be the error
Which is the general point the essay’s last section reaches for and can now be stated exactly.
A greedy rule that is provably close needs the number it ranks by and the number that bounds its loss to be the same number. Rank by and lose at most , with and different quantities, and there is nothing to stop a position where is small and is large — which is precisely a sente region, small in temperature and large in swing.
Temperature satisfies the condition by construction. It is defined as what moving in a region is worth relative to moving elsewhere, so the cost of moving elsewhere instead is bounded by it, and the argument closes on itself. The swing satisfies nothing: it is a measurement of the region taken alone, made without reference to what a move elsewhere would have been worth, and there is no reason a quantity computed in isolation should bound an error that is about a comparison.
So the two paragraphs of practical advice above are not a simplification of the theory. They are the theory, and the version with the swing in it is a different rule that happens to agree on the easy cases — which is the most dangerous way for a wrong rule to be wrong, and the reason the correction is worth making in a section rather than in a word.
It also sharpens what the demonstrations established. Strong players were not ranking regions badly; they were ranking them by a quantity that is right whenever a region has no follow-up, and Go endgames are full of regions that do. The theory told them something in exactly the positions where the two quantities disagree, which is a much narrower and much more interesting claim than “the mathematics beat the professionals”.
Who found it, and when
Conway’s theory dates from around 1970 and was published in 1976. The application to Go is Elwyn Berlekamp’s, from the 1980s and 1990s, with David Wolfe — Mathematical Go is the book, and its subtitle, Chilling Gets the Last Point, is a claim about a specific quantity.
The gap is the point of filing this essay here: the machinery was built for other reasons and validated externally about two decades later. The regions, the temperatures and the cooling operator were all in place before anybody had used them to beat a professional at anything.
Berlekamp’s demonstrations were constructed endgames — positions designed so that the decomposition is exact — and the results were unambiguous: players of very high rank, given the analysis, changed their move order.
The general form of the claim
Strip out the Go and what is left is a statement about greedy rules that is worth having on its own.
A greedy rule ranks the available actions by a number and takes the best. Greedy rules are usually either optimal — and then there is a theorem — or heuristic, and then there is nothing. Hottest-first is neither: it is provably suboptimal and provably close, with the closeness measured in the same units as the ranking.
That combination is rarer than it sounds and is the most transferable thing in this essay. It requires the ranking quantity to be the same quantity as the error, which is exactly what temperature is: the amount at stake, and therefore the amount that can be lost by moving elsewhere.
The reversed rule is the evidence that the condition is doing work rather than being satisfied by anything. Coolest-first ranks by the same quantity, in the opposite direction, and its loss is not measured in that quantity at all — the figure at the top of this essay shows it breaking the bound on fifty-four lines and giving up as much as five points. So the guarantee is not a property of greedy rules, nor of temperature as a ranking device; it belongs to one rule, and it belongs to it because that rule takes the component whose stake is the one the bound is stated in.
What it did not do
For completeness, because the claim in the title is narrow.
It did not produce a Go program. The strongest Go programs owe nothing to this theory, and the ones that eventually surpassed professionals did so by learning rather than by decomposing.
It did not change how the game is taught. Endgame technique in Go is still taught by counting in the traditional way, and the theory’s vocabulary has not been adopted.
And it did not extend to the middlegame, where the board does not decompose and the hypothesis fails outright.
So the achievement is a narrow one, and it is the only one of its kind this subject has: a case where the mathematics produced a correct answer that expertise had got wrong, in a game people actually play, in positions somebody could check.
Where the ladder goes next
The rungs below build the thermograph, read it, and use it to choose a move. This rung is what happened when the reading was handed to somebody who plays for a living. The direction onward is the coupling — what happens to the account when the regions are not quite independent, which is where the theory currently stops being exact and starts being a very good approximation with a bound.
Part 5 of 8
One argument about Temperature. The parts either side of it:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 10.
The objects named here
The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.
ApproximationCoolingDisjunctive sumGo endgameGreedyMean valueOrthodox accountingSwitchesTemperatureTempoThermograph
- The bound names the hottest part and the cost does not approximation, disjunctive sum, mean value, switches, temperature, thermograph
- The operator chosen for one game approximation, cooling, go endgame, mean value, temperature, thermograph
- A schedule instead of a number approximation, disjunctive sum, mean value, temperature, thermograph
- Cooling adds and heating does not cooling, disjunctive sum, mean value, temperature, thermograph
- Cooling by exactly one cooling, go endgame, mean value, temperature, thermograph
- How cold a sum of hot games can be cooling, disjunctive sum, mean value, temperature, thermograph