Temperature

A schedule instead of a number

Playing in the hottest component loses at most the largest temperature on the board — the classical guarantee, stated against one number. Sorting the temperatures and reading the guarantee one step further down gives a promise 47 per cent smaller that is never breached over 1,734 sums. Two steps down it fails 120 times, so the schedule has exactly one step of slack in it.

Assumes: A pool built to punish greed · A rule with a guarantee

Play in the hottest component is the rule with a theorem behind it. A rule with a guarantee states it: a player who always moves in a hottest part of the sum finishes within the largest temperature on the board of what optimal play would have got.

That is a promise about one number, and a pool built to punish greed closed by naming what should replace it:

The rung above is the ambient temperature as a schedule rather than as a single number. Both bounded rules are stated against the largest temperature on the board, and the sharper theory replaces that with a whole graded stack of temperatures, against which the guarantee becomes an integral rather than a subtraction.

The discrete version of that is a list. Sort the components’ temperatures, largest first, and ask which position in the list the guarantee can be read off. The answer is the second, and only the second.

How far down the stack the guarantee reaches. The temperatures of a sum's components, sorted largest first, with each position asked whether playing in the hottest component can lose more than the temperature sitting there. The first two never fail; the third fails on 681 sums.
Fig. 1 Over 1,734 sums, four positions in the sorted stack of temperatures asked whether playing in the hottest component can lose more than the temperature sitting there. The first two never fail; the third fails on 120 sums.

The census

The pool is the site’s standing list of hot components together with the trap components built to punish greed — positions offering a large immediate gain and handing over a larger follow-up. Every sum of two and three of them is built, with repeats allowed, giving 1,734 sums with at least one hot part in them.

Each is played twice: once with one side following move in a hottest component and the other playing exactly, and once with both playing exactly. The difference is what the rule costs, and the question is what bounds it.

The classical bound is never breached. One thousand seven hundred and thirty-four sums, no exception, and 61 of them attain it exactly — so the theorem is tight and is not tight often.

That is the shape every bound on this site turns out to have, and it is worth pausing on before the improvement. A theorem that had never been attained would be a theorem about the wrong quantity: the bound would be true and would have no relation to what happens. Three and a half per cent attainment is enough to say the number is the right number and few enough to say a player should not plan against it.

One step down

Sort the temperatures largest first. The classical guarantee reads the first entry. Read the second instead.

It is never breached either.

That is the finding, and its size is worth stating three ways. The second-largest temperature averages 2.07 against the largest’s 3.89, so the promise is 47 per cent smaller. It is exactly attained on 233 sums against 61, so it is tight nearly four times as often. And it costs nothing to compute: a player who knows every component’s temperature already has the list, and reading the second entry is not harder than reading the first.

What is promised and what is lost. The average size of the classical guarantee, of the guarantee read one step down the stack of temperatures, and of the loss actually incurred. The tighter guarantee is a third smaller and is still seven times the typical loss.
Fig. 2 The two promises against what is actually lost. Moving one step down the stack cuts the guarantee by more than a third, and the tighter guarantee is still seven times the average loss.

Two steps down fails

The third-largest temperature is breached on 120 of the 1,734, and the fourth on 221.

So the stack has exactly one step of slack. That is a much more definite answer than the guarantee becomes an integral suggested, and the reason it stops where it does is visible in the breaches.

The sum that refutes the third level. The worst breach of the guarantee read two steps down the stack: a sum whose temperature is carried entirely by two components, so that the third largest temperature is nought and the loss is not.
Fig. 3 The worst breach of the third level. Two components carry the whole temperature of the sum, so there is no third temperature to read a guarantee off, the bound is nought, and the loss is five.

The worst breach is a sum whose temperature lives in two components. Both are worth 5, the third component is cold, and the sorted list is 5, 5 with nothing after it. The third entry is nought, so the bound is nought, and the rule loses five.

That is the whole mechanism. The guarantee is about the reply. A player moving in the hottest component hands the opponent the board; what the opponent can then take is bounded by the hottest thing still standing, which is the second entry in the list. There is no third move in the argument, so there is no third entry to read.

Why the second and not the first

Stating the mechanism the other way round explains why the classical bound was ever stated at the first entry, and it is not an error.

The classical statement is proved without sorting anything: the loss is at most the ambient temperature, where the ambient temperature is the largest on the board. That is true, it is easy to state, and it does not require knowing which component is which.

The second-entry version is the same argument run one move further. A player moves in the hottest; that component is now cooler or gone; the opponent’s best reply is in whatever is hottest now, which is at most the old second entry. The loss the first player suffers is the value of that reply, so the second entry bounds it.

Written out like that the improvement looks obvious, and the census is what says it is true rather than plausible — because whatever is hottest now is not necessarily the old second entry. Moving in a component can leave it hotter than it was; a switch answered becomes another switch, and a component whose temperature rises after a move would break the argument. The census finds no such sum in 1,734, which is the evidence the sentence needs and which is not the same as a proof.

The six hundred and eighty-one

The breaches at the third level are worth a look as a class, because they say what shape of position the improvement cannot be pushed further on.

They are the sums whose heat is concentrated. A sum of three components with temperatures 5, 5 and nothing has a third entry of nought and can lose 5; a sum with temperatures 5, 4 and 3 has a third entry of 3 and loses less than that. The bound at the third level is what remains after two components have been dealt with, and a sum whose heat lives in two components has nothing left.

That is a fact about how heat is distributed rather than about how much of it there is, and it is invisible to any single number. Two boards with the same largest temperature and the same total temperature can be one breach and one not, decided entirely by whether the heat is in two components or spread across three.

Concentration is also what makes a position hard for a human. A board with one big fight and a lot of small ones plays itself; a board with two equal big fights is where the reading has to be done, and it is exactly the shape this page’s bound stops covering. The theory and the difficulty agree, which does not always happen.

Five rules over 120 sums built to punish greed. Each rule plays every sum against an opponent evaluating exactly, on a pool whose components are traps: a large immediate gain that hands the opponent a larger follow-up. The pool was built to punish the greedy rule and does not — that rule scores a move by the stop it leaves, and a stop already contains the follow-up. What the traps catch is the rule below it, which scores a move by the territory it takes and loses up to 16.
Fig. 4 The four rules run over the components built to punish greed, which is the pool half these breaches come from. A trap component is one whose largest temperature understates what a reply is worth, and it is the same feature that concentrates heat.

What the numbers say about guarantees generally

The third figure is the one to sit with, because it is about every bound in the subject and not only this one.

The average loss is 0.28. The classical guarantee averages 3.89 and the improved one 2.07. So even after a 47 per cent improvement, the promise is seven times what is typically lost.

That is not a criticism of either theorem. A guarantee is a worst case and worst cases are rare — 61 sums attain the classical bound out of 1,734, which is three and a half per cent. What the gap says is that a guarantee is the wrong instrument for deciding how to play: it tells a player what cannot happen, and the thing a player wants to know is what usually does.

And it says where a better guarantee would have to come from. Tightening from the first entry to the second closed nearly half the gap and there are two more entries in the stack, both of which fail. So the remaining factor of seven is not going to be recovered by reading further down; it would need a different quantity, and the census has no candidate.

How much room is under the promise. Every sum of two or three components from a pool of twenty, including the traps built to punish greed — 1734 boards, each played out with one side following the hottest rule and the other evaluating exactly. The rule never ends below the board's mean, so the guarantee's whole temperature of margin goes unspent; read instead as the bound on the loss it is, the promise is attained 61 times.
Fig. 5 The rung below’s measurement of how much room there is under the classical guarantee. This page’s improvement is one step into that room and the room is much larger than one step.

What a player should do with it

The practical form of this page is short and it is worth separating from the theory.

Look at the second-hottest fight, not the hottest. A player deciding whether take the biggest is safe here should be reading the size of what the opponent gets in reply, and that is the second entry in the sorted list. If the second entry is small, playing greedily is nearly free whatever the first entry is; if the two are the same size, the rule can cost the whole of one of them.

That is the reading of the breach above turned into advice. Two fives on the board is the dangerous shape, and one five and a one is not — even though the classical guarantee says five in both cases.

And it explains why the rule feels safer than it is promised to be. In an ordinary endgame the temperatures fall away steeply, so the second entry is much smaller than the first and the real exposure is small. The positions where the rule is dangerous are the ones with two equally hot fights, which is exactly where sente is a fact about the rest of the board rather than about the fight, and where the crossover has to be computed rather than guessed.

What the rule costs. Every sum of three components from a fixed pool, played out twice: once with one side following the rule "move where the stake is largest" and once with both sides evaluating exactly. The rule is not optimal, the gap is bounded, and the bound is the largest temperature on the board.
Fig. 6 The rule measured on the standing pool, where it was first shown to cost something. This page is the same measurement asked what bounds the cost rather than what the cost is.

Where the improvement comes from, arithmetically

There is a way to see the 47 per cent that does not involve any game theory, and it is worth having because it says the improvement is about the pool rather than about the theorem.

The classical bound is the maximum of a list; the improved one is its second largest. The gap between the two is the gap between the largest and the second largest of whatever list the pool produces, and that is a fact about the distribution of temperatures in the pool, not about Domineering or Go or switches.

Here the components are drawn from a list whose temperatures run from a half to five, and a sum of two or three of them has a top gap of about half its maximum on average. A pool whose temperatures were nearly all equal would have almost no gap and the improvement would be worth nothing; a pool with one very hot component and the rest cold would have a gap of nearly everything and the improvement would be enormous.

So the 47 per cent is not transferable and the zero breaches are. That is the right way round for a result: the structural claim — the second entry bounds the loss — is what a reader should carry, and the size of the saving is a fact about the board in front of them, computable on the spot by looking at their own two hottest fights.

The same caution applies in reverse to the seven-fold looseness. In a pool with steeply falling temperatures the average loss would be smaller still and the guarantee looser; in a pool of equal temperatures the loss would approach the bound. A guarantee’s tightness is a property of the position, and the only way to know it is to look, which is what makes a bound a bound rather than an estimate.

One step, and why one is the right number to expect

That the guarantee holds one step down the sorted list and fails two steps down is a tidy result, and it is worth asking whether the tidiness is a coincidence of this pool or the shape of the argument.

The classical proof charges the rule for one exchange going the wrong way. The player following it takes the hottest component; the opponent may then take the second-hottest; and the accounting that produces the bound allows for the follower losing the difference once. The bound is stated against the largest temperature because the largest is what the follower took, and stating it against the second largest would be stating it against what the opponent took instead.

So one step is exactly the number of exchanges the argument is paying for, and finding the promise good one step down and bad two steps down is finding that the argument is charging for one exchange and the play only ever costs one.

That is a much better position than the bound is loose by an unknown amount. It says the looseness is structural, that it has a size that can be named without a sweep, and that a sharper bound is a re-statement of the existing proof rather than a new argument.

It also says the improvement does not compound with board size, which the census confirms and which a reader might otherwise expect. A schedule with more entries does not give more slack, because the slack is one exchange whatever the board is — so the 47 per cent is a fact about how far apart the top two temperatures typically sit rather than about how many components there are.

What the census does not say

Four limits.

A pool, not a population. The components are the site’s standing list plus the traps, and sums of two and three of them. The same census over four-component sums gives the same answer at every level and costs half a gigabyte of the build to say so, which is why it is not the one run here. A pool built differently gives different averages; what it should not give is a breach at level two, and that is the claim.

The improvement is checked and not proved. The loss is at most the second-largest temperature is a statement about all sums, checked on 1,734. The argument in the section above is a sketch and its weak point is named there: it assumes no component gets hotter when moved in, and no sum in the pool does that.

Sorted temperatures, not a schedule. The rung below asked for a graded stack against which the guarantee is an integral, and this page has tested the discrete reading — a sorted list and a position in it. The continuous version, where a coupon stack supplies a temperature at every level, is a different and richer object and is where the environment comes in.

And one rule. Everything here is about play in the hottest. The other bounded rule reads the stops rather than the temperatures, and whether its guarantee also survives a step down the stack is a computation this page has not run.

The convention, named

Normal play, and the temperature is the standard one: a tax on moving, with the temperature the height at which neither player wants to move. A number is given temperature −1, so a sum’s stack of temperatures contains only its genuinely hot components.

Playing in the hottest means, at each turn, moving in a component of maximal temperature and taking the best move there; ties are broken by the stop the option leaves. The opponent plays exactly, by the full recursion over the sum.

The loss is the score under optimal play minus the score under the rule, both from the same starting position with the same player to move. A breach is a sum whose loss exceeds the bound being tested, by more than the arithmetic tolerance.

Where the ladder goes next

The strategy anchor has four rungs: a rule with a guarantee, a rule with none that beats it, the pool built to separate them, and now how far down the stack of temperatures the guarantee reaches.

The rung above is the continuous version. A coupon stack supplies a temperature at every level rather than at four, and the guarantee against it is the integral the rung below asked for; measuring where that stops holding needs an environment rather than a pool, and the machinery for one is already here. The prediction this page suggests is that the answer will again be one step, because the argument is about a reply and a reply is one move.

Two neighbours are worth the trip. Temperatures do not add is why the stack has to be sorted rather than summed, and it is the fact that makes the whole question a question about a list. And a rule with no promise at all is the rule that beats this one in practice and promises nothing, which is the standing reminder that a bound and a performance are different measurements.

Part 4 of 5

One argument about Strategy. The parts either side of it:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.

ApproximationBoundCounterexampleDisjunctive sumEnumerationGuaranteeHotstratMean valueStrategySwitchTemperatureThermograph