Temperature

A pool built to punish greed

The rung below found a rule with no theorem behind it beating the rule with one, and predicted that a pool of deliberate traps would reverse the result. It does not. The traps miss, because the rule called greedy scores a move by the stop it leaves and a stop already contains the follow-up — and the rule the traps do catch, losing sixteen points where the guaranteed rule loses five, had to be written to make the point.

Assumes: A rule with no promise at all · A rule with a guarantee

A rule with no promise at all set four move rules against each other on 220 sums and found the uncomfortable result: the greedy rule, which has no theorem of any kind attached, scored what perfect play scores on 205 of them against the guaranteed rule’s 196. It closed by naming the repair and predicting the outcome:

The first rung is an adversarial pool — components built to punish greed rather than drawn from a standing list. The obvious shape is a component with a large immediate gain and a larger follow-up for the opponent, which is what sente is, and a pool of those should separate the two rules the other way.

The pool was built. The prediction is wrong.

Five rules over 120 sums built to punish greed. Each rule plays every sum against an opponent evaluating exactly, on a pool whose components are traps: a large immediate gain that hands the opponent a larger follow-up. The pool was built to punish the greedy rule and does not — that rule scores a move by the stop it leaves, and a stop already contains the follow-up. What the traps catch is the rule below it, which scores a move by the territory it takes and loses up to 16.
Fig. 1 The same rules over 120 sums whose components are traps: a large immediate gain that hands the opponent a larger follow-up. The greedy rule loses six at worst against the guaranteed rule’s five, and stays inside a promise it was never given, on every one of the 120.

Greed is barely separated from the rule with the theorem, and it does not break the bound once.

The pool, and what was in it

Eight components. Five are traps of exactly the shape the prediction described — {4{012}}\{4 \mid \{0 \mid -12\}\}, {6{116}}\{6 \mid \{1 \mid -16\}\}, {8{220}}\{8 \mid \{2 \mid -20\}\} and two mirror images — where Left’s move looks like a gain of four, six or eight and hands Right a fight worth three times as much. Three are honest switches, {55}\{5 \mid -5\}, {33}\{3 \mid -3\} and {11}\{1 \mid -1\}, because a board made entirely of traps would be a board with no decision on it.

Every unordered triple gives 120 sums, and each is played out with one side following the rule and the other evaluating exactly.

{8 | {2 | −20}} beside one other fight. A local position and a single switch, played out together at each of several ambient temperatures. The middle columns are what optimal play does: whether it opens the local fight, and whether it answers when the opponent opens it. The answer stops being forced at a temperature the local position alone does not name.
Fig. 2 One of the traps, priced against a rising ambient temperature. Left’s move looks like a gain of eight and leaves a fight worth twenty; whether taking it is right depends entirely on how hot the rest of the board is, which is what makes the component a trap and not merely a bad move.

The guaranteed rules come through untouched, which was never in doubt: the bound is a theorem and a pool cannot break it. What the sweep checks — and refuses on — is that the guarantee holds here, so a pool that produced a breach would stop the build rather than being reported as a surprise.

Why the traps miss

The prediction failed for a reason that is more interesting than the prediction, and it is a fact about the rule rather than about the pool.

The rung below defined the greedy rule precisely: play the move with the largest incentive, where a move’s incentive is measured by the stop it leaves. For Left choosing among options oo of a component GG, the quantity maximised is RS(o)RS(G)\mathrm{RS}(o) - \mathrm{RS}(G).

A stop is the result of playing the component out. RS(o)\mathrm{RS}(o) is what Right gets by moving first in oo and fighting on until somebody is left facing a number — so it already contains every follow-up inside that component, in full, correctly evaluated.

So the greedy rule is not short-sighted about follow-ups at all. It is short-sighted about which component, and about nothing else. Handing it {4{012}}\{4 \mid \{0 \mid -12\}\} does not fool it: it prices Left’s move at the stop 12-12 rather than at the number 44, and declines it exactly as the guaranteed rule would.

The word was carrying more short-sightedness than the rule. That is the finding, and it is a finding about a piece of this site’s own machinery rather than about combinatorial game theory — which is why the repair is a new rule and not a wider pool.

The rule the traps do catch

A genuinely myopic rule scores a move by the territory it takes: the mean of the position it leaves, which is what that component will be worth once nobody wants to move in it any more. It ignores whose turn it is next, which is precisely the thing a follow-up is about.

That rule is what the traps were built for, and it is caught by them.

standing pool trap pool
worst loss, hottest rule 1 5
worst loss, greedy rule 1 6
worst loss, territory rule 9 16
sums where the territory rule breaks the guarantee 24 of 220 11 of 120
Five rules over 220 sums. Each rule plays every sum against an opponent evaluating exactly. Two of the rules come with a bound and two do not; the coldest rule is the control, and it violates the bound often enough to show that being inside it is a real constraint rather than a description of the pool.
Fig. 3 The standing pool with the fifth rule added. The territory rule loses nine points at its worst and violates the guarantee twenty-four times — on a pool nobody designed to hurt it, and where the greedy rule and the guaranteed rule are within a point of each other.

The comparison is worth reading twice. On the traps the territory rule loses sixteen where the guaranteed rule loses five, which is a factor of three; on the standing pool it loses nine where the guaranteed rule loses one, which is a factor of nine. The traps make the guaranteed rule worse faster than they make the bad rule worse, because a component with a big follow-up is a component where even a good rule can be caught out about the ambient temperature.

What the traps do and do not do

That last observation deserves its own sentence, because it is the second thing the prediction got wrong.

An adversarial pool does not simply widen the gap between good rules and bad ones. It raises the stakes for everybody: the largest temperature on a trap board is larger, and the guarantee scales with the largest temperature, so a rule promised a score within tt of the mean is promised much less on a board where tt is five than on one where tt is one.

A rule that is not optimal, and cannot be far wrong. Every line of play in the experiment as one point: the largest temperature on the board across, and what playing hottest-first cost against optimal play up. The diagonal is the bound the theory proves. Nothing gold sits above it; the magenta series is the same rule with its ordering reversed, and it sits above the line often — which is what makes the bound a claim rather than a description.
Fig. 4 The margin between the guaranteed rule and perfect play, and what a bound on it is worth. The bound is stated against the largest single temperature on the board, so a pool of hotter components is a pool where the promise says less — which is the mechanism by which a trap pool hurts the rule with the theorem as much as it hurts the rule without one.

So the honest summary is not the traps work or the traps fail. It is that a pool designed to punish one rule punishes every rule, and separating two rules requires a pool on which one of them specifically goes wrong — which means understanding what the rule is actually measuring, which is what the previous section had to do.

Three sums, and each rule’s worst

The arithmetic is easier to believe on single boards than in a table of 120, and the three worst cases in the pool say more together than any of them says alone.

The territory rule’s worst is three copies of one trap: {8{220}}\{8 \mid \{2 \mid -20\}\} three times over. The board’s mean is 66 and its largest temperature is 66, so the guarantee is a score of at least nought. Perfect play scores 1212; the hottest rule scores 1212; the greedy rule scores 1212; the territory rule scores 4-4. Sixteen points lost, and the guarantee broken by four. It takes the eight, three times, and is answered in the twenty each time.

One board, and what each rule scores. A single sum of three components, played out once by each rule against an opponent who evaluates exactly. Each row is the component that rule moves in first and the score it ends with. The board's mean and its largest temperature are computed from the components' thermographs, and the promise the two make is checked against every row rather than asserted once.
Fig. 5 One trap, three times over, with each rule’s opening and its final score. All three components are the same, so which component a rule opens in decides nothing and the whole difference between the rows is which option it takes inside one. The two rules that price a move by what the component comes to both play the board perfectly at 1212; the rule that prices a move by the territory it leaves takes the eight and is answered in the twenty, three times running, and ends at 4-4 — below a promise of nought.

The greedy rule’s worst is one trap and two honest switches: {8{220}}+{55}+{55}\{8 \mid \{2 \mid -20\}\} + \{5 \mid -5\} + \{5 \mid -5\}. Perfect play scores 88; the hottest rule and the territory rule both score 88; the greedy rule scores 22. Six lost, and it is lost on the ordering of the two switches rather than in the trap at all.

One board, and what each rule scores. A single sum of three components, played out once by each rule against an opponent who evaluates exactly. Each row is the component that rule moves in first and the score it ends with. The board's mean and its largest temperature are computed from the components' thermographs, and the promise the two make is checked against every row rather than asserted once.
Fig. 6 The greedy rule’s worst board, and the opening column is the whole story. The two rules that score 88 open in the trap; the greedy rule opens in a switch — not because the trap fooled it, since it prices the trap at the stop the trap leaves and declines it correctly, but because a plain five-point switch offers the largest immediate gain on the board and taking one of the two first is wrong here. The trap is a bystander to the six points.

And the hottest rule’s worst is the one worth pausing on: {4{012}}+{6{116}}+{55}\{4 \mid \{0 \mid -12\}\} + \{6 \mid \{1 \mid -16\}\} + \{5 \mid -5\}. Perfect play scores 66. The rule with the theorem scores 11 — five lost, exactly its whole promise. The greedy rule scores 66, which is to say it plays this board perfectly.

One board, and what each rule scores. A single sum of three components, played out once by each rule against an opponent who evaluates exactly. Each row is the component that rule moves in first and the score it ends with. The board's mean and its largest temperature are computed from the components' thermographs, and the promise the two make is checked against every row rather than asserted once.
Fig. 7 The board the guaranteed rule handles worst and the unguaranteed one handles perfectly. The mean is 11 and the largest stake is 55, so the promise is a score of at least 4-4; the rule ends at 11, five above its promise and five below perfect play. Both other rules open in {55}\{5 \mid -5\}, which is where perfect play opens; the guaranteed rule opens in {6{116}}\{6 \mid \{1 \mid -16\}\} instead, and the two components are tied at a temperature of five, so it is the tie-break rather than the measurement that sends it there.

Three boards, three different rules going wrong, and on the board that hurts the guaranteed rule most the unguaranteed one is exact. That is the shape of the whole comparison, and it is why the previous rung’s result was uncomfortable rather than wrong.

What “greedy” should have been called

The rule this page had to invent is worth one more paragraph, because naming it properly is most of the finding.

There are three quantities a rule can score a move by, and the site’s four standing rules use two of them. Temperature is what the hottest rule reads: how much is at stake in a component, before any move is made. The stop left behind is what the greedy rule reads: what the component comes to if both players fight it out from the position the move creates. And the mean left behind — the territory rule — is what the component settles at once nobody wants to move in it at all.

The first two are forward-looking in the only sense that matters. A temperature is computed from the whole thermograph and a stop is computed by playing the component to its end, so both have already priced every follow-up inside that component. The third is not: a mean is what the position is worth ignoring who has the move, which is precisely the information a follow-up carries.

So the honest labels are not guaranteed and greedy. They are two rules that read the component correctly and disagree about how to compare components, against one rule that reads the component wrongly. That reordering explains every number on this page at once — why the traps cannot separate the first two, why they catch the third, and why the third loses nine points on a pool nobody aimed at it.

It also explains why the first two stay within a point or two of each other everywhere. Both are looking at the same fully-priced components and differing only on a tie-break; the space between them is the space between two readings of one correct picture. The territory rule is looking at a different picture, and the gap is correspondingly an order of magnitude.

The lesson for building the next pool is the one this one paid to learn. An adversarial example must be adversarial to a stated quantity, and the quantity has to be identified from the rule’s definition rather than from its name. Five components were built to punish a rule that does not exist, and the rule that does exist had to be written afterwards.

Where the promise is exactly attained

The other half of the rung below’s closing section is the tightness question, and it has an answer.

A sum on which the hottest rule lands exactly on the bound is a certificate that the bound is the right one, and none of the 935 sums swept here is one.

There are two readings of lands exactly on the bound, and they come out opposite ways.

How much room is under the promise. Every sum of two or three components from a pool of twenty, including the traps built to punish greed — 1734 boards, each played out with one side following the hottest rule and the other evaluating exactly. The rule never ends below the board's mean, so the guarantee's whole temperature of margin goes unspent; read instead as the bound on the loss it is, the promise is attained 61 times.
Fig. 8 Both readings, over 1,734 sums of two and three components from a pool of twenty — the standing pool, the traps, and two small switches. The rule ends below the board’s mean nought times; it ends on the mean less the temperature nought times; and it loses exactly the largest temperature sixty-one times.

Read as a bound on the scoreat least the mean less the largest temperature — nothing comes near it, and the reason is much stronger than “the pool is easy”. Over 1,734 sums the rule never once ends below the board’s mean, and lands exactly on the mean 217 times. The guarantee gives away a whole temperature that the rule never spends.

That is stated as a conjecture the sweep supports, not a theorem, and it is refusable: a sum on which the rule ends below the mean stops the build.

Read as a bound on the loss against perfect playat most the largest temperature — the bound is attained, sixty-one times. The worst case is

{{46}4}  +  {6{116}}  +  {55},\{\{4 \mid -6\} \mid -4\} \;+\; \{6 \mid \{1 \mid -16\}\} \;+\; \{5 \mid -5\},

on which perfect play scores 11, the hottest rule scores 4-4, and the largest temperature on the board is 55. Five lost, five promised, nothing to spare.

So the certificate exists and the previous essay was looking for it in the wrong currency. The theorem is a statement about the loss, it is exactly right at that size, and reading it as a floor on the score is reading it enormously weakly.

The sixty-one witnesses are not scattered either. Fifty-eight of them contain a trap and three do not, which is the one place the adversarial pool did the job it was built for: the earlier sweep of 935 sums from the standing pool found none at all, and adding components with large follow-ups is what makes the bound reachable. So the prediction was half right after all — the traps did not separate the two rules, and they did make the guaranteed rule’s promise bite. A rule with a guarantee measured how far from optimal the rule can get and answered with a number it could not reach; this pool reaches it.

What a player should take from this

Three things, and the second is the one worth carrying to a real board.

A bound on the loss is not a floor on the score. They differ by the whole mean, and on a board with a positive mean the difference is the whole of what a player has. The guarantee that a hottest-first player never falls more than tt behind perfect play is compatible with their score being anything at all.

The dangerous mistake is counting territory, not chasing gains. Chasing the biggest immediate gain is safe here, because a gain measured by a stop has already looked ahead. Counting the biggest immediate territory is what loses sixteen points, and it is what a player does when they add up the board as though the fighting were over. The endgame, accounted for is the correct way to do that addition, and its whole apparatus exists because the naive version is this bad.

There is a fourth thing, and it is a caution about reading any of these numbers as advice. Every count on this page is a count against an opponent who evaluates exactly, and a real opponent does not. Sente is a fact about the rest of the board is precisely the observation that a threat is only worth making if the opponent answers it, and a rule tuned against perfect play may be the wrong rule against a fallible one. Nothing here measures that, and the first time it told somebody something is the essay about what happens when the theory is handed to a person who plays for a living.

And an adversarial example has to be adversarial to something specific. A component with a big follow-up is a trap for a rule that ignores follow-ups, and none of the four rules on this site’s standing table does. Building the pool was worth it for finding that out.

What the sweep does not settle

The pool is eight components at three sizes and the sums are triples, which is a small and deliberately shaped population. Nothing here says what the worst case is over all boards; it says what the worst case is over these, and the worst case over all boards is exactly what the theorem covers and the sweep cannot.

The conjecture that the rule never ends below the mean is supported by 1,734 sums and by no argument. It has the shape of something provable — the mover has a move at the hottest component and the mean is what the board settles to — but a proof would need the alternation argument done properly, and stating it as a conjecture is the honest form until somebody does.

And the territory rule is a rule this essay wrote in order to be beaten. It is a fair model of a common human error and it is not drawn from the literature, so its sixteen-point loss is a fact about a stated rule rather than a measurement of anything anybody advocates.

The convention, named

Normal play, disjunctive sum, and the score is the total of the components once every one of them is a number. The side following the rule moves first throughout; the other side uses an exact evaluator, so every “loss” above is a loss against play that cannot be improved on.

The temperature quoted in each guarantee is the largest temperature on the board at the start, not a running one. That matters: the ambient temperature falls as the board is played out, and a guarantee restated against a falling temperature would be a different and sharper promise. An environment made of coupons is where that sharper version lives.

Where the ladder goes next

The strategy anchor has three rungs to here: a rule with a guarantee, a rule with none that beats it, and now the pool built to separate them and what it found instead.

The rung above takes the third direction the rung below named — the ambient temperature as a schedule rather than as a single number — and it turns out to need far less machinery than the integral this page was expecting. A schedule instead of a number sorts the board’s temperatures and reads the guarantee one step down the list rather than at the top. That promise is 47 per cent smaller than the classical one and it is never breached across the same 1,734 sums this page’s tightness figure is built on. Two steps down it fails 120 times.

So the slack this page measures has a size, and the size is exactly one step. The classical bound is not weak by an unknown margin that a finer theory might close; it is weak by one entry of the sorted temperature list, and the next entry after that is where it stops being true. That is a much more useful thing to hand a player than the bound is loose, and it explains the shape of the sixty-one witnesses above: a bound attained sixty-one times out of 1,734 is a bound that is right at the top of the list and generous everywhere below it.

It also puts the conjecture on this page in better company. The rule never ends below the board’s mean is still unproved, and it now sits beside a second measured statement of the same kind — one step of slack, never two — which suggests the missing argument is about how the schedule is consumed as the board is played out rather than about any single component.

Two neighbours are worth the trip. A rule with no promise at all is the rung below, whose prediction this page tested and refuted, and whose definition of greedy is the thing that turned out to matter. And sente is a fact about the rest of the board is where the trap shape is defined properly, and where the reason a follow-up changes a move’s value is worked out rather than assumed.

Part 3 of 5

One argument about Strategy. The parts either side of it:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.

Ambient temperatureBoundCounterexampleDisjunctive sumExhaustive searchFollow-upGreedy playHeuristicHot positionIncentiveMean valueMove selectionSenteStopsStrategyTemperature