A pool built to punish greed
Assumes: A rule with no promise at all · A rule with a guarantee
A rule with no promise at all set four move rules against each other on 220 sums and found the uncomfortable result: the greedy rule, which has no theorem of any kind attached, scored what perfect play scores on 205 of them against the guaranteed rule’s 196. It closed by naming the repair and predicting the outcome:
The first rung is an adversarial pool — components built to punish greed rather than drawn from a standing list. The obvious shape is a component with a large immediate gain and a larger follow-up for the opponent, which is what sente is, and a pool of those should separate the two rules the other way.
The pool was built. The prediction is wrong.
Greed is barely separated from the rule with the theorem, and it does not break the bound once.
The pool, and what was in it
Eight components. Five are traps of exactly the shape the prediction described — , , and two mirror images — where Left’s move looks like a gain of four, six or eight and hands Right a fight worth three times as much. Three are honest switches, , and , because a board made entirely of traps would be a board with no decision on it.
Every unordered triple gives 120 sums, and each is played out with one side following the rule and the other evaluating exactly.
The guaranteed rules come through untouched, which was never in doubt: the bound is a theorem and a pool cannot break it. What the sweep checks — and refuses on — is that the guarantee holds here, so a pool that produced a breach would stop the build rather than being reported as a surprise.
Why the traps miss
The prediction failed for a reason that is more interesting than the prediction, and it is a fact about the rule rather than about the pool.
The rung below defined the greedy rule precisely: play the move with the largest incentive, where a move’s incentive is measured by the stop it leaves. For Left choosing among options of a component , the quantity maximised is .
A stop is the result of playing the component out. is what Right gets by moving first in and fighting on until somebody is left facing a number — so it already contains every follow-up inside that component, in full, correctly evaluated.
So the greedy rule is not short-sighted about follow-ups at all. It is short-sighted about which component, and about nothing else. Handing it does not fool it: it prices Left’s move at the stop rather than at the number , and declines it exactly as the guaranteed rule would.
The word was carrying more short-sightedness than the rule. That is the finding, and it is a finding about a piece of this site’s own machinery rather than about combinatorial game theory — which is why the repair is a new rule and not a wider pool.
The rule the traps do catch
A genuinely myopic rule scores a move by the territory it takes: the mean of the position it leaves, which is what that component will be worth once nobody wants to move in it any more. It ignores whose turn it is next, which is precisely the thing a follow-up is about.
That rule is what the traps were built for, and it is caught by them.
| standing pool | trap pool | |
|---|---|---|
| worst loss, hottest rule | 1 | 5 |
| worst loss, greedy rule | 1 | 6 |
| worst loss, territory rule | 9 | 16 |
| sums where the territory rule breaks the guarantee | 24 of 220 | 11 of 120 |
The comparison is worth reading twice. On the traps the territory rule loses sixteen where the guaranteed rule loses five, which is a factor of three; on the standing pool it loses nine where the guaranteed rule loses one, which is a factor of nine. The traps make the guaranteed rule worse faster than they make the bad rule worse, because a component with a big follow-up is a component where even a good rule can be caught out about the ambient temperature.
What the traps do and do not do
That last observation deserves its own sentence, because it is the second thing the prediction got wrong.
An adversarial pool does not simply widen the gap between good rules and bad ones. It raises the stakes for everybody: the largest temperature on a trap board is larger, and the guarantee scales with the largest temperature, so a rule promised a score within of the mean is promised much less on a board where is five than on one where is one.
So the honest summary is not the traps work or the traps fail. It is that a pool designed to punish one rule punishes every rule, and separating two rules requires a pool on which one of them specifically goes wrong — which means understanding what the rule is actually measuring, which is what the previous section had to do.
Three sums, and each rule’s worst
The arithmetic is easier to believe on single boards than in a table of 120, and the three worst cases in the pool say more together than any of them says alone.
The territory rule’s worst is three copies of one trap: three times over. The board’s mean is and its largest temperature is , so the guarantee is a score of at least nought. Perfect play scores ; the hottest rule scores ; the greedy rule scores ; the territory rule scores . Sixteen points lost, and the guarantee broken by four. It takes the eight, three times, and is answered in the twenty each time.
The greedy rule’s worst is one trap and two honest switches: . Perfect play scores ; the hottest rule and the territory rule both score ; the greedy rule scores . Six lost, and it is lost on the ordering of the two switches rather than in the trap at all.
And the hottest rule’s worst is the one worth pausing on: . Perfect play scores . The rule with the theorem scores — five lost, exactly its whole promise. The greedy rule scores , which is to say it plays this board perfectly.
Three boards, three different rules going wrong, and on the board that hurts the guaranteed rule most the unguaranteed one is exact. That is the shape of the whole comparison, and it is why the previous rung’s result was uncomfortable rather than wrong.
What “greedy” should have been called
The rule this page had to invent is worth one more paragraph, because naming it properly is most of the finding.
There are three quantities a rule can score a move by, and the site’s four standing rules use two of them. Temperature is what the hottest rule reads: how much is at stake in a component, before any move is made. The stop left behind is what the greedy rule reads: what the component comes to if both players fight it out from the position the move creates. And the mean left behind — the territory rule — is what the component settles at once nobody wants to move in it at all.
The first two are forward-looking in the only sense that matters. A temperature is computed from the whole thermograph and a stop is computed by playing the component to its end, so both have already priced every follow-up inside that component. The third is not: a mean is what the position is worth ignoring who has the move, which is precisely the information a follow-up carries.
So the honest labels are not guaranteed and greedy. They are two rules that read the component correctly and disagree about how to compare components, against one rule that reads the component wrongly. That reordering explains every number on this page at once — why the traps cannot separate the first two, why they catch the third, and why the third loses nine points on a pool nobody aimed at it.
It also explains why the first two stay within a point or two of each other everywhere. Both are looking at the same fully-priced components and differing only on a tie-break; the space between them is the space between two readings of one correct picture. The territory rule is looking at a different picture, and the gap is correspondingly an order of magnitude.
The lesson for building the next pool is the one this one paid to learn. An adversarial example must be adversarial to a stated quantity, and the quantity has to be identified from the rule’s definition rather than from its name. Five components were built to punish a rule that does not exist, and the rule that does exist had to be written afterwards.
Where the promise is exactly attained
The other half of the rung below’s closing section is the tightness question, and it has an answer.
A sum on which the hottest rule lands exactly on the bound is a certificate that the bound is the right one, and none of the 935 sums swept here is one.
There are two readings of lands exactly on the bound, and they come out opposite ways.
Read as a bound on the score — at least the mean less the largest temperature — nothing comes near it, and the reason is much stronger than “the pool is easy”. Over 1,734 sums the rule never once ends below the board’s mean, and lands exactly on the mean 217 times. The guarantee gives away a whole temperature that the rule never spends.
That is stated as a conjecture the sweep supports, not a theorem, and it is refusable: a sum on which the rule ends below the mean stops the build.
Read as a bound on the loss against perfect play — at most the largest temperature — the bound is attained, sixty-one times. The worst case is
on which perfect play scores , the hottest rule scores , and the largest temperature on the board is . Five lost, five promised, nothing to spare.
So the certificate exists and the previous essay was looking for it in the wrong currency. The theorem is a statement about the loss, it is exactly right at that size, and reading it as a floor on the score is reading it enormously weakly.
The sixty-one witnesses are not scattered either. Fifty-eight of them contain a trap and three do not, which is the one place the adversarial pool did the job it was built for: the earlier sweep of 935 sums from the standing pool found none at all, and adding components with large follow-ups is what makes the bound reachable. So the prediction was half right after all — the traps did not separate the two rules, and they did make the guaranteed rule’s promise bite. A rule with a guarantee measured how far from optimal the rule can get and answered with a number it could not reach; this pool reaches it.
What a player should take from this
Three things, and the second is the one worth carrying to a real board.
A bound on the loss is not a floor on the score. They differ by the whole mean, and on a board with a positive mean the difference is the whole of what a player has. The guarantee that a hottest-first player never falls more than behind perfect play is compatible with their score being anything at all.
The dangerous mistake is counting territory, not chasing gains. Chasing the biggest immediate gain is safe here, because a gain measured by a stop has already looked ahead. Counting the biggest immediate territory is what loses sixteen points, and it is what a player does when they add up the board as though the fighting were over. The endgame, accounted for is the correct way to do that addition, and its whole apparatus exists because the naive version is this bad.
There is a fourth thing, and it is a caution about reading any of these numbers as advice. Every count on this page is a count against an opponent who evaluates exactly, and a real opponent does not. Sente is a fact about the rest of the board is precisely the observation that a threat is only worth making if the opponent answers it, and a rule tuned against perfect play may be the wrong rule against a fallible one. Nothing here measures that, and the first time it told somebody something is the essay about what happens when the theory is handed to a person who plays for a living.
And an adversarial example has to be adversarial to something specific. A component with a big follow-up is a trap for a rule that ignores follow-ups, and none of the four rules on this site’s standing table does. Building the pool was worth it for finding that out.
What the sweep does not settle
The pool is eight components at three sizes and the sums are triples, which is a small and deliberately shaped population. Nothing here says what the worst case is over all boards; it says what the worst case is over these, and the worst case over all boards is exactly what the theorem covers and the sweep cannot.
The conjecture that the rule never ends below the mean is supported by 1,734 sums and by no argument. It has the shape of something provable — the mover has a move at the hottest component and the mean is what the board settles to — but a proof would need the alternation argument done properly, and stating it as a conjecture is the honest form until somebody does.
And the territory rule is a rule this essay wrote in order to be beaten. It is a fair model of a common human error and it is not drawn from the literature, so its sixteen-point loss is a fact about a stated rule rather than a measurement of anything anybody advocates.
The convention, named
Normal play, disjunctive sum, and the score is the total of the components once every one of them is a number. The side following the rule moves first throughout; the other side uses an exact evaluator, so every “loss” above is a loss against play that cannot be improved on.
The temperature quoted in each guarantee is the largest temperature on the board at the start, not a running one. That matters: the ambient temperature falls as the board is played out, and a guarantee restated against a falling temperature would be a different and sharper promise. An environment made of coupons is where that sharper version lives.
Where the ladder goes next
The strategy anchor has three rungs to here: a rule with a guarantee, a rule with none that beats it, and now the pool built to separate them and what it found instead.
The rung above takes the third direction the rung below named — the ambient temperature as a schedule rather than as a single number — and it turns out to need far less machinery than the integral this page was expecting. A schedule instead of a number sorts the board’s temperatures and reads the guarantee one step down the list rather than at the top. That promise is 47 per cent smaller than the classical one and it is never breached across the same 1,734 sums this page’s tightness figure is built on. Two steps down it fails 120 times.
So the slack this page measures has a size, and the size is exactly one step. The classical bound is not weak by an unknown margin that a finer theory might close; it is weak by one entry of the sorted temperature list, and the next entry after that is where it stops being true. That is a much more useful thing to hand a player than the bound is loose, and it explains the shape of the sixty-one witnesses above: a bound attained sixty-one times out of 1,734 is a bound that is right at the top of the list and generous everywhere below it.
It also puts the conjecture on this page in better company. The rule never ends below the board’s mean is still unproved, and it now sits beside a second measured statement of the same kind — one step of slack, never two — which suggests the missing argument is about how the schedule is consumed as the board is played out rather than about any single component.
Two neighbours are worth the trip. A rule with no promise at all is the rung below, whose prediction this page tested and refuted, and whose definition of greedy is the thing that turned out to matter. And sente is a fact about the rest of the board is where the trap shape is defined properly, and where the reason a follow-up changes a move’s value is worked out rather than assumed.
Part 3 of 5
One argument about Strategy. The parts either side of it:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.
Ambient temperatureBoundCounterexampleDisjunctive sumExhaustive searchFollow-upGreedy playHeuristicHot positionIncentiveMean valueMove selectionSenteStopsStrategyTemperature
- The answer that starts another fight ambient temperature, counterexample, exhaustive search, follow-up, mean value, move selection, sente, stops, temperature
- The quantity that does not order a board bound, counterexample, disjunctive sum, follow-up, heuristic, mean value, sente, strategy, temperature
- What a move nobody makes is worth ambient temperature, disjunctive sum, exhaustive search, follow-up, mean value, move selection, sente, stops, temperature
- A rule that beats the hottest bound, disjunctive sum, follow-up, heuristic, mean value, sente, strategy, temperature
- How cold a sum of hot games can be ambient temperature, counterexample, disjunctive sum, exhaustive search, follow-up, mean value, stops, temperature
- A bound with one number too many bound, counterexample, disjunctive sum, follow-up, mean value, stops, temperature