A rule that beats the hottest
Assumes: The quantity that does not order a board · A rule with a guarantee
The quantity that does not order a board took the quantity that decides when players leave a coupon environment — the larger of a position’s own temperature and its hottest answer’s — and used it to order a board. It lost, badly: 124 boards played exactly against playing in the hottest component’s 196, and nine points lost on its worst board against one.
That page closed by proposing the opposite rule and betting against it:
The rung above is the reverse rule. If a follow-up is a liability when an opponent is watching, then a component should perhaps be discounted by its answer’s temperature — played later rather than sooner … The prediction worth recording before the sweep is that it will not beat Hotstrat either, because Hotstrat is exact on 89 per cent of these boards and there is very little left to win.
The prediction was wrong.
The rule, in one line
Every rule on this ladder chooses a component and then plays the best move inside it, so a rule is a scoring function on components and nothing else. Playing in the hottest component scores a component by its temperature . The rung below’s rule scored it by , where is the temperature of its hottest answer. The rule here scores it by
and plays where that is largest. It is the previous rule with a minus sign, and it is exact on 201 of the 220 three-component boards against the hottest rule’s 196.
The two rules lose the same amount when they lose: at most one point, on any board, at any size. So the difference between them is entirely in how often, and that is worth saying because it settles what kind of improvement this is. It is not a rule that avoids a disaster the hottest rule walks into; it is a rule that picks up single points the hottest rule leaves on the table, a little more often than the hottest rule picks up the ones it leaves.
Where they disagree
The two rules make the same choice on most boards — 203 of 220 at three components — so the whole comparison is decided by seventeen. On those seventeen the discount rule is the better one eleven times and the hottest rule six. The same shape holds at two components (2 to 1) and at four (46 to 21), which is what makes the result a property of the rule rather than of seventeen boards.
The board shows the mechanism in one move. The two large components are much the hotter and the hottest rule opens in one of them; Right answers inside it, collecting the follow-up, and the board ends a point short of perfect play. The discount rule scores the large components at their temperature less the four-point answer sitting inside them and scores the small switch at one, so it takes the small switch first and leaves the large fights until the opponent has to open one.
That is the sente reading of a follow-up applied to the ordering question. A component with a large answer is one where whoever moves second collects something, so moving into it first is spending a move to hand the opponent a gift. A component with no answer is safe to open. The rung below had the same insight the right way up and got the sign wrong, which is why the minus sign was worth trying.
The gap widens with the board
At two components the discount rule is exact 98.2 per cent of the time against 96.4; at three, 91.4 against 89.1; at four, 88.3 against 84.8. The gap is 1.8 points, then 2.3, then 3.5.
A widening gap is the reading a rule about ordering should have. On a board of two components there is one decision to get right and both rules usually get it; on a board of four there are several, and a rule that chooses better compounds its advantage. It also means the measurement is understating the difference for a real board — a played Go or Amazons position has a dozen live components, and this sweep reaches four.
The trend is three points and the pool is one pool, so it is a direction rather than a rate. What can be said is that the improvement does not shrink as the boards get bigger, which is the failure mode a rule of this kind usually has.
Which half of it is working
A rule with two terms in it invites the question of whether it needs both, and this one has a clean answer.
Dropping the component’s own temperature — playing simply where the answer is smallest — scores 495 of 715 at four components, far below either the discount rule’s 631 or the hottest rule’s 606. So the discount rule is not avoid follow-ups: a player who only avoids follow-ups plays in cold components with nothing at stake, which is the coldest rule with extra steps.
And the answer taken the right way up — the rung below’s rule — is the worst of the four at 284. The three readings sit in the order the mechanism predicts: the answer is a liability, so subtracting it helps, ignoring it is worse than subtracting it, and adding it is worse than ignoring it.
The honest summary is that the temperature does the work and the answer is a correction, in the same sense as half of the smaller temperature — a second-order term that improves a first-order rule and cannot replace it.
Why the same insight went two ways
The two rungs are worth reading together, because they used the same fact about follow-ups and reached opposite rules, and the difference between them is a question about whose decision is being made.
How big the answer is found that players leave a coupon environment at the larger of a position’s two temperatures. The reading is about when a fight is worth taking at all: a fight whose answer is large is worth entering early, because whoever enters it collects the answer as well as the fight, and a player who waits will find the opponent has taken both.
Ordering a board asks a different question. Here the fight is going to be taken by somebody, and the choice is which of several to open now. Opening one hands the answer to the opponent, because the opponent replies inside it. So the same fact — a big answer means a big second helping — makes a component urgent when the alternative is the coupon stack and makes it dangerous when the alternative is another component.
That is why the sign flips, and it is a distinction worth stating carefully because both rungs describe themselves as ordering by the second temperature. They are not ordering the same things. One is ordering a position against an environment and the other is ordering components against each other, and no amount of care about the quantity would have saved the rung below from the sign: it had to be swept.
The guarantee it was not given
A rule with a guarantee is the reason the hottest rule is the standing rule on this site, and the reason is not its score. A player following it cannot finish more than the board’s largest temperature below the board’s mean — a promise that holds on every board rather than most of them, and one that a rule with a better average and no promise does not replace.
The discount rule stays inside that bound on every board of every size swept here. That is a measurement and not a theorem: the bound is proved for the hottest rule and observed for this one, and 990 boards from a pool of ten positions is not a proof. But it is the right measurement to make, because a rule that scored better and broke the promise would be a rule with no promise at all — interesting, and not an improvement.
The control is what makes the check mean anything. Playing in the coldest component breaks the bound on 339 of the 715 four-component boards, so a bound nothing violates on this pool would be a bound the pool cannot test. It can.
What a rule of this shape is for
There is a temptation, on finding a rule that beats the standing one, to say the standing one has been replaced. It has not, and the reason is worth being precise about because it is the difference between the two things this ladder measures.
A guarantee is a promise about the worst case: whatever the board, a player following the hottest rule finishes within the largest temperature of the mean, and that is a theorem. A score is how often a rule matches perfect play on a population, and it is a statement about the population. The discount rule wins on the second and matches on the first only as far as this sweep can see.
For a player at a board those are two different kinds of use. A guarantee is what lets a player stop calculating: knowing the loss is bounded means the position can be left alone. A score is what makes a rule worth following when there is nothing better available. The right reading of this page is that a player who already follows the hottest rule should subtract the answer’s temperature when two components are close — it costs one more thermograph per component and it picks up a point roughly one board in ten.
And there is a third thing neither measures, which a bound instead of an answer is careful about: how far a rule is from perfect play when it is not exact. Both rules here lose at most one point on any board in the sweep, so on this pool that question has the same answer for both — and on a pool with larger components it would not.
Beating the hottest is not the same as being safe
The result here is a rule that outscores the one with the theorem, and this site has met that shape once before with an uncomfortable ending. It is worth saying what is different this time and what is not.
The greedy rule also outscored the guaranteed rule, on a different pool, and the lasting objection to it was that it had no bound: nothing said its worst case over all boards was finite, so a sample however large could not promise anything about a board nobody had tried.
This rule is in a better position on exactly that point. It stays inside the guarantee proved for playing in the hottest component — so its worst case is bounded, by an argument about the other rule, and a reader adopting it is not giving up the promise to get the average.
What it does not have is a bound of its own. The guarantee it respects was proved about a different rule, and this rule never exceeds that rule’s bound is an observation over 220 boards rather than a theorem. It could fail on a board outside the pool while the other rule’s own bound stands, because the bound is about the other rule.
So the honest statement has two halves and both matter. On everything measured this rule is better and safe; the safety is measured and the improvement is measured, and neither is proved. That is a materially stronger position than the greedy rule was ever in, and it is not the position the guaranteed rule is in — which is the trade a reader is actually being offered.
What this does not say
Ten positions, and they were chosen for something else. The pool was built by an earlier rung to punish a strategy that ignored follow-ups, so it is a pool with unusually large answers in it — which is exactly the property the discount rule exploits. A pool of plain switches would give both rules the same score, since a switch has no answer to discount. The right way to read the result is that the discount rule is better where follow-ups matter, and how often that is on a real board is not measured here.
It is not a proof of anything. The hottest rule’s guarantee is a theorem with a proof; this is a rule with a better score on 990 boards. What a proof would need is a statement of the form a player scoring by finishes within of the mean, and nothing here suggests what would be — the observed worst case is one point, which is also the hottest rule’s.
The margin is small and the pool is one. Five boards in 220, twenty-five in 715. A different pool of ten positions could reverse it, and the argument that it would not is the mechanism rather than the numbers: the follow-up is a liability, and a rule that prices it should beat a rule that ignores it wherever follow-ups exist.
And the discount is unweighted. charges the answer at full rate, and there is no reason it should be — half of the smaller temperature found a crossover correction that is exactly half the answer’s temperature, so is at least as natural a rule as this one and was not swept.
The convention, named
Normal play throughout, and every score computed by playing the board out.
A board is a sum of components, and the sweep is every unordered multiset of two, three or four components from this site’s ten-position strategy pool. A rule chooses which component to move in; the move inside the component is the same for every rule — the option leaving the best stop — so the rules differ only in the choice of component.
A rule is scored by playing it against optimal play: one side follows the rule and the other plays exactly, and the score is what the board finishes at. Plays exactly means the rule’s score equals what perfect play would have got.
A component’s answer is its hottest option that is not a number, and its temperature is the height at which that option’s own thermograph walls meet. A component with no such option has an answer of nought, so the discount rule reduces to the hottest rule on a board of plain switches.
Hotstrat’s bound is the board’s mean less its largest temperature. It is proved for the hottest rule; for every other rule on this page it is measured, and the tables say which is which.
Where the ladder goes next
The coupons anchor has six rungs to here, and this one has just found a ranking rule that beats playing in the hottest component. The rung above finds that its coefficient was the worst available.
The worst value in its own interval sweeps the discount weight and finds every value strictly between nought and one scoring the same, and all of them beating the choice of one at every board size. The reason applies to any rule of this shape and is worth more than the improvement: a ranking rule’s behaviour depends only on the order its scores put the components in, and that order changes only where two scores cross. So the score against perfect play is a step function of the coefficient, constant on each interval between crossings — and one is exactly a crossing, the point at which two components tie.
That is an uncomfortable place for a rule to have been standing. At a crossing the ranking is settled by whatever breaks ties in the implementation — the order components were built, the order they appear in a list — so the rule as measured here was partly a rule nobody wrote.
The instruction that follows generalises past this ladder and is short. Never choose a round number for a coefficient of this kind. Round numbers are where quantities coincide, coinciding quantities are where a ranking is decided outside the rule, and the safe choice is the middle of an interval — which a sweep like that one is exactly the way to find.
Part 6 of 9
One argument about Coupons. The parts either side of it:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.
AmbientApproximationBoundCoupon stackDisjunctive sumEnumerationFollow-upHeuristicMean valueSenteStrategyTemperature
- A pool built to punish greed bound, disjunctive sum, follow-up, heuristic, mean value, sente, strategy, temperature
- A schedule instead of a number approximation, bound, disjunctive sum, enumeration, mean value, strategy, temperature
- A second level of stops approximation, bound, enumeration, follow-up, heuristic, mean value, temperature
- A subtraction, not a factor ambient, bound, enumeration, follow-up, sente, strategy, temperature
- Half a follow-up out approximation, bound, enumeration, follow-up, heuristic, mean value, temperature
- Which end a sum lands at approximation, bound, disjunctive sum, enumeration, follow-up, mean value, temperature