The worst value in its own interval
Assumes: A rule that beats the hottest · The quantity that does not order a board
A rule that beats the hottest scored a component by its own temperature less its hottest answer’s, played where that is largest, and beat Hotstrat at two, three and four components. It closed on the coefficient it had assumed without argument:
The rung above is the weight. The rule charges the answer at full rate and there is no argument that it should — the sente ladder found a correction of exactly half an answer’s temperature in a different setting, and a rule scoring has a free parameter that the same sweep can fit. Whether the best is one, a half, or something the pool decides is a measurement of exactly the shape this page already runs.
The pool decides, and the coefficient the rung below used is the worst value in its own range.
Every weight inside the interval scores the same
Sweep over the same boards the rung below used — every unordered multiset of components from this site’s ten-position pool, at two, three and four components — and the result is a plateau:
| 2 components | 3 | 4 | |
|---|---|---|---|
| (play in the hottest) | 53 of 55 | 196 of 220 | 606 of 715 |
| any in | 55 | 205 | 641 |
| (the rung below’s rule) | 54 | 201 | 631 |
| 53 | 196 | 616 |
A quarter, a half and three quarters give identical scores at all three sizes. So the question is it one, a half, or something the pool decides has an answer, and the answer is that it is not a number at all: it is an interval, and the sweep cannot tell one point of it from another.
That is not an evasion. It is what a fitted parameter looks like when the thing being fitted is a ranking.
Why the score is a step function
The rule does not use ; it uses the order of across the components on the board. Two components with temperatures and answers swap places at
so the whole set of coefficients at which the rule can behave differently is finite and computable from the pool without playing a single game. On this pool it is .
Between crossings the rule is the same rule, so the score is constant. That is why a fine grid finds steps rather than a curve, and it is why fitting a coefficient here is a different activity from fitting one to a continuous quantity: what is being located is an interval between crossings, and reporting its midpoint as the answer would be inventing precision.
It also explains the dip. At two components score exactly equal, and a tie is broken by whichever comes first in the list — which is an arbitrary choice made by the code rather than by the rule. So the rung below’s rule sits precisely on the one coefficient in its own interval where the rule stops being a rule and becomes a coin toss. Four boards at three components and ten at four turn on it.
The moral is small and general: a ranking rule with a free coefficient should never be evaluated at a crossing, and the crossings can be computed before any game is played.
At two components it never loses a point
One result is stronger than a share and deserves separating.
At two components, every inside the interval plays exactly on all 55 boards, with a worst case of nought — it never gives away a single point. Hotstrat’s worst is one point, and the rung below’s full-rate rule is also one.
On this pool and at this size, the weighted rule is optimal play. That is a claim about 55 boards built from ten positions and not a theorem, and the two-component case is the one where a rule has least room to go wrong. But no other rule this ladder has measured has a worst case of nought at any size, and it is worth noticing which one does.
At three and four components the worst case is back to one point, which is the same as Hotstrat’s. So what the weight buys is a larger share of boards played exactly, not a better guarantee — the rule is right more often and no less wrong when it is wrong. That is the ordinary shape of a strategy result on this site, and when to leave the environment is where the distinction between the two was first drawn.
The other side of nought
The interval is bounded below as well as above, and the lower bound is a disaster measured two rungs ago.
A negative promotes a component by its answer’s temperature — play the fights with the biggest follow-ups first — which is the rule the quantity that does not order a board tried and lost with. At it plays exactly on 115 of 220 three-component boards, loses nine points on the worst board, and falls outside Hotstrat’s guarantee on 13 boards at three components and 118 at four.
That last is the sharpest thing in the whole sweep. Every non-negative weight, from nought to two, finishes within the bound proved for playing in the hottest component on every board at every size. The negative ones do not. So the guarantee is not a property of the pool being forgiving; there is a rule in easy reach that violates it, and the sweep found it.
What the weight means, if it means anything
A coefficient with no interpretation is a fitted number, and this ladder has been careful about the difference, so it is worth asking what would mean if the sweep had picked one out.
The rule’s score is the stake, less something for the answer. The stake is what a player takes by moving here now. The answer is what the opponent takes back immediately — so a component with a hot answer is one where the gain is temporary, and is how much of the answer the mover expects to lose.
At the mover expects to lose none of it, which is Hotstrat’s implicit assumption and is why Hotstrat is beaten. At the mover expects to lose all of it, which over-corrects: the opponent has a whole board to play on and will not always take the answer immediately. Somewhere between is the opponent takes the answer some of the time, and every value in between is consistent with that reading.
So the plateau is not an embarrassment for the interpretation; it is what the interpretation predicts. The share of the time an opponent actually answers depends on what else is on the board, which is a fact about the rest of the board rather than about the component, and a rule that reads only the component cannot know it. A single correct would exist only if that share were constant, and there is no reason it should be.
That is a better account of the plateau than the pool is coarse, and the two are not exclusive. The pool being coarse is why the plateau is the whole interval rather than part of it; the interpretation is why an interval is the right kind of answer.
What this says about the half
The rung below hoped for a coefficient of a half turning up here and on the sente ladder, and named the reason: a λ that came out at a half on both ladders would be worth more than either finding alone.
A half is inside the plateau. Nothing here contradicts it, and nothing here supports it either — a quarter and three quarters score identically, so the measurement has no opinion between them. A sweep that cannot separate two hypotheses is not evidence for whichever one arrived first, and the honest report is that the coincidence remains available and unconfirmed.
There is a way it could be confirmed, and it is worth stating because it is what the pool is missing. The crossings are set by the pairs the pool holds, and this pool holds ten of them with temperatures in and answers in . A pool whose crossings fell densely across the unit interval would narrow the plateau, possibly to a point; this one has three crossings and two of them are below a quarter. The plateau is wide because the pool is coarse, which is a fact about the instrument and not about the rule.
The two numbers at the top is where the sente ladder’s half comes from, and it is worth reading beside this page for the contrast in what a pool can settle. There the claim was about which inputs a formula reads, and a grouping test answered it outright; here the claim is about a coefficient’s value, and a pool of ten positions cannot resolve it better than to an interval.
Three rules, one line apart
Standing back, the whole of this ladder’s strategy work is now one expression with one number in it, and that is worth writing down because it was three separate rules until this page.
is Hotstrat, which is the textbook rule and has the only proved guarantee. is the rule the quantity that does not order a board built out of the departure temperature, and it is the worst thing on the ladder. is the discount rule, which beat Hotstrat and is a tie boundary. And anything strictly between nought and one is better than all three.
Three essays’ worth of separate rules turn out to be three points on one line, which is a small piece of consolidation and is the kind that only appears once somebody parameterises. It also reorders the ladder’s own history: the rung two below was not a wrong idea but a sign error, and the rung below was a right idea with the coefficient pushed one step too far.
The measurement that would have found this at the second rung rather than the seventh is exactly the one run here, and it costs a sweep the ladder had already paid for. That is worth carrying: when a rule is a quantity, corrected by another quantity, the correction has a coefficient whether or not anybody wrote one, and sweeping it is cheaper than proposing a second rule.
Why a ranking rule’s score is a step function
The finding here looks like a curiosity about one coefficient and it is a general fact about scoring rules of this shape, worth stating in the form that applies to any of them.
A ranking rule with a coefficient scores each component by some expression in and plays in the highest-scoring one. What the rule does on a given board depends only on the order the scores put the components in, and that order changes only when two scores cross. Between crossings the order is constant, so the rule makes identical moves, so its score against perfect play is identical.
So the rule’s performance is not a smooth function of that could be optimised; it is a step function whose steps are at the crossing points. Every value strictly between two consecutive crossings gives exactly the same play on every board in the pool, which is why an entire open interval scores alike.
And that says why the endpoints are the dangerous places to stand. A crossing is where two components tie, and a tie has to be broken by something — the order they appear in, the order they were built, whatever the implementation happens to do. A coefficient sitting exactly on a crossing is a coefficient whose behaviour is decided by a tie-break rather than by the rule, and a tie-break chosen by accident is a rule nobody wrote.
The practical instruction is short and applies well beyond this ladder: never choose a round number for a coefficient of this kind. Round numbers are exactly where quantities coincide, and coinciding quantities are exactly where a ranking is decided by something outside the rule. Pick the middle of an interval, and the interval is what a sweep like this one is for finding.
What this does not say
The plateau is this pool’s, not the rule’s. Every number here comes from ten positions chosen, in an earlier essay, to punish a strategy that ignores follow-ups. A different pool would have different crossings and a narrower or wider plateau, and might place somewhere harmless.
A crossing is a property of a pair, not of a board. The three crossings are computed over all pairs from the pool; on a particular board only the pairs actually present matter, so most boards have no crossing at all inside the unit interval and the rule is the same rule on them for every weight.
Nothing here is a bound. The rule’s worst case is one point at three and four components, exactly as Hotstrat’s is, and the fact that every non-negative weight stayed inside Hotstrat’s guarantee on every board of this sweep is a measurement of 990 boards rather than a proof about any of them. Playing the hottest is where that guarantee is actually proved, and it is proved for one rule.
And a step function is not a fit. Reporting the plateau’s endpoints as and is a statement about the grid; what the crossings say is that the true endpoints are nought and one, open at both ends. The grid is here to show the shape, and the shape is explained by the pool rather than measured from it.
The convention, named
Normal play throughout, and every game played out exactly with one side following the rule and the other playing optimally.
A component’s temperature is the height at which its thermograph’s walls meet, and a number is counted as cold, at temperature nought. Its answer is its hottest option that is not a number, and the answer’s temperature is .
The weighted rule plays in whichever component maximises , and then makes the move in it that leaves the best stop. is playing in the hottest component; is the rung below’s discount rule; a negative is the rule two rungs below.
A board is played exactly when the rule’s final score equals what optimal play achieves. The worst board is the largest number of points the rule gives away anywhere in the sweep.
Hotstrat’s guarantee is that a player following the hottest rule finishes no more than the board’s largest temperature below the board’s mean. It is proved for that rule; whether another rule stays inside it is a measurement, and it is reported as one.
The pool is ten positions and the boards are every unordered multiset of two, three or four of them — 55, 220 and 715 boards.
Where the ladder goes next
The coupons anchor has seven rungs: the environment, when to leave it, two games in one environment, how big the answer is, the quantity that does not order a board, the rule that beats the hottest, and now the rate it charges.
The rung above is a pool with dense crossings. Everything about the plateau’s width is a property of the ten positions, and the crossings are computed from the pairs before any game is played — so a pool can be designed to have crossings wherever they are wanted, and a pool with twenty crossings spread across the unit interval would narrow the plateau to whichever cell of it is best. That is a construction rather than a sweep, which is a kind of work this ladder has not done, and it is the only route to a number rather than an interval.
Two neighbours are worth the trip. How big the answer is is where the answer’s temperature first became a quantity worth reading, and it is the reason a rule reads two temperatures rather than one. And a schedule instead of a number is the other place on this ladder where a rule turned out to need a parameter, and the two together are the site’s whole account of what a strategy costs beyond the temperature.
Part 7 of 9
One argument about Coupons. The parts either side of it:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.
ApproximationBoundCoupon stackEnumerationEnvironmentError termFollow-upHeuristicHot positionMove selectionStrategyTemperature
- A pool built to punish greed bound, follow-up, heuristic, hot position, move selection, strategy, temperature
- A rule with a guarantee error term, follow-up, heuristic, move selection, strategy, temperature
- A rule with no promise at all approximation, error term, heuristic, move selection, strategy, temperature
- A second level of stops approximation, bound, enumeration, follow-up, heuristic, temperature
- Half a follow-up out approximation, bound, enumeration, follow-up, heuristic, temperature
- A subtraction, not a factor bound, enumeration, follow-up, strategy, temperature