Temperature

A second pool, designed differently

One designed pool put the rule's best coefficient between a quarter and a third, and all three board sizes agreed. A second pool, built by the identical greedy criterion from different material, has no cell that is best at every size — so the coefficient is a property of the pool and there is no number to find.

Assumes: A pool built to have an answer · The worst value in its own interval

A pool built to have an answer took a rule with a free parameter — score a component by t − λa and play where that is largest — and did something to make the parameter measurable. On an ordinary pool every λ strictly inside the unit interval scores identically, because the rule reads only an ordering and an ordering changes at finitely many places. So that page designed a pool: seven positions chosen greedily to make the ordering change as often as possible, twelve crossings, thirteen cells, and one cell that beat the rest at two components, three and four alike.

The cell was between a quarter and a third. And that page said, in as many words, what it did not know:

One designed pool gives one answer, and the way to find out whether the answer is about the rule or about the pool is to design another … If the best cell of the second pool overlaps (1/4, 1/3) the coefficient is a property of the rule; if it does not, then the rule’s parameter is a per-pool quantity and there is no number to find, which would be a cleaner negative result than any amount of further sweeping.

The second pool has been built. It does not overlap, and it does something worse than not overlapping.

The second pool

A second pool, built the same way from different material. The first designed pool drew its positions from temperatures 1 to 12 against answers 0 to 6. This one draws from temperatures 3 to 15 against answers 0 to 7, chosen by the identical greedy criterion, so that any disagreement between the two is about the material and not about the construction.
Fig. 1 The seven positions of the second pool, with the temperature and hottest answer each contributes.
Thirteen cells, thirteen scores. The rule scored in every cell of the unit interval on the designed pool, at three components.
Fig. 2 The first designed pool’s cells, from the rung below: three sizes converging on one cell of width one twelfth, which is what made its answer look like an answer.

The construction is deliberately identical: the same family of positions {T{AA}}\{T \mid \{A \mid -A\}\}, whose temperature and hottest answer are read off T and A; the same greedy rule, taking at each step the candidate that adds the most distinct crossings; the same tie-break; the same size of seven.

Keeping the construction fixed is the whole design of the experiment. If the second pool used a different criterion — crossings spread evenly rather than crossings maximised, say — a disagreement could be about either the material or the criterion, and there would be nothing to conclude. Changing exactly one thing is what makes the answer readable.

What changes is the grid the candidates come from. The first pool drew on temperatures 1 to 12 against answers 0 to 6; this one draws on temperatures 3 to 15 against answers 0 to 7. The resulting pools have no position in common and no (t,a)(t, a) pair in common.

They land on twelve crossings each, which is the greedy criterion doing its job, and seven of the twenty-four crossings coincide. That coincidence is worth naming so it is not read as similarity: both pools are made of ratios of small integers, and small rationals collide. A third, a half and two thirds turn up in any such pool.

Where they differ is at the bottom. The first pool’s lowest crossing is a sixth; the second’s are a twelfth, an eighth and a seventh, so the second pool cuts the region below a sixth into four cells where the first pool has one. That matters for what follows, because three of the second pool’s winning cells are down there — a region the first pool could not see into at all, since a cell is only as fine as the crossings that bound it.

What the sizes say

The three sizes do not agree. The first pool's answer was credible because two, three and four components all pointed at the same cell of width one twelfth. On the second pool, built the same way, they point at three cells that do not overlap — so this pool has no answer of its own to compare with the first one's.
Fig. 3 The rule scored at the midpoint of every cell of the second pool’s interval, at two, three and four components, with the cells that reach the best score.

At two components the rule plays exactly on all twenty-eight boards, and it does so in four places: the cells (5/12, 3/7), (3/7, 1/2) and (1/2, 2/3), and at λ = 1/2 itself.

At three components the best is sixty-nine of eighty-four, reached in the cells (0, 1/12) and (1/8, 1/7) — down at the bottom of the interval, nowhere near a half.

At four components the best is a hundred and fifty-nine of two hundred and ten, reached in (1/7, 1/6) alone.

The scores themselves are worth noting alongside the cells. On the first pool the winning cell played exactly on eighty of eighty-four boards at three components and a hundred and ninety of two hundred and ten at four — better than nine in ten. Here the bests are sixty-nine of eighty-four and a hundred and fifty-nine of two hundred and ten, around three quarters. So the second pool is harder for the rule as well as less decisive about it, which is what a different greedy walk through a different grid has no reason not to be: the criterion maximises crossings, not difficulty, and the two come apart.

No cell is best at more than one size, and no two of the winning regions overlap. Two components wants about a half, three wants nearly nothing, four wants a seventh. The second pool selects nothing.

That is the finding, and it is stronger than the test was designed to detect. The rung below asked whether the second pool’s answer would agree with the first’s; the second pool has no answer to disagree with. What made the first pool’s number credible was precisely that three independent board sizes converged on one cell of width one twelfth, and that convergence does not survive a change of material.

The two-component row deserves a second look, because it is the one that would be easiest to over-read. Twenty-eight of twenty-eight is a perfect score, and it is reached in four places at once — three adjacent cells and a crossing. A perfect score reached across a wide region is not evidence about where the optimum is; it is evidence that two components are too few to distinguish anything, which is the same reason the ladder went to three and four in the first place. The rows that carry information are the second and third, and those disagree.

The first pool’s answer, seen from here

It would be easy to over-read the disagreement, so the first pool’s answer is priced on the second pool’s boards directly.

The first pool's answer, priced here. The coefficient the first designed pool selected is scored on the second pool's boards. It does well and it is never the best cell at any size, which is the outcome that neither confirms nor refutes it — and is what a per-pool quantity looks like from outside its own pool.
Fig. 4 λ between a quarter and a third, scored on the second pool, against the second pool’s own best at each size.

A quarter to a third lands in the second pool’s cell (1/6, 1/3), and there it scores twenty-five of twenty-eight, sixty-eight of eighty-four, and a hundred and fifty-one of two hundred and ten. The pool’s own bests are twenty-eight, sixty-nine and a hundred and fifty-nine.

So the first pool’s coefficient is not bad here. It is three boards short at two components, one short at three, eight short at four — respectable, and never best.

That is the outcome that neither confirms nor refutes, and it is what a per-pool quantity looks like from outside its own pool. A coefficient that were a property of the rule would win here too; a coefficient that were an artefact would do badly. Doing well and not best is the signature of a parameter whose optimum moves with the material while the rule itself stays roughly right across the whole middle of the interval.

It is also worth checking the comparison the other way, and the shape of the answer is the same. The second pool’s own best cells — around a half at two components, near nought at three, a seventh at four — are three different λ, and each of them would score somewhere between well and badly on the first pool depending on which one was chosen. There is no pair of numbers here, one per pool, that could be averaged into a compromise; there is one number from a pool that produced one, and three from a pool that did not.

What is actually established

Two pools, and no coefficient. The previous rung asked whether the rule's best coefficient is a property of the rule or of the pool it was measured on, and said that a second designed pool would settle it. It does. The second pool's three sizes disagree with each other, so it selects nothing, and the first pool's answer is close but never best on it.
Fig. 5 The two pools’ verdicts side by side.

There is no number to find. The rule’s parameter has no pool-independent optimum in the sense this ladder was looking for, and the search for one should stop. That is the cleaner negative result the rung below named in advance, and naming it in advance is what makes it a result rather than a failure to find something.

The difference between those two is not rhetorical. A page that had gone looking for the coefficient, failed, and concluded that more pools were needed would be a page reporting an absence of evidence. The rung below specified the experiment, said what each outcome would mean, and ran it — so this page reports evidence of absence, on a test it did not design after seeing the result.

The rule is not thereby damaged. Everything the ladder established about the rule stands: it beats playing the hottest component, λ = 1 is the worst value in the interval on both pools, and the whole middle of the interval scores well on both. What has failed is the attempt to pin a best λ, not the claim that some λ in the middle is much better than the ends.

The rule’s ordering, not its arithmetic, is what is being fitted. That is worth restating because it explains why the parameter can be so unstable while the rule is so robust. λ enters only through which component scores highest, so two λ in the same cell are literally the same rule — and the cells are cut by the pool. A parameter that only ever appears inside an ordering has no meaning apart from the set of things being ordered, which is the same difficulty a value has apart from its company and is not a new kind of trouble.

And a designed pool is a discriminating instrument, not a sample. That was said on the first pool and it is the thing this page turns from a caveat into a measurement. Both pools were built by choosing positions to make the rule’s ordering change as often as possible, which is exactly what makes them unrepresentative — and two instruments built to be maximally discriminating discriminate differently. Between them they say the discrimination is about the instrument.

What is not established is that no coefficient exists in any sense. Two pools is two, and both are designed the same way. It remains possible that averaging over many pools would show a mode, or that a pool built from positions of a different family — not {T{AA}}\{T \mid \{A \mid -A\}\} — would behave differently. What is established is that the specific evidence the ladder had for a quarter to a third does not reproduce, and evidence that does not reproduce is not evidence.

Why three sizes agreeing was the wrong evidence

The first pool’s case rested on an agreement across board sizes, and it is worth saying why that looked like strong evidence and was not.

Two, three and four components on the same pool are not independent experiments. They are drawn from the same seven positions; a four-component board contains three-component boards; and a cell that orders those seven positions well will order every board made of them well. So the three sizes agreeing is close to one experiment reported three times, and its agreement measures the pool’s internal consistency rather than the rule’s.

The second pool makes that visible in the sharpest way available: on it, the three sizes disagree. Whatever produced the first pool’s convergence was not a general feature of the construction, since the same construction on different material produces the opposite.

That is a lesson about the instrument rather than about coupons, and it is the kind this site keeps meeting — a criterion fitted to a small sample looks exact until the sample changes. The difference here is that the sample was designed rather than found, which makes the failure sharper and the diagnosis easier: nothing was overlooked, the pool was doing exactly what it was built to do, and what it was built to do turns out not to answer the question.

There is a general form of the mistake and it is worth stating, because it is available to any measurement on this site. Agreement between measurements is only evidence when the measurements could have disagreed for reasons independent of what is being measured. Board size is not such a reason on a fixed pool: a bigger board is more of the same positions. Two errors that cancel on the Domineering anchor is the same shape from the other direction — two quantities agreeing because they are built from the same thing rather than because both are right.

What survives, priced

It is worth ending on what a player is left with, since the ladder began with a rule somebody might use.

The coefficients, compared. The best cell against the two coefficients the ladders have proposed and against playing the hottest.
Fig. 6 The two coefficients the ladders have proposed, scored against the best cell on the first pool. Both are beaten, and λ = 1 is beaten badly.

The rule score a component by t − λa and play where that is largest is worth using, for any λ in the middle of the unit interval. Both pools agree that the ends are bad: λ = 1, which charges the answer’s temperature in full and is the coefficient the sente ladder proposed, is the worst value in the interval on both, and λ near nought reduces the rule to playing the hottest component, which the ladder already beat.

What neither pool supports is any particular middle value. The honest instruction is charge something for the answer, less than the whole of it, and the ladder cannot say how much.

That is less than it hoped for and it is not nothing. A rule with a free parameter whose whole middle range works is a more robust rule than one that needs a tuned constant, and a player who does not have to remember a fraction is better off than one who does.

What two designed pools can and cannot settle between them

Both pools here were built to discriminate, and a pool built to discriminate is not a sample of anything.

That is not a flaw in either. A discriminating pool is the right instrument for the question can this coefficient be pinned down at all, because a pool of ordinary positions would answer with a wide interval and no information about whether the width is the rule’s or the population’s. What two such pools buy, disagreeing, is the knowledge that the width belongs to the rule.

What they cannot buy is the number a player would want. Neither pool is representative and neither was meant to be, so the interval each produces is an interval about that pool. Reporting them together is the honest form of the result: one designed pool narrows the coefficient and a second, built by the identical criterion from different material, does not agree with it — which settles that the coefficient is not a constant of the rule and leaves open what it is on a board somebody might meet.

Where the ladder goes next

The coupons anchor has nine rungs: the environment, when to leave it, two games in one environment, how big the answer is, the quantity that does not order a board, the rule that beats the hottest, the rate it charges, the pool that can measure the rate, and now the second pool that says the rate is not a number.

Both pools are instruments rather than positions, and the instrument they are calibrating is a rule. A rule with a guarantee is the bound the whole anchor is measured against, and an environment made of coupons is where the environment itself was first written down as a game.

The rung above is the ordinary pool. Both pools here are designed to discriminate and neither is representative, so the question they cannot answer is the one a player would ask: over positions somebody might actually meet, is any λ in the middle of the interval better than any other? An unrepresentative pool cannot say and a plateau makes an ordinary pool unable to say either — but a large ordinary pool has crossings too, thinly spread, and averaging over many such pools would give a distribution of best cells rather than a single one. Whether that distribution has a mode, and where, is a measurement neither of these two pools can be asked for.

Two neighbours are worth the trip. A rule that beats the hottest is where the rule was shown to be worth having at all, and it is what survives this page intact. And the worst value in its own interval is where the plateau was first measured and where λ = 1 was found to be the worst choice available — a finding that holds on both pools, and the only thing about the coefficient that does.

Part 9 of 9

One argument about Coupons. The parts either side of it:

The objects named here

The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.

CouponsDisjunctive sumEnumerationExhaustive searchHeuristicHot gameMean valueStrategyTemperatureThermograph