Where it stops

The clause that turns the class off

Three rungs failed to find the dead-ending class doing measurable work, and each time the population was blamed. Toads and Frogs with and without the jump is the matched pair the anchor wanted — the same board with the class switched on and off — and on it the test the class licenses gains less from the class than a control that has never heard of it.

Assumes: Wrong in one direction only · What the class does not buy

Wrong in one direction only found every misère option test sound in one direction and wrong in the other, found the end proviso helping on one ruleset and hurting on two, and closed by saying why nothing could be concluded from that:

Three negative results on four rulesets is as far as this population can be pushed, and the fault is the population rather than the instruments: three dead-ending rulesets against one that is not cannot separate a hypothesis from a coincidence. What is wanted is one game with a clause that turns dead-endedness on and off.

Toads and Frogs has one. It is the jump, and it turns the class on and off exactly.

Who gains, and how much. How much each test improves when dead-endedness is turned on, with the class-specific test beside the others.
Fig. 1 How much each test improves when dead-endedness is turned on, with the class-specific test beside the others.

The clause does what it was hoped to do

Dead-endedness is the property that a player with no move never gets one back. It is checked here on strips rather than on values, because it is a property of the ruleset’s positions and two strips can share a value.

One clause, and the class. Whether a stuck player ever regains a move, with the jump allowed and forbidden.
Fig. 2 Whether a stuck player ever regains a move, with the jump allowed and forbidden.

Over every strip of at most seven squares: with the jump, 498 of 2,360 stuck positions later give the player a move back. Without it, nought of 3,184.

So the jumpless game is dead-ending and the jumping game is not, with the same squares, the same pieces, the same directions of travel and the same ending convention. One clause, and nothing else.

Where a stuck player recovers. Positions in which a player with no move later gets one, all of them by the jump.
Fig. 3 Positions in which a player with no move later gets one, all of them by the jump.

The mechanism is visible in the positions. A toad immediately in front of a frog is blocked: the toad cannot move right into the frog, and the frog cannot move left into the toad. Without the jump that is permanent — nothing else can come between them. With the jump the toad can leap the frog when the square beyond is empty, so a blocked player becomes unblocked, and the class fails.

That the clause is the mechanism rather than merely correlated with it is what makes the pair an experiment rather than a comparison of two rulesets that happen to differ.

The census, run twice

The census, on both. The three option tests scored against the quantifier on both versions of the game.
Fig. 4 The three option tests scored against the quantifier on both versions of the game.

The measurement is the rung below’s, unchanged. For each version: 244 strips of at most five squares, each pair of them compared by the quantifier — does ABA \geq B hold against every member of the ruleset’s own universe? — and by three tests.

The plain recursion is the normal-play comparison test transplanted: ABA \geq B unless some Right option of AA is at most BB, or some Left option of BB is at least AA. The end proviso is that recursion with the clause dead-endedness is supposed to license added: a Left end may only be at most another Left end. And the outcome control compares the two misère outcomes and nothing else — it knows nothing about ends, options or dead-endedness.

Over 59,292 ordered pairs on each version, the agreement rates are 66666{\cdot}6, 82682{\cdot}6 and 61561{\cdot}5 per cent with the jump, and 71771{\cdot}7, 83783{\cdot}7 and 62962{\cdot}9 without it.

Every test does better on the dead-ending version.

Two things about the population are worth stating, because the comparison is between two censuses rather than inside one.

The strips are the same words. Both versions are evaluated on the identical 244 strings over toad, frog and empty square, so the only difference between the two censuses is how a string is turned into a game. Nothing has been chosen for either side.

And each version is measured against its own universe. A ruleset’s universe is the set of positions it can produce, and misère comparison is a quantifier over that set — so measuring the jumpless game against the jumping game’s universe would be measuring something neither ruleset is about. The consequence is that the two agreement rates are not directly comparable as levels, which is why the page reports the change in each test rather than the height of it.

Which is exactly the wrong shape

Against a control that knows nothing. The class-specific test's gain against a control with no clause about ends.
Fig. 5 The class-specific test’s gain against a control with no clause about ends.

Everything improves is not evidence for the hypothesis; it is what a simpler game does to every instrument pointed at it. The jumpless game has a smaller universe — 210 positions against 276 — shorter lines of play and simpler values, and a test of any kind will find it easier.

What the hypothesis predicts is a differential: the end proviso is the test the class licenses, so it should gain more from the class than the tests that do not use it.

It gains least. The plain recursion gains 525{\cdot}2 points, the outcome control 151{\cdot}5, and the end proviso 121{\cdot}2.

A class-specific instrument gaining less than a class-agnostic control, on the one population built to isolate the class, is as clean a negative as this anchor can produce. It is not that dead-endedness fails to help; it is that the test built out of dead-endedness benefits from dead-endedness less than a test that has never heard of it.

What the quantifier says about the two games

One column of the census is worth reading on its own, because it is a fact about the games rather than about the tests.

Of the 59,292 ordered pairs, the quantifier accepts 14,470 on the jumping version and 17,032 on the jumpless one. So removing the jump makes positions more comparable: about a sixth more pairs are ordered.

That is the expected direction and it is worth having as a number. A misère universe with fewer positions has fewer witnesses with which to separate two games, so comparisons that fail in a large universe succeed in a small one — and dead-ending rulesets have well-behaved comparison could be nothing more than dead-ending rulesets tend to be small. Nothing on this anchor has ever separated those two readings, and the matched pair does not separate them either, because the clause that switches the class also shrinks the universe.

What that suggests is a test the anchor has never tried: hold the universe’s size fixed and vary the class. Whether that is possible at all is unclear — the class and the size are entangled by the same clause here — and it is the shape the next experiment would need.

And one thing does transfer

What does survive. The false-negative count on both versions, which is nought.
Fig. 6 The false-negative count on both versions, which is nought.

Over all 118,584 ordered pairs, on both versions, no test refuses a comparison the quantifier accepts.

That is the rung below’s finding surviving a change nobody had been able to make. One-sidedness was measured on four rulesets, three of them dead-ending, and could have been a fact about the class; it is not. It holds on a ruleset with the class and on the same ruleset without it, which is the strongest form of this is about misère comparison and not about dead-endedness available here.

So the anchor’s one usable result is the one that has nothing to do with its subject. A program can run any of these tests first and put the quantifier only to the pairs the test accepts, at no risk of a missed comparison, whether or not the ruleset is dead-ending. Equal in this company is where the quantifier’s cost is measured, and it is the work this filter removes.

Why the plain recursion gains most

The largest gain belongs to the test with no clause about ends in it at all, and it is worth asking why, because the answer is about the jump rather than about the class.

The plain recursion is a statement about options: ABA \geq B unless one of AA’s Right options is small enough or one of BB’s Left options is large enough. It is a normal-play theorem, and what breaks it under misère is that an end is a win rather than a loss, so the recursion’s base case is inverted relative to what the options say.

Removing the jump makes options scarcer and lines of play shorter, so a position has fewer options for the recursion to be wrong about and fewer levels for the wrongness to accumulate through. That is a reason for the recursion to improve which has nothing to do with ends, and it predicts exactly what the census finds: the option-reading test gains four times what the end-reading test gains.

Read that way the result is not merely a null. It is that the improvement has an identifiable cause, the cause is the number of options, and dead-endedness is a side effect of the same clause rather than the operative part of it. The rung below’s experiment was designed to hold everything but the class fixed, and the clause that switches the class turns out to switch something else as well.

That is the honest limitation of a matched pair with one clause: a single clause changes one thing only if the ruleset is simple enough that it does, and this one is not.

Four rungs, four negatives

It is worth putting the anchor’s record in one place, because a run of four is a different object from four separate failures.

Nobody comes back defined the class and found it a clean, checkable property of a ruleset. What the class does not buy measured misère comparison inside four universes and found the class explaining none of the counts. Wrong in one direction only found the class’s own test one-sided and helping on one ruleset in four. And this page finds it, on the matched pair, gaining less than a control.

Each rung blamed the population and proposed a better one. This is the better one, and it agrees with the other three.

At some point a run of negatives stops being a series of failed experiments and becomes a result. The result is: dead-endedness is a real property of a ruleset that does not measurably improve any comparison test on this site. That is a much duller sentence than the hypothesis and it is what four rungs of measurement have established.

What it does not say is that the class is useless in general. The literature’s interest in it is largely about misère quotients — the monoid of a ruleset’s positions modulo indistinguishability — and nothing here measures a quotient. Misère quotients is what the subject does instead of a general theory, and it is where the class might still earn its keep.

What the figures show and what they cannot

Every figure here is a table of rates, and there is a picture missing that this site would usually draw.

It is a strip. TF. with the frog stuck, and then the same strip with the jump allowed, where the toad leaps and the frog is stuck no longer. Two rows of three squares would carry the entire mechanism — why the jump is dead-endedness’s off switch — better than the 498 in the table does.

What the tables carry instead is the census, and there the count is the argument: one position recovering would not make the class fail, and 498 of 2,360 is a ruleset in which recovery is ordinary. The same strip without the jump draws these strips and is where a reader should go for the picture.

The second thing not drawn is the quantifier. Does ABA \geq B hold against every member of the universe is a loop over 210 or 276 positions, each a sum and an outcome, and there is no way to draw the loop that is not a list of two hundred outcomes. That is a genuine limit rather than a choice, and it is why this whole anchor reports percentages.

What this does not settle

Two rulesets, and they are the same ruleset. That is the experiment’s strength and its limit: everything is held fixed, and everything held fixed is Toads and Frogs. Whether the differential is nought on some other switchable pair is untested, and there may not be another one — a clause that turns dead-endedness on and off while changing nothing else is a rare thing, which is why the anchor spent three rungs without one.

Strips of at most five squares. 244 positions and 59,292 ordered pairs on each version, which is two orders of magnitude more than the rung below’s 492 pairs and is still a small game. The quantifier’s cost is what bounds it: it compares two positions against every member of the universe, and the universe grows as the square of the reachable set.

The universes are different. Each version is measured against its own universe, which is the only meaningful comparison — a ruleset’s universe is made of its own positions — and it means the two censuses are not measuring the same quantifier. That is unavoidable and it is the reason the differential rather than the level is what the page reports.

And the end proviso is one proviso. It is the condition the rung below wrote from the class’s definition, and a different condition drawn from the same class might behave differently. What has been refuted is that this test, which is the natural one, gains from the class.

The differential is small in absolute terms. 121{\cdot}2 against 151{\cdot}5 points is three tenths of a point, on 59,292 pairs. What makes it a result is the direction and not the size: the hypothesis predicts a positive differential and the measurement gives a negative one, and a hypothesis that predicts the sign wrong is not rescued by the magnitude being small. What a larger population would buy is confidence that the sign is real, and that population is the same census on six-square strips.

And the dead-endedness check is over strips, not values. Two strips with the same value can differ in whether a stuck player recovers, because recovery is about the moves and the value is about the outcomes of sums — so the property is checked on the 3,184 strips of at most seven squares rather than on the games they produce. That is the right population for a statement about a ruleset, and it means the 498 recoveries are a count of positions rather than of values.

Misère play throughout, and normal play mentioned only as the source of the test being transplanted. Under normal play the recursion over options is a theorem and there is nothing to measure.

One sentence a reader can take away

Strip the anchor down to what it has actually established and there is one sentence in it.

A misère option test is a sound filter and not a characterisation, on every ruleset tried, whether or not it is dead-ending. It never refuses a true comparison; it accepts many false ones; and the share it accepts is 41 to 64 per cent of the pairs, so putting the quantifier only to those removes between a third and three fifths of its work.

That sentence has a use, it has been tested on six rulesets now — four from the rung below and this matched pair — and it says nothing about the class the anchor is named after. Which is a fair summary of four rungs: the useful thing found is not the thing looked for, and the thing looked for has been given every chance it could be given.

Where the ladder goes next

The dead-ending anchor has four rungs: nobody comes back, what the class does not buy, the test it was supposed to license, and now the clause that turns it off.

The rung above is the quotient. Every measurement on this anchor has been about comparison — whether one position is at least another in a universe — and the class’s actual reputation is about quotients, which are a coarser object: two positions are identified when no sum distinguishes their outcomes. A quotient can be small where comparison is hard, and the matched pair is exactly the population on which to ask whether dead-endedness makes it smaller. That is the same two rulesets, the same strips, and a different quantity, and it is the last question this anchor has that its own subject is actually about.

Two neighbours are worth the trip. The same strip without the jump is where this pair of rulesets is set out under normal play and where the clause’s effect on values is measured, and it is worth reading beside a page that uses the same pair for a completely different question. And two misère outcomes are not enough is why any of this is hard, and it is the page that makes a filter worth having at all.

Part 4 of 5

One argument about Dead-ending. The parts either side of it:

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.

ComparisonCounterexampleDead-endingEnumerationHeuristicMisere playMisère quotientNormal playOutcomesPartial orderToads and FrogsUniverses