The criterion that cannot exist
Assumes: The reading that survives too much · Nothing worth fighting over
Push is Shove with a wall instead of a cliff: a coin moved one square left takes the contiguous run in front of it along, and the move is legal only if that run has an empty square to move into. The reading that survives too much tested the count a newcomer writes down — the empty squares each coin can still travel over, signed by colour — and found a condition that guarantees it: the reading is right whenever no position either player can force ever brings two coins of opposite colour together.
That condition is sound on all 2,186 strips of at most seven squares and covers 254 of the 1,114 positions the reading handles. The other 860 mix and come out right anyway, and the page closed by asking why:
The reading survives mixing three quarters of the time, and the plausible reason is that a blocking exchange is often symmetric in the room it costs the two sides. Turning often into a condition on the strip is the piece of work.
The piece of work turns out to be impossible in the form it was asked for, and three strips of four squares are the whole proof.
Three strips
.LLR — an empty square, two of Left’s coins, one of Right’s — reads as one. It is worth .
.LRL reads as one. It is worth .
.RLL reads as one. It is worth .
Every quantity a criterion of the proposed kind could use is identical across the three. They are four squares long. They contain two coins of one colour and one of the other. They have one run, of length three, and it is mixed. Their readings are one. Their errors are , and .
So no function of the length, the reading, the coin counts, the run lengths and the count of mixed runs can say how far out the reading is — because such a function would have to give three answers to one input.
How general the obstruction is
The three strips are not a lucky corner. Group all 2,186 strips by that whole bundle of local statistics and there are 805 classes; 207 of them contain more than one error.
A quarter of the classes is enough. A criterion that worked on the other 598 and gave the wrong answer on 207 would not be a criterion; it would be the reading again, with a second reading of the same kind bolted to it.
The census asserts this. If every class carried a single error, a local criterion would not be ruled out and this page would be about finding one rather than about the impossibility, so a sweep in which the split vanished stops the build.
What is left over, and it is the interesting part
What separates .LLR from .LRL is the order of the colours inside the run, and that is not a count of anything.
Read the three strips as words — LLR, LRL, RLL — and the natural object is a binary expansion. Hackenbush is a numeral is the essay about a row of two colours spelling a number in binary, where the first colour gives the integer part and every colour change after it halves the step. That is a construction in which the order is everything and the counts are nothing, which is exactly the shape the errors here have.
This page does not claim Push values are Hackenbush numerals. They are not: a Push strip packed against the wall is worth nought whatever its colours, because no coin can move, and the numeral of LLR is . The empty squares are load-bearing in a way the Hackenbush word has no room for.
What the comparison establishes is the kind of object a repaired reading has to be. A count of room can be computed left to right, one coin at a time, and forgets everything except a running total; a numeral cannot, because each colour change halves the significance of everything after it. The rung below’s proposed criterion was of the first kind, and the answer is of the second.
How badly it fails
The other half of the question is the size, and the answer is that the reading is not a near miss.
At three squares the reading is wrong on 2 of 18 strips, and never by more than a half. At seven it is wrong on 790 of 1,458 — a majority — and the worst error is 6.14.
Both numbers move in the same direction at every length, which is the combination that rules out the two comfortable readings of a heuristic. It is not coarse but nearly right, because the errors grow. It is not right on the cases that matter, because the share wrong grows too.
The worst case is worth looking at, because it shows the mechanism at full size. ...LLLR gives Left three coins with three squares of room each and Right one coin behind them. The reading counts in Left’s favour. The value is .
Right’s single coin is behind Left’s three, and pushing it moves all four. The run in front of Right’s coin is Left’s entire block, so every push Right makes spends Left’s room as well as Right’s own — and the reading, which credits each coin with its own room, has no way to express that a move by one player consumes the other player’s assets.
What Right’s move actually does
The worst case is worth drawing rather than describing, because the mechanism is visible in one move.
Right has exactly one legal move from ...LLLR, and it shifts the whole block: Right’s coin and Left’s three, one square towards the wall. Right spends one square of Right’s own room and one square of each of Left’s.
The reading has no vocabulary for that. It was built by walking the strip from the wall outward and charging each coin the empty squares nobody ahead of it has claimed, which is a correct accounting if room is private. Under Shove it is private, and the reading is a theorem. Under Push a run is a shared object and one player’s move spends the other’s assets, and no left-to-right accounting can express a liability that has not been incurred yet.
That also explains why the error is a fact about order. Which coins get dragged along depends on which colours sit where inside the run, and nothing about the run’s length or its composition says who is in front of whom.
The one thing that does hold
There is a bound, and it is the only statement about the error that survives the whole census.
The error is at most the total room on the strip — the empty squares in front of each coin, added up unsigned — on all 2,186 strips with no exception. The census asserts it.
It is also nearly useless, and honestly so. The total room is roughly the size of the reading itself, so the bound says the reading may be wrong by about as much as it claims. The closest any strip comes to attaining it is 67 per cent, on .LLLLLR, so it is not tight either.
That is worth stating rather than suppressing. A rule that is never right and cannot be far wrong is this site’s account of what a useful approximation looks like — an error term small enough to plan against — and the point of quoting the bound here is that Push’s reading fails that test in the most complete way available: the only bound is the size of the answer.
Where the failures concentrate
One more reading of the same census is worth having, because it says the failures are not a large-strip phenomenon that a player would notice.
The condition that guarantees the reading covers a shrinking share of the strips, which is what a condition about no reachable interaction has to do: longer strips have more lines of play and more chances for the colours to meet. What is striking is that the reading’s accuracy falls much more slowly than the guarantee does, which is the slack the rung below noticed and named.
This page’s finding is what that slack is made of. It is not a weaker condition waiting to be found; it is the difference between two quantities that are not functions of each other. The guarantee is about the game tree, the reading’s accuracy is about a binary expansion, and neither predicts the other except by accident.
Why Shove is different, in one clause
The comparison that makes all of this legible is the same reading on the same strips under the other rule, where it is a theorem.
In Shove a coin takes every coin to its left along and whatever stands on the edge falls off, so a move is always legal and a coin’s room is never spent by anybody else — it is spent by the coin’s own owner or it is not spent at all. Under that rule the count of room per coin is exactly right on every strip, and the values are all numbers with a formula.
Push changes one clause — a wall instead of a cliff — and the change makes room shared. Everything on this page follows from that: the reading fails, it fails by an amount that depends on the order of the colours in a shared run, and the amount is bounded only by how much room there was to share.
What the census does not say
Four limits.
Seven squares. The impossibility argument needs only four, so the length bound is not doing any work there. What it does bound is the growth claims — the share wrong and the worst error at each length — and those are seven points on a curve rather than a trend established.
The statistics are a chosen list. No local criterion here means no function of the six quantities named. A criterion using something else about the strip — the positions of the empty squares, say, or the pattern of runs from the wall outward — is not ruled out, and one of those is what a numeral reading would be.
The bound is the only rule and it is checked. The error being at most the room holds on 2,186 strips and is not proved. Its plausible mechanism is that every push moves at least one coin one square, so no line of play can spend more room than there is, which is a sketch and not an induction.
And every Push value in the census is a number. That is why the reading can be compared to it at all. A game whose values were not numbers would need the error measured against a stop interval instead, which is what the domineering ladder does and what makes its failures harder to price.
Three strips are enough, and a thousand would not be
The refutation here rests on three four-square strips, and it is worth saying why that is a complete argument rather than a small sample — because the instinct in a subject full of censuses is to want more of them.
A statistical criterion is a function of a stated list of quantities: length, coin count, run count, colour counts. Two positions that agree on every quantity in the list are, to any such criterion, the same input. So a criterion assigns them the same answer, necessarily, whatever the criterion is.
.LLR, .LRL and .RLL agree on every quantity in the list and their readings are wrong by , and . One input, three required outputs. No function does that, so no criterion of that shape exists, and the argument is finished — a fourth example would add nothing and a thousand would add nothing.
That is a different kind of evidence from everything else on this ladder, and the difference is worth recognising. A census supports a claim by failing to find an exception and is always provisional; a counterexample to a functional dependence is a proof, and it is available from three positions.
It also says exactly how to widen the list if somebody wants to. The refutation kills every criterion over these quantities, and adding a quantity that separates the three strips — anything reading the order of the colours — escapes it immediately. So the finding is not that Push is unreadable but that the reading must be order-sensitive, which is precisely what the rung above delivers.
The 207 statistical classes carrying more than one strip are then the size of the problem rather than the proof of it: they say the collision is common, not that it happens.
What a player should do with this
Two things, and they are both negative, which is the honest content of the page.
Do not carry the reading past four squares. It is exact to three, wrong on a fifth of the four-square strips, and wrong on a majority of the seven-square ones — and the errors are not small enough to play against. A player who wants a rule of thumb for Push does not have one, and the nearest thing available is the guarantee from the rung below: if the colours cannot be brought together by any line of play, the count is right and it is right exactly.
And do not look for the missing correction term where the rung below suggested. How much room the interaction costs each side is a well-defined quantity and it does not determine the answer, because three strips agreeing on every such quantity disagree on the value. The repair, if there is one, is a different kind of object: an expansion in which position matters, of the sort a two-coloured row already spells out in a game where nothing has to move out of the way.
That is a smaller conclusion than the rung below expected and it is a firmer one. A criterion that does not exist is settled; a criterion nobody has found yet is not.
The convention, named
Normal play: the player who cannot move loses. Both players push left. Left moves an L coin, Right an R coin, and a push moves the coin together with the contiguous run of coins immediately in front of it, one square, into an empty square that must exist.
The reading is the count a player writes down: for each coin from the wall outward, the empty squares in front of it that no earlier coin has claimed, signed positive for Left and negative for Right.
The error is the value minus the reading, and the census reports its magnitude, because the two colours are mirror images and the signed errors come in cancelling pairs — a table of signed errors would show a symmetry rather than a size.
The room is the same count taken unsigned: every empty square in front of every coin, added up regardless of colour.
Where the ladder goes next
The push anchor has three rungs to here, and this one has just ruled out a whole shape of answer. The three above supply the shape that works, find where it stops, and locate the information.
A numeral in the empty squares is the reading that survives the refutation, and its two ingredients are the reverse of the guess. The colours pick a fraction — , , , — and the empty squares supply the binary precision, so a run of coins before one of the other colour, with gaps, is worth exactly . That is a statement about order, which is what this page shows any answer has to be.
The cliff a cut invents then finds why it does not compose. A strip cut at a gap is not two positions: a Push strip is a line with a wall at one end, and cutting hands the back half a wall it never had. Widen the gap and the value converges geometrically to a limit that is not the sum — and Shove, whose reading is exact everywhere and adds across genuine sums, fails at the same cut on 89 of 93 strips.
Read from the back forwards closes the anchor by locating the information. The convergence rate belongs to the rearmost run, at every gap and every distance, and it does not compound across runs — while Shove, one clause away, compounds.
So the ladder ends where this page’s three four-square strips point: the value of a Push strip is decided by the order of its colours against a wall, and every attempt to read it as a statistic, a sum, or a per-run quantity fails for the same reason.
Part 3 of 8
One argument about Push. The parts either side of it:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
The objects named here
The third axis, after the field and the series: the games, values and theorems themselves, and every essay that touches each one.
ApproximationBinaryBoundCounterexampleEnumerationHackenbushHeuristicInvariantNumberPartizanPushValue
- Half the difference in odd runs approximation, bound, counterexample, enumeration, heuristic, invariant, number, value
- One domino every three cells approximation, bound, enumeration, heuristic, invariant, number, partizan, value
- The moves a player can be talked out of approximation, bound, counterexample, enumeration, heuristic, number, partizan, value
- Two strips that end the same way approximation, counterexample, enumeration, invariant, number, partizan, push, value
- Cut small unless you are behind counterexample, enumeration, heuristic, invariant, number, partizan, value
- The birthday is a floor bound, counterexample, enumeration, hackenbush, invariant, number, value