Peeking until the answer is yes
Worth reading first: A result at five per cent that favours the null · Two barriers and a fair game.
A result at five per cent that favours the null compared a p-value and a Bayes factor computed on the same data with the sample size fixed in advance, and found them disagreeing more and more as the sample grew. That disagreement is about interpretation: both numbers were computed correctly and they answer different questions. This essay changes one thing — the sample size is no longer fixed, but chosen by watching the data — and the disagreement stops being a matter of interpretation. One method’s guarantee collapses entirely, and the other’s survives intact.
The practice is common and usually innocent-looking. An experimenter runs a study, checks the result, finds it not quite significant, collects a few more observations, and checks again. An online experiment is monitored on a dashboard that updates every hour. A clinical trial has interim analyses. Each time the data are examined, there is a chance of stopping on a fluctuation, and the question is how much that chance accumulates.
Watching a fair coin until it looks unfair
The figure runs the experiment two thousand times. Each coin is fair. After every toss from the tenth, two numbers are computed: the z-score of the count of heads — significant at 5% when it is beyond — and the Bayes factor comparing a coin of unknown bias, spread evenly over , with a fair coin. The first curve records, at each point, how many of the coins have already crossed the 5% line at least once; the second, how many have already shown a Bayes factor of twenty to one against fairness.
At any single, pre-chosen sample size, a fair coin is significant at 5% exactly 5% of the time. That is what the threshold means. But a coin examined after every toss gets a fresh 5% chance at each look, and although successive looks are strongly correlated, the chances accumulate. By a thousand tosses, 47% of the fair coins have been declared biased at some moment. By a hundred thousand, 70% have. Peter Armitage, Charles McPherson and Brian Rowe tabulated this in 1969 for repeated tests on accumulating data, and the effect had been pointed out decades earlier, by William Feller among others, against experiments in extrasensory perception whose authors stopped when the results looked good.
The Bayes factor’s curve is different in kind. It rises at first, as a few coins wander far enough early on to produce twenty to one, and then it flattens and stays below one in twenty — 3.4% after a hundred thousand tosses — however long the watching continues.
One coin, significant twenty times over
A single coin shows how both things happen at once.
The coin’s z-score wanders. It is a random walk rescaled by the square root of the number of tosses, and in the limit it behaves like a Brownian motion observed on a logarithmic clock, which the walk that becomes a curve introduced. It leaves the band twenty times in a hundred thousand tosses. Any of those exits, reported as the result of the study, would be a significant finding of bias in a fair coin.
The Bayes factor computed from exactly the same tosses never reaches even two to one. When the z-score leaves the band at 2.0 or 2.1 after several thousand tosses, the Bayes factor is, by the arithmetic of the paradox, well below one: a result at the edge of significance in a large sample favours the fair coin. And as the tosses accumulate, the Bayes factor drifts towards zero, because a fair coin is what the data keep confirming. The two numbers are computed from the same information and point in opposite directions throughout.
Why the p-value fails with certainty
The figure’s 70% is not the end. With no limit on the tosses, every fair coin is eventually declared significantly biased, with probability one.
The reason is the law of the iterated logarithm, proved by Aleksandr Khinchin in 1924. A fair coin’s surplus of heads, divided by , is a z-score that by the central limit theorem has a bell-shaped distribution at each fixed . Khinchin’s law describes the whole path: the z-score keeps returning to values as large as , and no larger, for ever. The average settles and the wobble does not named the law as the sharper statement beyond the central limit theorem, and this is a place where it is needed. Since grows without bound — slowly, passing 1.96 near a thousand tosses — the z-score must cross any fixed threshold infinitely often. A rule that says “stop when ” stops with probability one.
So a significance test applied with optional stopping does not merely inflate its error rate from 5% to something larger. It can be made to reject a true null hypothesis with certainty, by an experimenter with enough patience — a practice that came to be called sampling to a foregone conclusion.
Repairs that keep the p-value
Statisticians who run clinical trials know all this, and the standard repair keeps the significance test by planning the looks. In a group sequential design the number of interim analyses is fixed in advance, and the threshold at each is raised so that the chance of ever crossing, under the null, is 5% in total. Stuart Pocock’s design of 1977 uses one raised threshold for every look: with five looks it is each time, rather than 1.96. Peter O’Brien and Thomas Fleming’s design of 1979 makes the early thresholds very strict and relaxes them later — with five looks, 4.56 at the first and 2.04 at the last — so that a trial stopped early has overwhelming evidence and the final analysis is nearly the fixed-sample one. Gordon Lan and David DeMets showed in 1983 how to spend the 5% across looks whose number and timing need not be fixed, as long as the spending schedule is.
These designs work, and they are the reason interim analyses in trials do not produce the figure’s 70%. But they work by making the plan part of the analysis. The same data, with the same final z-score, are significant or not according to how many looks were planned, and a look taken outside the plan — a dashboard checked one extra time — breaks the guarantee. The error rate is a property of the procedure, and the procedure includes what the analyst would have done in situations that never arose.
The evidence does not depend on the plan
A Bayes factor does not care about the plan, and the cleanest illustration is a coin that comes up heads nine times and tails three times. If the experimenter had decided to toss twelve times, the one-sided p-value for an excess of heads is the chance of nine or more heads in twelve, , not significant. If the experimenter had instead decided to toss until the third tail, the same sequence gives the chance of nine or more heads before the third tail, , significant. Dennis Lindley and Lawrence Phillips used this example in 1976. The tosses are identical; only the intention differs, and the p-value changes because it sums over sequences that did not happen, and which sequences could have happened depends on the intention.
The Bayes factor depends on the tosses only through the probability of the sequence actually observed under each hypothesis, , and that is the same whichever rule was used, since the rule’s own factor — the number of ways the sequence could have been arranged — cancels in the ratio. This is the likelihood principle, and Ward Edwards, Harold Lindman and Leonard Savage drew the consequence for stopping in 1963: for a Bayesian analysis, the stopping rule is irrelevant. Ville’s inequality is the frequency guarantee that makes that irrelevance safe — the reason a rule that ignores the plan cannot be exploited by choosing the plan.
Ville’s bound on watching a Bayes factor
The Bayes factor’s protection is a theorem, and it is quantitative.
Under a fair coin, the Bayes factor for the biased coin is a martingale: whatever has happened so far, its expected value after the next toss equals its value now. The reason is short. The Bayes factor after tosses is the ratio of the probability of the tosses under the biased coin to their probability under the fair one, and the expected ratio after one more toss, computed under the fair coin, is the sum over heads and tails of the fair probability times the ratio — which cancels to the current ratio. It starts at one, before any tosses, and it can never go negative.
Jean Ville proved in 1939 that a non-negative martingale starting at one exceeds at any time, ever, with probability at most . It is the inequality bounding how far from the average a thing can be, in Markov’s form, extended from one moment to a whole path: a fair game that starts at one pound cannot give a player better than a one-in- chance of ever holding pounds, whatever rule the player uses for when to stop. Two barriers and a fair game used the same fairness to find the chance of ruin, and Ville’s inequality is that argument with the lower barrier removed.
So the probability that a fair coin, watched for ever, is ever reported as twenty-to-one biased by its Bayes factor is at most one in twenty. The measured shares in the figure are all below the line, and well below at small , because the Bayes factor rarely jumps exactly to and the finite run gives fewer chances than an infinite one. Richard Royall called the quantity the probability of misleading evidence, and its bound holds whatever stopping rule is used — which is what the p-value lacks.
A fair game whose typical value sinks
How can a quantity with average one be so rarely large? The exact distribution answers it.
The exact mean is one at every sample size: summing the Bayes factor against the fair probability over all gives . But the median falls steadily, from 0.44 at ten tosses to 0.005 at a hundred thousand, and the chance of exceeding one falls to 0.1%. The mean is held at one by rare enormous values — the extreme counts of heads, which a fair coin almost never produces and which give the biased coin huge support. A typical fair coin’s Bayes factor sinks towards zero, which is the Bayes factor learning the truth, and the rare excursions upward are paid for, in Ville’s bookkeeping, by all the mass that sank.
The p-value has no such structure. At every sample size it is spread almost evenly over under the null — exactly evenly but for the coin’s discreteness, and as likely to be below 0.05 at a million tosses as at ten — so it does not sink as evidence accumulates, and repeated looks keep drawing fresh chances from the same distribution.
A boundary that outruns the wandering
Translated into the z-score, the Bayes factor’s stopping rule is a threshold that rises with the sample.
From the formula in the companion essay, a Bayes factor of against the fair coin needs , a boundary growing like . The fair coin’s excursions grow like , which is smaller by far: outgrows , so beyond some point the boundary is permanently out of reach. The fixed line at 1.96 is crossed by the iterated logarithm near a thousand tosses and then lies below it for ever. Every rule of the form “stop when ” fails under optional stopping, and every rule whose threshold grows faster than has a chance below one of ever stopping a fair coin. Herbert Robbins and Donald Darling built confidence sequences on exactly this in 1967: intervals valid at every sample size at once, whose widths shrink a little more slowly than the fixed-sample interval to pay for being watched.
What the guarantee costs a real effect
A boundary that a fair coin rarely reaches might be one that a biased coin takes too long to reach. The last figure measures it.
The cost is modest. A coin with bias 0.55 is stopped at twenty to one after a median of 979 tosses; a fixed-sample test with 80% power at the 5% level needs 778, and offers a far weaker guarantee — rejection at one in twenty under the null, rather than a Bayes factor of twenty. The reason is the same arithmetic as before, read the other way: under a real bias, the log of the Bayes factor grows in proportion to the number of tosses, at a rate set by the divergence between the true coin and the fair one, while the boundary rises only like a logarithm. Every biased coin in the simulation was eventually stopped. Evidence measured in decibans met the same economy in Wald’s sequential test for two definite hypotheses; the Bayes factor against a spread-out alternative extends it to the case Wald’s test did not cover, where the size of the bias is unknown.
What the figures cannot show
The simulations use 2,000 coins per curve, so the shares carry sampling error of about one percentage point at 50% and a few tenths of a point at 3%, and the run stops at a hundred thousand tosses where the theorems speak of infinite runs. The p-value’s approach to certainty is a theorem and is only begun in the figure. Ville’s bound is a theorem too, and the figure shows it holding with room to spare; it does not show that the bound is nearly attained for continuous-time processes, which is also true.
Nor do the figures say that a Bayes factor is immune to every abuse. Its guarantee is against stopping rules, not against choosing the alternative after seeing the data, which can still inflate it to the ceiling of that the companion essay computed. And a Bayes factor protected against optional stopping is still only as good as the model behind it.
Still open: which guarantee to ask for
The mathematics here is classical, and its consequences for practice are being worked out now. E-values — non-negative quantities whose expected value under the null is at most one, of which a Bayes factor is the leading example — have become the basis of anytime-valid testing, in work by Glenn Shafer and Vladimir Vovk on game-theoretic probability and by Peter Grünwald and others on safe testing. They keep Ville’s guarantee for any stopping rule and can be multiplied across studies, which p-values cannot. What is not settled is how to choose them well: an e-value with good power against one family of alternatives can be poor against another, and the optimal choice when the alternative is composite and the stopping rule unknown is an active question. Whether the practice of science should move from p-values to e-values, and how a field would agree on the alternatives they need, is not a mathematical question at all.
A fair game cannot be watched into a win
A p-value is a fresh chance at every look: spread evenly under the null at every sample size, so repeated looks keep drawing from the same distribution, and the law of the iterated logarithm guarantees that a fair coin watched long enough crosses any fixed line. A Bayes factor under the null is a fair game: its average is one at every sample size, its typical value sinks as the evidence accumulates, and Ville’s inequality bounds the chance of ever seeing to one by , whatever rule decides when to stop. Of two thousand fair coins watched for a hundred thousand tosses, 70% were significant at some moment and 3.4% ever reached twenty to one. The difference is not about interpretation. It is whether the number that measures evidence is a fair game, and only one of them is.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A long enough chain never locks — both name central limit theorem, random walk
- A walk that follows its own footsteps — both name martingale, random walk
- An urn forgets its start only below one half — both name central limit theorem, random walk
Named objects
A dashed tag is an object no other essay names yet.
Bayes factorCentral limit theoremLikelihoodMartingaleP valueRandom walkStopping time