A result at five per cent that favours the null
Worth reading first: Evidence measured in decibans · Certain of a coin that is not there.
Evidence measured in decibans wrote Bayes’ theorem as addition, with every observation contributing the logarithm of a likelihood ratio, and certain of a coin that is not there showed what that arithmetic does when the list of hypotheses is wrong. Both compared a handful of definite hypotheses: a coin with bias 0.3, 0.6 or 0.8. The question most experiments actually put is different. It is whether a coin is exactly fair, against the vague alternative that it is biased by some amount nobody can specify in advance.
There are two standard ways to answer it, and on large samples they disagree flatly. A significance test computes the p-value: the chance, if the coin is fair, of a count of heads at least as far from half as the one observed. A small p-value, conventionally below 0.05, is read as evidence against fairness. A Bayes factor computes how much more probable the observed count is under one hypothesis than under the other, with the biased coin’s probability averaged over the biases it might have. Hold the p-value fixed at 0.05 and let the sample grow, and the Bayes factor moves steadily towards the fair coin, without limit. A result that a significance test calls significant can be strong evidence for the very hypothesis the test throws out.
This is Lindley’s paradox, named for Dennis Lindley’s paper of 1957, although Harold Jeffreys had noted the effect in 1939. It is not a paradox in the sense of a contradiction. Both computations are correct; they answer different questions, and the figures here locate exactly where the answers part company and why.
The same p-value at every sample size
The figure fixes the p-value and varies the sample. For a fair coin tossed times, the count of heads has mean and standard deviation , and by the bell curve assembled out of coin flips a count standard deviations above the mean has a two-sided p-value read from the normal tail. A p-value of 0.05 is . So the figure places the count of heads at for every — always exactly at the edge of significance — and computes the Bayes factor there.
For a coin whose bias is spread evenly over , the Bayes factor in favour of fairness has a closed form. The probability of heads under a fair coin is , and under the biased coin, averaged over , it is — every count from to equally likely, which is the uniform prior’s signature. The ratio is
and at a count standard deviations from the mean, Stirling’s formula turns it into . The exponential factor depends only on , so only on the p-value; the square root grows with without bound.
At the factor is , and the two hypotheses are at even odds near 76 tosses. Below that, a significant result favours the biased coin, modestly. Above it, a significant result favours the fair coin, and increasingly: at a million tosses, by 117 to one. A smaller p-value only postpones the crossing. At the crossing is near 1,200 tosses and at near 83,000, and every line in the figure eventually rises through even odds and keeps going.
A hundred million trials and a significant result
The circled point is a real data set. From 1979 the Princeton Engineering Anomalies Research laboratory asked volunteers to influence the output of a random-event source by intention alone, and over twelve years it recorded 104,490,000 trials, of which 52,263,471 came out in the intended direction — 18,471 more than half. That is , a two-sided p-value of , and by any conventional standard a highly significant departure from chance. William Jefferys computed the Bayes factor in 1990 and found it favoured the hypothesis of no effect by about twelve to one, which the figure reproduces: 11.9.
Neither number is in error. The p-value says, correctly, that a fair device would rarely produce a surplus as large as 18,471. The Bayes factor says, correctly, that a device with some bias spread over the whole range would produce such a small surplus even more rarely. A bias that makes a surplus of likely is a bias within about of one half, and a hypothesis that spreads its belief evenly from 0 to 1 gives that narrow window a tiny weight. The data are unusual under both hypotheses, and less unusual under the fair one.
The same peak over a narrower strip
The next figure shows the mechanism directly.
The curve is the likelihood ratio: for each possible bias , how much more probable the observed count is under that bias than under a fair coin. Its peak is at the observed proportion of heads and its height there is , which for is — the same at every sample size, because a fixed p-value is a fixed number of standard errors and the height depends on nothing else. What changes is the width. The likelihood is concentrated within a standard error of the peak, and the standard error is , so a hundredfold larger sample makes the peak ten times narrower.
The Bayes factor for the biased coin is the average of this curve over the prior — here over evenly, which is the area under the curve. The same peak over a strip a tenth as wide covers a tenth of the area, and that is the whole of Lindley’s paradox in one picture. The p-value measures the height of the peak in standard-error units. The Bayes factor measures how much of the alternative’s belief sits under it, and a vague alternative’s belief sits almost entirely away from it once the sample is large enough to locate the bias precisely.
The most a p-value can say
It is natural to blame the uniform prior. A biased coin that might equally have any bias from 0 to 1 is not what anyone testing a random-event source believes; a plausible alternative puts most of its weight near one half. So the question becomes how favourable to the alternative the prior can be made, and the answer has a ceiling.
The most favourable alternative of all puts its entire weight on the bias that best fits the data — chosen after looking at them. Its Bayes factor is the peak of the likelihood ratio, , which at is . No prior can do better, because an average cannot exceed its maximum. Ward Edwards, Harold Lindman and Leonard Savage pointed this out in 1963: a result at the 5% level can never, under any alternative hypothesis whatever, be more than about seven to one against the null, and that bound is reached only by an alternative chosen to fit the result exactly. The phrase “one chance in twenty” invites a reading of twenty to one, and no honest calculation supports it.
Thomas Sellke, Maria Bayarri and James Berger gave in 2001 a bound that applies to the alternatives a reasonable analyst might actually hold, those whose p-values would be spread with a decreasing density: the odds against the null are at most . At that is 2.5 to one. Their reading is that a result at 5% is weak evidence even before the sample size enters, and that the size of the sample, through Lindley’s effect, can only weaken it further.
The bar rises with the sample
Turned around, the computation says what a result must be to count as evidence.
Setting gives . The threshold rises with the sample, slowly — the square root of a logarithm — but without limit: 2.04 at a hundred tosses, 2.96 at ten thousand, 3.66 at a million, 4.24 at a hundred million. A test at a fixed significance level holds its threshold at 1.96 regardless. Between the dashed line and the curve lies a region of results that are significant at 5% and favour the null, and the region widens as the sample grows.
This is the precise sense in which a fixed significance level is too lenient for a large sample. Jeffreys’s own approximate tests of 1939 behave exactly this way, with a threshold growing like ; the Schwarz criterion for choosing between statistical models, which penalises each extra parameter by , is the same rule written for model selection. The proposal by Daniel Benjamin and seventy-one co-authors in 2018, to move the default threshold for a new discovery from 0.05 to 0.005, is a coarser version of the same correction, fixed at one sample size rather than growing with it.
Tuning the alternative buys at most two to one
Between the uniform prior and the post-hoc choice lies a family of reasonable alternatives: biases spread around one half with some chosen width.
A very narrow spread makes the alternative indistinguishable from the null, and the odds sit at one. A very wide spread is Lindley’s case, and the odds for the null grow in proportion to the width. In between, each curve has a minimum, and the minimum is the same at every sample size: about , or 2.1 to one in favour of a bias. The algebra for a bell-shaped alternative gives it exactly, at . What changes with the sample is where the best spread lies — 0.083 at a hundred tosses, 0.0009 at a million — which is to say that the alternative most favourable to a significant result is one that expected an effect of just the size that happened to be seen.
This is the reply to the objection that Lindley’s paradox is an artefact of a silly prior. No bell-shaped prior centred on fairness makes a 5% result worth more than about two to one, and any prior fixed before the data — not tuned to a sample size nobody knew in advance — is worse than that at all but one sample size, and arbitrarily worse at large ones. Maurice Bartlett observed in 1957, in reply to Lindley, that the Bayes factor depends on the prior’s width without limit. That is true, and it cuts both ways: the evidence a significant result carries is not a property of the result alone, and a test that reports it as if it were is answering a question nobody asked.
Harmless for estimating, decisive for testing
The prior’s width does something in a test that it never does in estimation, and the contrast explains why the paradox surprises people who use Bayesian methods routinely. Suppose the question is not whether the coin is fair but how biased it is. Then the posterior for the bias after a million tosses is a narrow bell around the observed proportion, of width one standard error, and it is nearly the same whatever smooth prior was used, because the likelihood is a thousand times narrower than any reasonable prior and swamps it. The average settles and the wobble does not described that narrowing: the wobble of an average shrinks like . For estimation, a vague prior is a safe default.
For testing, the prior enters not through its shape near the peak but through its height there — through how much belief it placed in the narrow window the data eventually picked out — and that height is inversely proportional to the prior’s width. Double the width and the Bayes factor for the null doubles. A coin’s bias cannot be spread wider than , but the mean of a measurement can, and in the limit of a perfectly flat prior over all real numbers the null wins every time, whatever the data. The improper prior that is harmless for estimation is fatal for testing, a cousin of the trouble that the envelope that always looks better traced to a prior that spreads its weight too thinly to be normalised. A test needs an alternative stated sharply enough to be wrong.
The tail probability does not escape the issue either. How far from the average a thing can be bounded the chance of a result standard deviations out using nothing but the mean and the spread, and the bound — one in — holds for every distribution. A p-value is such a tail probability, computed under one hypothesis. It is a statement about where the result sits in one distribution, and no statement about one distribution compares two.
Which question each number answers
The p-value answers a question about the null hypothesis alone: if the coin is fair, how surprising is this? It needs no alternative and no prior, which is its great practical virtue, and it guarantees that a fair coin is rejected at most 5% of the time by a test at the 5% level. It does not say how much more surprising the result would be under some other hypothesis, and so it cannot say which hypothesis the data favour.
The Bayes factor answers the comparative question, and it pays for the answer by needing the alternative stated precisely enough to compute with. Lindley’s paradox is the observation that the two questions have answers that diverge without limit as the sample grows. A tiny surplus of heads in a vast sample is surprising if the coin is fair, and far more surprising under almost any definite account of how it might be biased. Bayes’ theorem as a picture of a square showed the same structure in miniature: a positive result on an accurate test can still leave the disease unlikely, because the result’s probability under the alternative matters as much as its improbability under the null.
The paradox also explains a familiar pattern in practice. Very large studies — genetic association scans, randomised trials with hundreds of thousands of participants, online experiments on millions of users — routinely report significant effects that are tiny. Some are real and tiny. But a significant result in a very large sample is also exactly the case in which a fixed threshold is most lenient, and a Bayes factor computed against any plausible alternative may well favour no effect at all.
What the figures cannot show
Every figure here uses the normal approximation to locate the count at a given z-score, and the exact binomial computation for the Bayes factor, so the curves are accurate to within the approximation’s error, which is negligible beyond a few dozen tosses. The PEAR computation follows Jefferys’s choice of a uniform alternative; a narrower alternative gives smaller odds for the null, and the tuning figure shows how much smaller — about 2.1 to one for a bias at the very best. None of this shows whether the source was biased, which is a question about the experiment and not about the arithmetic.
The figures also hold the p-value fixed while the sample grows, which is the idealised setting of the paradox. A real biased coin produces a z-score that grows like , and both methods eventually detect it; the paradox concerns results that stay near the edge of significance, which under a real effect become rare as the sample grows and under the null do not.
Still open: what a significance threshold should be
The mathematics of the disagreement is settled. What is not settled is what to do about it. One school holds that the p-value should be retired as a measure of evidence and replaced by Bayes factors or posterior probabilities, with alternatives stated in advance; another that p-values remain the right tool for controlling error rates and that the fault lies in reading them as evidence. Proposals between the two include thresholds that shrink with the sample, as Jeffreys’s tests imply; reporting the bound beside every p-value; and e-values, quantities that behave like Bayes factors under the null and can be computed without committing to a single prior. Which of these, if any, should become standard is argued in the statistical literature and in the editorial policies of journals, and there is no consensus. The companion problem, of what happens when the sample size is not fixed in advance but chosen by watching the data, is the subject of peeking until the answer is yes, and it is the one setting where the two methods’ difference is not a matter of interpretation at all.
Surprising under one, more surprising under the other
A count of heads exactly at the edge of significance is surprising if the coin is fair, and the surprise is the same in standard-error units however many times the coin was tossed. Under a coin whose bias was uncertain before the tossing began, the same count becomes more and more surprising as the sample grows, because a larger sample pins the bias into a narrower window and the vague alternative had little belief to spare there. So the same p-value is evidence against fairness at fifty tosses, neutral near seventy-six, and evidence for fairness at a million, and the twelve-to-one verdict on a hundred million trials at is the same arithmetic at scale. The p-value is computed correctly throughout; what it cannot do is compare.
What links here
Computed from the collection, not written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- One coin, counted by runs and by wakings — both name bayes' theorem, likelihood
- The door that was not opened — both name bayes' theorem, likelihood
- Two children and the sentence about one of them — both name bayes' theorem, likelihood
Named objects
A dashed tag is an object no other essay names yet.
Bayes factorBayes' theoremCentral limit theoremLikelihoodModel misspecificationP valuePosterior