Probability

Certain of a coin that is not there

Bayes' theorem is exact, and it is only as good as the list of hypotheses it is given. Hand it a list that leaves out the truth and it does not hesitate: it becomes certain of the entry that is least wrong, in a precise sense — the one closest in Kullback–Leibler divergence — and if two entries are equally wrong it never settles at all. Hand it a model that assumes independence where there is none, and its intervals shrink as fast as they would for honest data while covering the truth less and less often.

Worth reading first: Evidence measured in decibans · The envelope that always looks better.

Evidence measured in decibans wrote Bayes’ theorem as addition: prior odds in decibans, plus the weight of each observation, gives posterior odds, and the weights are logarithms of likelihood ratios. It ended with a question the arithmetic cannot answer by itself. Every weight is computed from a model — a statement of how likely each observation would be under each hypothesis — and when the model is wrong, the weights are wrong. How wrong can the conclusion then be?

The question has a precise answer in two common situations, and neither is reassuring. The first is a list of hypotheses that does not contain the truth. Bayes’ theorem still produces a posterior, and the posterior still becomes certain; it becomes certain of the wrong thing, and which wrong thing is decided by a single number. The second is a model that assumes observations are independent when they are not. The posterior still concentrates at the usual rate, its intervals still shrink as the data accumulate, and they cover the truth less and less often.

Neither is a failure of the theorem. Bayes’ theorem is a rule for combining a prior with a likelihood, and it combines them exactly. What fails is the tacit claim that the likelihood describes the world. This essay computes what happens when it does not, on the simplest object there is: a coin.

A fair coin and three wrong guesses

A fair coin judged by someone who thinks it is biased. Posterior probabilities of three wrong coin biases over two thousand tosses of a fair coin, for five runs; the bias 0.6 wins every run.
Fig. 1 A fair coin tossed 2,000 times, judged by someone who believes its bias is 0.3, 0.6 or 0.8 with equal odds — none of them right — and the posterior probability of each after every toss, for five runs, on a logarithmic scale of tosses. Every run settles on 0.6 and ends more than 99% sure of it.

Someone is sure that a coin is biased and knows only that its chance of heads is 0.3, 0.6 or 0.8. They start with equal odds on the three and update by Bayes’ theorem after every toss. The coin is in fact fair.

For the first few tosses the posterior lurches about, as it should: three heads in a row favour 0.8, a tail favours 0.3. After a few dozen tosses one hypothesis pulls ahead, and after a few hundred it has won. In all five runs drawn, and in essentially every run there could be, the winner is 0.6. By two thousand tosses the believer is more than 99% certain that the coin’s bias is 0.6, a number the coin does not have, and further tosses only make them more certain.

Nothing has gone wrong in the calculation. Each update multiplied the odds by the ratio of likelihoods, exactly as it should, and the posterior is the correct posterior for someone who believes what this person believes. The failure is that the belief put zero prior weight on the truth, and Bayes’ theorem cannot move weight onto a hypothesis that has none: a prior of nought times any likelihood is still nought. What it can do is choose among the hypotheses it was given, and it chooses decisively.

The number that picks the winner

How far each wrong coin is from the true one. A curve of the Kullback–Leibler divergence from a fair coin to a coin of bias q, with three hypotheses marked and the closest one highlighted.
Fig. 2 The Kullback–Leibler divergence from a fair coin to a coin of bias qq, with the three hypotheses marked: 0.0872 for 0.3, 0.0204 for 0.6 and 0.2231 for 0.8. The posterior goes to whichever listed hypothesis sits lowest on the curve.

Why 0.6? Each toss adds to the log odds between two hypotheses the logarithm of the ratio of their likelihoods, and when the tosses come from the true coin, the average of that logarithm has a name. For a true bias pp and a hypothesis qq it is the difference of two quantities of the form

D(p ∥ q)=plog⁡pq+(1−p)log⁡1−p1−q,D(p \,\|\, q) = p \log\frac{p}{q} + (1 - p)\log\frac{1 - p}{1 - q},

the Kullback–Leibler divergence from the truth to the hypothesis. It is the average amount of evidence per toss by which the truth would beat the hypothesis if both were on the list, and it is never negative, and nought only when q=pq = p. Between two wrong hypotheses q1q_1 and q2q_2 the average weight of a toss, in favour of q1q_1, is D(p ∥ q2)−D(p ∥ q1)D(p\,\|\,q_2) - D(p\,\|\,q_1): the evidence favours whichever is closer to the truth in this particular sense.

For a fair coin the divergence to 0.6 is 0.0204 nats a toss, to 0.3 it is 0.0872, and to 0.8 it is 0.2231. So 0.6 is the closest of the three, and the posterior goes there. This is a theorem, due in this generality to Robert Berk in 1966: when the truth is not among the hypotheses, the posterior concentrates on the hypotheses that minimise the Kullback–Leibler divergence from the truth. The divergence is the same quantity that the whole histogram deviating found governing the probability of a whole histogram’s deviation, and it plays the same role here: it measures how quickly the data can tell two distributions apart.

The divergence is not a distance in the ordinary sense. It rises steeply towards biases of nought and one, so a hypothesis of 0.9 is more than twice as far from a fair coin as one of 0.8, and it is not symmetric in its two arguments. The ranking it gives is the ranking the evidence enforces, and it may not be the ranking the believer would have guessed.

Evidence as a walk with a slope

Evidence between two wrong hypotheses, toss by toss. Six random walks of log posterior odds between two wrong coin biases, all climbing along a straight line whose slope is the difference of two divergences.
Fig. 3 The log odds, in decibans, of bias 0.6 against bias 0.3 after each toss of a fair coin, for six runs of 3,000 tosses, with a straight line at 0.290 decibans a toss: the difference of the two divergences.

The same picture can be seen as a walk. Between two hypotheses, each toss adds one of two fixed amounts to the log odds — here log⁡(0.6/0.3)\log(0.6/0.3) for a head and log⁡(0.4/0.7)\log(0.4/0.7) for a tail — and which one it adds is decided by the coin. The log odds therefore perform a random walk with steps of two sizes, and the walk’s average step is the difference of divergences: 0.290 decibans a toss in favour of 0.6.

A walk with a positive drift goes to infinity. Its fluctuations grow like the square root of the number of steps and its drift grows like the number itself, so after enough tosses the drift wins and the log odds are as large as anyone cares to name. That is the posterior becoming certain. It is also why certainty arrives on a schedule set by the divergences: a small difference between them means a long walk before the drift dominates, and two nearly equally wrong hypotheses can take thousands of observations to separate, although the final verdict is never in doubt.

The envelope that always looks better found expectation misleading when it was infinite. Here the expectations are finite and perfectly well behaved, and they mislead for a different reason: they are expectations under a model that is not the world, and the walk they describe climbs towards a hypothesis whose only merit is that it is less wrong than the alternatives.

When two hypotheses are equally wrong

Two equally wrong hypotheses: a posterior that never settles. Four runs of the posterior probability of one of two equally wrong coin biases over twenty thousand tosses, swinging between near-certainties indefinitely.
Fig. 4 A fair coin tossed 20,000 times, judged by someone who believes its bias is 0.3 or 0.7 — equally wrong in both directions — with the log odds of 0.3 against 0.7, in decibans, after every twentieth toss for three runs. The odds wander to strong conviction for one bias and then for the other; the runs changed their minds 68, 12 and 26 times.

The theorem has an edge case, and Berk found it too. If two hypotheses are exactly equally far from the truth — a fair coin judged by someone choosing between 0.3 and 0.7 — the drift is nought, and the log odds perform a walk with no drift at all. Such a walk, as the walk that comes home showed, returns to every level infinitely often and also leaves every bounded region infinitely often.

So the posterior never settles. It swings to strong conviction that the coin favours tails, then drifts back through indifference to strong conviction that it favours heads, and back again, for as long as the tosses continue. The three runs drawn changed their minds 68, 12 and 26 times in twenty thousand tosses, and each of them spent long stretches more than a hundred decibans — odds of ten billion to one — in favour of one of two false hypotheses.

This is not a curiosity of symmetric examples. Any time the true distribution sits at the same divergence from two or more hypotheses, the posterior has no limit. In richer models, where the hypotheses form a continuum, the same thing appears as a posterior that concentrates on a set rather than a point and wanders over it.

A model that ignores dependence

The second failure is subtler, because the truth is in the family. Take a coin that is fair in the long run but sticky: after each toss it repeats its last result with probability (1+ρ)/2(1 + \rho)/2 and switches with probability (1−ρ)/2(1 - \rho)/2. Over many tosses heads and tails each come up half the time, and the true long-run bias is exactly one half. The dependence parameter ρ\rho is the correlation between successive tosses.

Now analyse the tosses as if they were independent, with a uniform prior on the bias. The model family contains the right bias, one half, and the posterior concentrates on it: the posterior mean converges to one half as the tosses accumulate. What goes wrong is the width.

Credible intervals from a model that ignores stickiness. A plot of how often a 95% credible interval covers the truth when dependent coin tosses are treated as independent, falling from 95% towards a third as the dependence grows.
Fig. 5 A sticky coin tossed 1,000 times and analysed as if the tosses were independent: the share of 1,500 runs whose 95% credible interval contains the true bias ½, against the stickiness ρ\rho (dots), with the curve 2Φ(1.96(1−ρ)/(1+ρ))−12\Phi\big(1.96\sqrt{(1-\rho)/(1+\rho)}\big) - 1. At ρ=0\rho = 0 the coverage is 95%; at 0.4 it is 83%; at 0.8 it is 45%.

The independent model thinks a thousand tosses carry a thousand tosses’ worth of information, and its 95% interval has the width appropriate to that. A sticky coin’s thousand tosses carry much less: they come in runs, and a run of ten identical results is closer to one observation than to ten. The true spread of the proportion of heads is larger than the model believes by a factor of (1+ρ)/(1−ρ)\sqrt{(1+\rho)/(1-\rho)}, and the interval, being too narrow by that factor, misses the truth more often.

The figure measures it. At ρ=0.4\rho = 0.4 the nominal 95% interval covers 83% of the time; at ρ=0.8\rho = 0.8, 45%. And this does not improve with more data. The coverage depends on the ratio of the true spread to the believed one, which does not change as the number of tosses grows, so at a million tosses the interval is a thousand times narrower and still covers the truth less than half the time. The posterior is converging to the right answer and is systematically overconfident about how close it is.

Counting each toss as part of a toss

Counting each sticky toss as a fraction of a toss. A plot of coverage against the tempering power: low at one, back to 95% where the power matches the dependence, above it for smaller powers.
Fig. 6 The same sticky coin at ρ=0.6\rho = 0.6, the likelihood raised to a power η\eta before it is combined with the prior, and the coverage of the resulting 95% interval against η\eta. Ordinary Bayes, η=1\eta = 1, covers 68%; the coverage returns to 95% near η=(1−ρ)/(1+ρ)=0.25\eta = (1-\rho)/(1+\rho) = 0.25.

One repair is to discount the likelihood: raise it to a power η\eta less than one before combining it with the prior, so that each observation counts as a fraction η\eta of an observation. The result is called a tempered or fractional posterior. For the sticky coin the right fraction is known exactly — each toss is worth (1−ρ)/(1+ρ)(1-\rho)/(1+\rho) of an independent toss — and tempering with that fraction widens the interval by exactly the missing factor. The figure shows the coverage climbing from 68% at η=1\eta = 1 to 95% near η=0.25\eta = 0.25, and past it for smaller η\eta.

The repair is honest only because the right η\eta was computed from knowledge of the dependence, which is exactly the knowledge the misspecified model lacked. In practice η\eta has to be chosen from the data, and how to choose it is an active question. Peter Grünwald’s “safe Bayes” chooses it by a cross-validation-like procedure that measures how well tempered posteriors predict data they have not seen, and it provably protects against some forms of misspecification; other proposals calibrate η\eta by bootstrapping, or by matching the posterior’s curvature to the sampling variability of the estimate. None is universally right, because the correct amount of discounting depends on how the model is wrong, which is not known.

The same fault in older puzzles

The sticky coin’s error has a familiar shape. It counts correlated observations as if each were separate evidence, and one coin counted by runs and by wakings met the same choice in the Sleeping Beauty problem: whether two awakenings that follow from one coin toss are two pieces of evidence or one. The answer there depended on what was being counted, and the sticky coin makes the cost of the wrong answer measurable — a credible interval that is too narrow by exactly the square root of the overcounting.

The first failure has an older relative too. The door that was not opened turned on what the host was allowed to do, and two children and the sentence about one on how the information was obtained: in both, the same observation supports different conclusions under different models of the process that produced it. Choosing the wrong model of the protocol is choosing a hypothesis list that omits the truth, and the posterior then converges, confidently, to the best answer to a question that was not asked.

What the coin adds to those puzzles is the limit. The puzzles are about one observation; here the observations accumulate, and the error does not wash out. A wrong model of one door gives one wrong probability. A wrong model of ten thousand tosses gives a probability that is wrong and approaches certainty, because the fluctuations that would reveal the error shrink like the wobble of an average while the error itself stays put.

What the wrong model still gets right

It is worth saying what survives. In both examples, the posterior converges to something definite and meaningful. In the first, it converges to the best approximation to the truth the list allows, in the Kullback–Leibler sense, and for many purposes — prediction, compression, decision — that is the right target: a model closest in divergence predicts future tosses better, on average, than any other on the list. In the second, the posterior mean converges to the true long-run bias, because the independent model’s estimate of the proportion of heads is still a consistent estimate of it.

What does not survive is the posterior’s account of its own uncertainty. A posterior probability of 99% for a hypothesis is a statement about the model’s world, and it transfers to the real world only if the model is right. That distinction is invisible from inside the calculation, which is why it is easy to forget: the square picture and the deciban scale both show exactly what the theorem does, and neither can show whether the likelihoods fed into it are true.

What the figures cannot show

The hypotheses here are finite lists and the models are coins, chosen because every quantity can be computed exactly and every divergence written down. In models with continuous parameters the same results hold — Berk’s theorem covers them, and the Bernstein–von Mises theorem for misspecified models, due to Kleijn and van der Vaart, gives the shape of the concentrating posterior — but the posterior’s limiting width then depends on two different matrices, one describing the curvature of the model’s likelihood and one the actual variability of the data, and the credible interval is too narrow or too wide according to how they differ. The sticky coin is the simplest case of that mismatch.

The coverage figures are simulations, 1,500 runs at each dependence and 1,200 for the tempering curve, and each coverage is uncertain by about a percentage point. The agreement with the formula is checked within four points, which is the size of the effect of the normal approximation to the posterior at a thousand tosses.

And a sticky coin is a mild failure of independence. Data with long-range dependence — where correlations fade too slowly to sum to a finite number — can make the effective number of observations grow more slowly than the actual number, and then no fixed tempering fraction repairs the interval at every sample size.

Still open: choosing the discount without knowing the fault

Tempering works when the right fraction is known, and the open question is how to find it from data when it is not. The difficulty is circular: the fraction measures how wrong the model is, and the model is what is being used to judge the data. Proposals that choose η\eta by predictive performance on held-out data work well in some misspecified settings and are known to fail in others, and there is no general theorem saying which procedure gives honest uncertainty for which kinds of misspecification.

A related question concerns the first failure. When the list of hypotheses omits the truth, the posterior’s certainty is certainty about the least wrong entry, and nothing in the posterior signals that every entry is wrong. Detecting that a model is misspecified from the posterior alone — without comparing against an alternative model that happens to be right — is not possible in general, and how much can be learned about the adequacy of a model from its own predictions is a question that model-checking methods address partially and case by case.

Exactness is not the same as truth

The picture that emerges is of a machine that is perfectly reliable and entirely literal. Given a list, it finds the best entry on the list and becomes certain of it; given two equally good entries, it oscillates between them forever; given a model that counts each observation too heavily, it becomes confident at the rate the model prescribes, whatever the data are really worth. In every case it does exactly what Bayes’ theorem says, and in every case the output is a statement about the model’s world.

The decibans of the essay on evidence were lengths laid end to end. What this essay adds is that each length was measured with a ruler the model supplied, and a ruler that is miscalibrated by a constant factor gives lengths that are consistently wrong and consistently confident. The arithmetic of evidence is exact. The evidence is only as good as the model that weighed it.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

Bayes' theoremCoverageCredible intervalIndependenceModel misspecificationPosteriorRandom walkRelative entropy