Evidence measured in decibans
Worth reading first: One coin, counted by runs and by wakings · Bayes' theorem is a picture of a square.
The square picture showed why a test that is 99% accurate can leave a positive result under one chance in five: the condition is rare, and the few true positives are swamped by false positives from the many who do not have it. Every puzzle that followed — the host’s door, the sentence about one child, the sleeping coin — turned on the same move: the prior multiplied by how much more likely the evidence is under one hypothesis than under the other.
That move has a form in which it is not a multiplication at all. Write the probability as odds — the chance that something is so divided by the chance that it is not — and Bayes’ theorem becomes
Take logarithms and the product becomes a sum. On a logarithmic scale of odds, each piece of evidence is a fixed length, and updating a belief is laying lengths end to end.
Two positive tests and a negative
The unit on the axis is the deciban: ten times the base-ten logarithm of the odds. Even odds are 0 decibans; odds of 10 to 1 are +10; odds of 1 to 100 are −20. The scale is chosen so that a factor of ten in the odds is ten units, and a doubling is about three.
The prior, 1%, starts the walk at −20. The first test’s positive result has likelihood ratio — a person with the condition is 19.8 times as likely to test positive as a person without it — and , so the result is worth thirteen decibans. That lands at −7, a probability of 16.7%: the “under one in five” of the square picture, arrived at by one addition instead of four rectangles.
A second positive, from a test with the same accuracy whose errors are independent of the first, adds thirteen more and lands at +6, about 80%. A negative from a third test, weaker — its negative result is five times as likely among the healthy as among the ill, a likelihood ratio of 0.2 — subtracts seven and ends at −1, a probability of 44.2%. The figure computes the same number by multiplying the whole joint table out, and the two agree to rounding.
The decibel scale makes something visible that the probability scale hides. From 1% to 16.7% looks like a big change and from 16.7% to 80% looks like a bigger one; in decibans they are the same step, thirteen units each, because each is the same evidence. Probabilities crowd together near 0 and near 1, and a fixed amount of evidence moves them by amounts that depend on where they started. Log-odds do not crowd: a piece of evidence moves a belief by the same distance wherever it starts.
From the square to the line
The square and the line are the same calculation. In the square, the column of people with the condition is split by the test’s sensitivity and the column without it by its false-positive rate, and the posterior is the orange area divided by the orange and green areas together. Dividing both of those areas by the green one changes nothing, and turns the fraction into odds: the orange area over the green is the prior odds times the ratio of the two splits — which is the likelihood ratio. So the square’s answer, read as odds, is already a product of two numbers, one belonging to the population and one to the test.
The square cannot show a second test without becoming a cube, and a third without becoming something that cannot be drawn. The line shows any number of them, because a product of any number of factors is a sum of their logarithms, and a sum is a sequence of arrows. The odds form does not add anything to Bayes’ theorem. It changes what can be drawn, and what can be drawn is what can be reasoned about at a glance.
What a result is worth
Each result has a weight, and it belongs to the test, not to the person tested.
The two halves of each bar are generally unequal, and the asymmetry is the practical content of the picture. A test that is highly specific — almost never positive in the healthy — makes a positive result decisive: the 99%-specific test’s positive is worth +19. A test that is highly sensitive — almost never negative in the ill — makes a negative result decisive: the 99%-sensitive test’s negative is worth −19.8. That is why a screening test, used on a population where most people are well and the aim is to rule the condition out, is built for sensitivity, and why a confirming test, used afterwards to rule it in, is built for specificity. Each is designed to make one of its two answers heavy.
A test that is 70% sensitive and 70% specific is worth only 3.7 decibans either way. Three such results pointing the same way are worth about as much as one good one, and that is a fair description of what weak evidence is.
The order does not matter
Addition does not care about order, and so neither does the belief at the end.
A doctor who sees the negative result first is, for a while, much more confident that the patient is well — the probability drops to 0.2% — than one who sees the two positives first, who is briefly at 80%. At the end they agree. That agreement is not a matter of luck or of good judgement; it is the commutativity of addition, and it holds whenever the pieces of evidence are independent given the hypothesis. It is also the reason the prior matters exactly as much as each piece of evidence: it is one more length in the sum, and there is no order in which it could be forgotten.
The converse is a useful test of reasoning. If two people reach different conclusions from the same evidence and the same prior only because they saw the evidence in different orders, one of them has made an error — has treated an early result as more or less important than its weight, or has let an early belief change how a later result was read. Anchoring on first impressions is exactly such an error, and on the decibel line it shows up as a path whose end depends on its order.
When the lengths do not add
The rule has one hypothesis, and it fails exactly when that hypothesis fails. The likelihood ratios multiply — the lengths add — only if the pieces of evidence are independent given each hypothesis: among the ill, the two tests’ results are unrelated, and among the well, likewise.
The figure’s two tests have exactly the same sensitivity and false-positive rate in every column; what changes is how their false positives are correlated. When the tests’ errors are independent, two positives are worth 25 decibans between them. When every false positive of one is also a false positive of the other, the second positive carries no new information — a healthy person who fooled the first test will certainly fool the second — and the two results together are worth only as much as one.
Adding the lengths anyway is the error of double counting, and it is the characteristic mistake of reasoning in this form. Two witnesses who heard the same rumour are one witness. Two laboratory tests on the same sample share the sample’s contamination. Two studies that used the same flawed method share the flaw. In each case the evidence looks like two lengths and is really one, and the decibel scale, which makes combining evidence effortless, makes over-combining it effortless too. The remedy is not to abandon the scale but to ask, before adding a length, whether it has already been added under another name.
Turing’s unit
The deciban was named, and the scale used in earnest, at Bletchley Park. Alan Turing’s method for attacking the German naval Enigma, called Banburismus after the town where the long perforated sheets it used were printed, compared intercepted messages two at a time and accumulated evidence about how they were aligned, letter by letter. Each coincidence of letters was a small piece of evidence, worth a fraction of a deciban, and the evidence was added up on the sheets until it passed a threshold. Turing called the unit of of the odds a ban, and a tenth of it a deciban — which I. J. Good later described as roughly the smallest change in the weight of evidence that intuition can register.
Good, who worked with Turing, called the quantity the weight of evidence, and it is the same number that Claude Shannon’s information theory measures in bits: one bit is about three decibans, the evidence of a single fair coin toss’s worth. The two scales differ only in the base of the logarithm.
The choice of a tenth of a ban was practical. The evidence from any one pair of letters was tiny, and the sheets had to be scored by hand by people working quickly; a unit small enough that each coincidence contributed a few whole units, and a threshold of a few dozen of them, turned a probabilistic inference into a column of additions. The method did not require anyone doing the scoring to know what a likelihood ratio was.
How much a single toss is worth, on average
The biased coin gives a way to ask how informative an experiment is before it is run. A head is worth decibans for “biased”, a tail . On a coin that really is biased, heads come 60% of the time, so the expected weight of evidence from one toss is
That small positive number is the drift of the walk in the next figure, and it predicts how long the walk takes: twenty decibans at 0.086 a toss is about 232 tosses, against the 229 the simulation measured. The expected weight of evidence per observation is a quantity with a name in information theory — the Kullback–Leibler divergence between the two hypotheses’ distributions — and it is the exact exchange rate between the size of an effect and the number of observations needed to detect it. A coin biased to 60% is close to fair, the divergence is small, and hundreds of tosses are needed; a coin biased to 90% would give about 1.6 decibans per toss, and the question would be settled in about thirteen. The same quantity limits how many fair bits can be squeezed out of a biased coin: a coin close to fair carries a lot of randomness and very little evidence about its own bias, and one far from fair carries the reverse.
Stopping when the evidence is decisive
If each observation adds a length, the natural procedure is to keep observing until the total is long enough — and to stop as soon as it is.
The running total is a random walk: each toss moves it up by 0.79 or down by 0.97, and on a coin that really is biased it drifts upward. The test stops the walk at the first barrier it reaches, which is the gambler’s-ruin problem with two barriers in a new setting.
The guarantee is a property of the odds, not of the coin. Under the hypothesis that turns out to be false, the odds in its favour — the ratio of the two likelihoods — form a fair game — the same device of fair bets placed on a coin that settles pattern races: their expected value after the next toss is what they are now. A fair game started at 1 reaches 100 before 0 with probability at most one in a hundred, so the chance of stopping at +20 decibans when the coin is really fair is at most about 1%, and the same argument bounds the other error. The simulation, 0.90%, sits just under the bound, which is loose only by the overshoot past the barrier on the last toss.
Abraham Wald proved in 1945 that this sequential probability ratio test needs, on average, fewer observations than any fixed-sample test with the same error rates; Turing had used the same idea at Bletchley, unpublished, some years before. Both saw the same thing: when evidence is a sum, the sensible time to stop is when the sum is big enough, and the size of “big enough” is exactly the error rate that can be tolerated.
What the scale leaves out
The likelihood ratios are given, not measured. Every figure takes the sensitivity and specificity of its tests as exact numbers. In practice they are estimates from studies, with their own uncertainty, and a posterior computed from uncertain likelihood ratios is itself uncertain in a way the single number at the end of the arrows does not show.
Independence is assumed except where it is the subject. The deciban scale, the orders and the sequential test all assume that the pieces of evidence are independent given each hypothesis. The dependence figure shows what goes wrong without it for one particular kind of correlation; real dependence can take many forms, and no picture here measures it.
Averages hide the tail. The sequential test’s 229 tosses is an average; individual runs in the figure took from about fifty to nearly four hundred, and some runs of the simulation took far longer. A procedure that stops when the evidence is sufficient has no fixed length, and planning around its average can be badly wrong for a single run.
Two hypotheses only. Odds compare one hypothesis with one alternative. With three or more — ill with one disease, ill with another, well — the decibel scale becomes a vector of log-odds, one for each hypothesis against a reference, and the arithmetic, though still additive, can no longer be drawn on one line.
Still open: how much evidence a result carries when the model is wrong
The arithmetic is exact once the likelihoods are known. The hard questions are about the likelihoods. When the model that assigns them is wrong — when the tests’ error rates differ between the population studied and the population tested, or the hypotheses considered do not include the truth — the posterior can be confidently wrong, and there is no general account of how wrong. Robust versions of Bayesian updating, which temper each likelihood ratio before adding it, exist in several forms, and which of them is right for which kind of misspecification is actively argued.
A related question is the sequential test’s. Wald’s test is optimal for two simple hypotheses; with a continuum of alternatives — a coin whose bias is unknown rather than one of two values — the optimal stopping rule is known only in special cases, and the practice of stopping a clinical trial early when the evidence looks strong has statistical properties that are still a subject of methodological dispute. Trials stopped early for benefit tend to overstate the benefit — the walk crossed the barrier partly because it happened to be running high — and how best to correct for that is not agreed.
Evidence as a length
The whole of Bayes’ theorem, for two hypotheses, is a single operation on a single line: start at the prior and add the weight of each piece of evidence. The operation is exact, it does not care about order, and it makes plain what a result is worth before it is known. What it cannot do is notice when two lengths are the same length, drawn twice, and nearly every error in reasoning of this kind is that one.
That is also the thread back through the puzzles of this subject. The host who opens a door, the parent who mentions a son, the coin that wakes a sleeper twice: each was a question about what the likelihood ratio really was — about what the evidence would have looked like under each hypothesis, given how it came to be reported. Once that number is right, the arithmetic is a single addition. Getting it right is the whole of the difficulty, and the decibel line makes the difficulty visible by making everything else trivial.
What links here
Computed from the collection, not written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A majority wiser than its members — both name bayes' theorem, independence
Named objects
A dashed tag is an object no other essay names yet.
Bayes' theoremConditional probabilityIndependenceMartingaleRandom walk