A majority wiser than its members
Worth reading first: How often the majority goes in a circle · A bell curve assembled out of coin flips.
Everything in the study of voting rules so far has treated a ballot as a preference: what a voter wants, with nothing for it to be right or wrong about. Condorcet’s Essai of 1785, the same book that introduced the majority cycle, also treats a ballot the other way — as a judgement about a question that has an answer. Is the accused guilty. Will the bridge hold. Is this proposal better for the town than that one, in some sense everyone agrees on and nobody can check in advance.
Read that way, a vote is a noisy measurement, and the question is no longer whether majority rule is coherent but whether it is accurate. Condorcet’s answer is the most optimistic theorem in the theory of voting. If each voter is independently more likely to be right than wrong, however slightly, then the chance that a majority is right grows with the number of voters and tends to certainty. A crowd of people who are each barely better than a coin is, as a body, nearly infallible.
The theorem is a law of large numbers, and it has exactly the strengths and the weaknesses of one. The figures below compute it exactly, measure how large a crowd has to be, and then take away its hypotheses one at a time — equal competence, independence, equal weight — to see which ones it can live without.
The binomial, and one very small jury
Take voters, an odd number so that a majority always exists, each right with probability and wrong with probability , independently of the others. The number who are right is binomial, and the majority is right when more than half are:
Three voters at make this small enough to do by hand. The majority is right if all three are, with probability , or exactly two are, with probability . Together that is : three jurors who are each right 60% of the time make a majority that is right 64.8% of the time. The improvement comes entirely from the cases where one juror errs and the other two outvote the error.
The figure is the theorem in one picture. Every curve for above one half rises towards certainty and every curve below falls towards zero — the mirror image is part of the theorem and is usually forgotten. A majority amplifies whatever its members have: a slight tendency to be right becomes near-certainty and a slight tendency to be wrong becomes near-certain error. At exactly nothing happens at all, and the majority is a coin however large it grows.
The reason is the averaging that turns coin flips into a bell curve. The fraction of voters who are right has mean and a spread that shrinks like . Once that spread is small compared with the gap between and one half, almost the whole distribution lies on the right side of one half, and the majority is right.
How large a crowd has to be
The same picture says how many voters are needed, and the answer is sobering at the low end. The spread of the right-fraction is , which for near one half is about . To put the whole distribution a safe number of spreads above one half requires , or
with for 95% reliability and for 99%.
The exact search agrees with the rough formula: at 51% competence a jury needs 6,763 members to be right 95% of the time, at 60% it needs 67, at 80% it needs seven. The dependence is on the square of the edge over a coin, so halving each voter’s edge quadruples the jury. A population of voters who are each right 50.1% of the time would need nearly seven hundred thousand of them to reach 95% — possible in a national referendum, absurd in a committee.
This is the same arithmetic that governs how many random samples an estimate needs: the error falls like one over the square root of the count, so each further digit of accuracy costs a hundred times the effort. A jury is a Monte Carlo estimate of which side of one half the truth lies, with each voter a sample.
The vote that decides, and why it almost never does
The same arithmetic has an uncomfortable consequence for each individual voter. A single vote changes the outcome only when the others are split exactly evenly, and the chance of that is one term of the binomial: for voters, the probability that the other divide equally is .
At this falls slowly, like : about one chance in eighteen of being decisive in a body of 201, one in forty in a body of a thousand. But as soon as moves away from one half the factor takes over, and it falls exponentially. For voters who are each right 60% of the time, a single vote in a body of 201 is decisive about once in a thousand questions; in a body of a thousand, about once in thirty billion; in a body of two thousand, about three times in a hundred billion billion. The majority is reliable precisely because no individual vote is ever close to mattering.
That exponential fall is the large-deviation tail of the binomial — the probability that an average lands far from its mean decays exponentially in the number of terms, at a rate set by how far — and it is the reason the jury theorem and the individual voter’s position pull in opposite directions. A body is trustworthy when its members are individually negligible. It also sharpens the strategic problem of the section below: a voter who is almost never decisive can afford to reason about the rare occasions on which they are, and those occasions carry information of their own.
Voters worse than a coin
The theorem is usually stated for voters of equal competence, and that hypothesis looks essential: surely a jury containing members who are worse than a coin is dragged down by them. It is not, provided the average survives.
What decides the outcome is the expected number of right votes, , against the threshold , together with the spread about that expectation. Mixing competences changes the expectation not at all — it is fixed by the average — and it reduces the spread, since each voter’s variance is largest for voters near one half and smaller for voters near the extremes. So a jury of mixed competences is, if anything, more reliable than a uniform jury with the same average. The spread claim is a one-line consequence of the curve being concave: its average over the voters is at most its value at the average competence. Wassily Hoeffding turned it into a statement about the probabilities themselves in 1956 — among independent voters with a given expected number of right votes, when the majority threshold sits far enough below that expectation, the uniform jury is the least reliable of all.
The figure’s jury is extreme on purpose. Four in ten of its members would do better to vote the opposite of their judgement, and the body as a whole is still nearly infallible at two hundred members. The requirement is not that each voter be better than a coin. It is that the voters, on average, are.
Voters who share their mistakes
The hypothesis the theorem cannot live without is independence, and it is the one real juries, committees and electorates violate most obviously. Voters read the same newspapers, hear the same witness, defer to the same persuasive colleague. When one of them is misled, others are misled in the same direction.
A simple model captures this. Suppose the voters share a common factor: some question-specific competence , drawn once for the whole electorate, and then each voter is right with probability independently. On average is , but on some questions it is higher and on some lower, and the correlation between two voters being right measures how much varies. At the voters are independent and the theorem applies.
With any correlation at all, the majority’s accuracy has a ceiling below certainty, and the ceiling is easy to state: it is the chance that the shared factor leaves voters better than a coin, . On a question where the common factor makes everyone likelier wrong than right, a large majority will be wrong, and no number of additional voters can help, because they are all being misled by the same thing. The large jury does not estimate the truth; it estimates , and reports which side of one half fell on.
The ceilings are low for modest correlations. Two voters whose chances of being right are correlated at only — a dependence so weak that no pair of votes would reveal it — hold the majority of an arbitrarily large body to 73.6%, barely better than a single voter at 60%. A correlated electorate behaves like a small number of independent voters, roughly of them, however many individuals it contains. This is the voting version of a lesson from the theory of concentration: averages concentrate when no single input can move them far, and a common factor is an input that moves everyone at once.
Counting heads against weighing them
The last hypothesis is that every vote counts equally. When the voters’ competences are known and differ, simple majority is not the best rule, and neither is deferring to the best voter.
The right weights are the logarithms of each voter’s odds of being right, , and the rule is to add the weights on each side and choose the heavier. This is not a heuristic. With the two answers equally likely beforehand, Bayes’s theorem says the best any rule can do with a given split of votes is to back whichever answer makes that split more probable, and the log of that likelihood ratio is exactly the sum of the log-odds weights on one side minus the other. Shmuel Nitzan and Jacob Paroush stated it for committees in 1982; the figure checks it by computing the Bayes-optimal accuracy directly over every split and finding it equal to the weighted rule’s.
The rule needs the competences, and that is its weakness as a procedure. When the competences are equal the log-odds weights are equal and weighted majority is simple majority, so simple majority is the best rule precisely when nobody knows who is better — which is one honest defence of counting heads. When the competences differ but are estimated badly, weighting by the estimates can do worse than counting, and the figure’s optimum assumes the numbers are exact.
The weights explain when to defer to an expert. A voter who is right 90% of the time carries a log-odds of , more than five voters at 60% put together, whose log-odds of each add up to only , so the optimal rule for that panel simply follows the expert. Add a sixth 60% voter and the crowd can outvote the expert — but only when all six agree. A vote’s weight grows like the log of its odds, which is slow at first and then fast: two voters at 90% outweigh any number of voters at 60% up to ten.
Condorcet’s resolution of his own paradox
The two halves of Condorcet’s Essai meet in a way that deserves to be better known. He did not regard the majority cycle as a defect of majority rule. He regarded it as evidence.
Suppose there is a true ranking of the candidates, and each voter judges each pairwise comparison correctly with probability above one half. Then the jury theorem applies to every pair separately: in a large electorate each pairwise majority is almost certainly right, and so the majority relation is almost certainly the true ranking — transitive, with a winner. A cycle can only appear when some pairwise majority is wrong. So a cycle, when it occurs, is evidence that the voters are unreliable on at least one pair, and the sensible response is to reverse the pairwise verdict that is least securely held.
That is precisely a rule, and two centuries later it was recognised. H. Peyton Young showed in 1988 that Condorcet’s recommendation, done properly, is the Kemeny rule: choose the ranking that disagrees with the fewest pairwise ballot judgements, which under Condorcet’s model is the most likely true ranking. It is the same rule that the nearest consistent verdict reaches from judgement aggregation, by asking for the consistent verdict closest to the judges. The same object arrives from two directions — as the most probable truth behind noisy judgements, and as the least-changed repair of an inconsistent majority.
With three candidates the recommendation is simple and was Condorcet’s own: a cycle has three pairwise majorities, and the most likely true ranking reverses the one carried by the smallest margin. With more candidates it is not simple at all. Finding the ranking that disagrees with the fewest pairwise judgements is a search over every ordering, and John Bartholdi, Craig Tovey and Michael Trick showed in 1989 that no shortcut is known to be possible: deciding the Kemeny winner is as hard as the hardest problems of its kind. The most principled answer to Condorcet’s paradox is also the most expensive to compute.
What the curves cannot show
They cannot show that the model applies. Every figure assumes that the question has a right answer, that each voter is right with a stated probability, and that the voters’ errors are independent or share a stated common factor. None of these can be read off a tally. In particular, independence cannot be checked from the votes themselves: a unanimous vote is equally consistent with a hundred independent experts and with one expert copied a hundred times, and the theorem’s conclusion differs enormously between them.
They cannot show strategic voting. The theorem assumes each voter votes their honest judgement. David Austen-Smith and Jeffrey Banks showed in 1996 that honest voting need not be rational: a juror who reasons that their vote only matters when the others are evenly split should condition on that event, which carries information about the question, and may then rationally vote against their own judgement. A lie can pay among voters with judgements as it can among voters with preferences.
And they cannot show the questions a binary vote cannot ask. The theorem is about a yes-or-no question. With three or more options, independent reliable judgements about each pair can still produce a cycle in a finite electorate, and which option a large body would choose depends on the rule that breaks the cycle — which is where the section above began.
Still open: how much correlation real judgements carry
The theory is complete for the models drawn and has an open end at exactly the point where it meets practice. The ceiling on a correlated body’s accuracy depends on how the common factor is distributed, and there is no general way to estimate it from votes alone, since a body’s votes on a single question are one sample of the common factor. Across many questions with known answers the correlation can be measured, and any scheme that pools judgements has to do something of the kind to know what its pooling is worth; but the measurement is specific to the questions used, and nothing in the mathematics says it transfers to the next one.
On the theoretical side, the effect of deliberation is the unsettled part. Discussion before a vote can share information, which improves each voter’s competence, and it can share mistakes, which raises the correlation, and whether a given procedure does more of the first than the second is not decided by any theorem here.
A law of large numbers, with its conditions
If voters judge a yes-or-no question independently and each is more likely right than wrong, a simple majority is right with a probability that rises to certainty as the body grows — a binomial law of large numbers, with a mirror image for voters more likely wrong. The body needed grows like the inverse square of each voter’s edge over a coin: 6,763 voters at 51% for 95% reliability, 67 at 60%.
The theorem needs only the average competence to beat one half; mixing competences does not hurt. It does not survive correlation: voters who share a common factor have a ceiling, the chance that the factor favours the truth, and behave like about independent voters however many there are. When competences differ and are known, weighing votes by their log-odds beats both counting them and following the best voter, and it is the best any rule can do. Condorcet used the same theorem to read a majority cycle as evidence that some pairwise majority is wrong, and the repair it implies is Kemeny’s rule.
A crowd is wise exactly to the extent that its errors are independent — the size of the crowd is the less important number.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Two children and the sentence about one of them — both name bayes' theorem, independence, law of large numbers
- A walk that always comes home, until it does not — both name binomial distribution, independence
- Averaging down the triangle — both name binomial distribution, law of large numbers
- One coin, counted by runs and by wakings — both name bayes' theorem, law of large numbers
- The average settles and the wobble does not — both name independence, law of large numbers
- The door that was not opened — both name bayes' theorem, independence
Named objects
A dashed tag is an object no other essay names yet.
Bayes' theoremBinomial distributionCondorcet jury theoremCorrelationIndependenceLaw of large numbersPairwise majorityProbability modelWeighted majority