Which vote finds the truth more often
Worth reading first: Deciding the premises or the conclusion · A majority wiser than its members.
The court that contradicts itself put three judges in front of a case decided by two premises. The defendant is liable if there was a valid contract and it was breached. Each judge reasons consistently, and still the majority finds a contract, the majority finds a breach, and the majority finds the defendant not liable. Deciding the premises or the conclusion set out the two ways a body can escape: vote on each premise and let the verdict follow, or vote on the verdict and let the premises look after themselves. The two procedures reach opposite answers on exactly the troubled profiles, and the impossibility theorems of the essay after it say neither can be made to satisfy every reasonable demand.
All of that treats the judges’ views as preferences to be combined. A court is not a committee choosing a menu. There is a fact of the matter — there was or was not a contract — and the judges are trying to find it. That changes the question from which procedure is fairest to which procedure is right more often, and the second question, unlike the first, has an exact answer once a model of the judges is fixed. This essay fixes the simplest one and computes the answer, and the answer has a surprising shape: it depends on how competent the judges are, on how many there are, and on how often defendants are in fact liable, and the conclusion vote turns out to be quietly doing something that courts usually do on purpose.
The model, stated in full
Each judge, on each premise, is right with the same chance — the competence — independently of every other judge and of the other premise. The judge then reasons consistently: votes liable exactly when voting yes on both premises. That is the whole model, and it is the one a majority wiser than its members used for a single question, carried to two.
The procedures are then two different majority votes. The premise vote takes a majority on the contract, a majority on the breach, and finds liability when both majorities say yes. The conclusion vote takes a majority on each judge’s own verdict. There are four possible worlds — contract and breach each real or not — and in each world each procedure has a definite chance of reaching the right verdict, which is a sum of binomial terms. For three judges the figure computes those sums and checks every one against brute enumeration of the ways six individual judgements can each be right or wrong.
The opening figure already shows the shape of everything that follows. When the defendant is liable, the premise vote finds it 61.5% of the time and the conclusion vote only 48.5% — worse than a coin. When the defendant is not liable, in any of the three ways, the conclusion vote is the better. Averaged over the four worlds weighted equally the two are within a fifth of a percentage point: 80.7% against 80.9%.
Why the conclusion vote leans towards no
The asymmetry has a one-line reason. A judge votes liable only by getting both premises right in a world where both hold — chance — or by making exactly the errors that manufacture a false yes elsewhere. With , a judge in the liable world says liable with chance : each judge is individually more likely to acquit a liable defendant than to convict one, even though the same judge is right about each premise seven times in ten. Multiplying two good chances produces a poor one, and the conclusion vote is a majority of judges each of whom is, on the conclusion, worse than a coin.
In the worlds where the defendant is not liable the same multiplication helps. A judge wrongly says liable only by being wrong about whichever premise fails — and, if both fail, wrong about both — so the chance of a false yes is or . On the conclusion, those judges are better than they are on either premise, and their majority is very reliable.
The premise vote does not multiply individual chances; it multiplies majority chances. Each premise is decided by a majority of judges who are each right seven times in ten, which the jury theorem says is better than any one of them — 78.4% for three. The verdict is right in the liable world when both majorities are right, . The premise vote aggregates first and combines second; the conclusion vote combines first and aggregates second, and the order matters because combining two uncertain things always loses information and aggregating can win it back.
Three judges, and a crossover at exactly the square root of a half
Averaged over the four worlds weighted equally, which procedure is right more often? For three judges the answer can be written in one line. The difference between the premise vote’s average chance and the conclusion vote’s is
which the figures check numerically at several competences. Every factor but the last is positive, so the sign is the sign of : the conclusion vote is right more often when , and the premise vote when . At exactly the two tie — the competence at which a single judge is right about both premises exactly half the time.
For larger courts the crossover moves down, and the curves separate. With fifteen judges the premise vote is better above a competence of 0.576, and with a hundred and one above 0.527. The premise vote, being two applications of the jury theorem, climbs to certainty as the court grows at any competence above a half. The conclusion vote does something odd: for a large court it stalls near three in four, because it becomes certain in the three worlds without liability and certainly wrong in the fourth — until competence passes , when it too climbs to certainty.
The world in which a larger court is worse
That stall deserves its own picture, because it is the most striking fact here and it holds for any court size. Restrict attention to the world in which both premises are true and the defendant is liable.
In this world each judge says liable with chance , so the conclusion vote is a jury-theorem majority of judges whose competence on the question asked is . The jury theorem cuts both ways: a majority of voters each right more than half the time becomes certain to be right as the group grows, and a majority of voters each right less than half the time becomes certain to be wrong. Between a half and , every judge is more often right than wrong about every premise, and the larger the court, the surer it is to acquit a liable defendant. At competence 0.69 a court of 201 finds liability less than four times in ten; at 0.72 it finds it more than six times in ten. The premise vote, at competence 0.55, already finds it more than three times in four.
The figure’s dashed curves all cross at one half at exactly , which is the jury theorem’s threshold transplanted: . It is a clean example of a thing the impossibility theorems cannot see. They say the two procedures must sometimes disagree; this says that when the truth is liability, the conclusion vote is systematically and increasingly wrong for a whole band of competence.
The liable world, worked by hand
The three-judge numbers are small enough to write out, and doing so shows exactly where the two procedures part. In the liable world a premise majority is right when at least two of three judges are right about it, which has chance
and the premise vote is right when both majorities are, with chance . The conclusion vote is right when at least two judges are right about both premises, each with chance , so its chance is . At these are against ; at , against .
So the comparison in the liable world is between and — squaring after the majority, or before it — and the majority function is the jury theorem’s amplifier, pushing chances above a half up and chances below a half down. Squaring first sends below a half whenever , and the amplifier then works against the court. Squaring after keeps the amplifier on the right side of a half. In the three worlds without liability the roles reverse, because there the squared quantity is the chance of a wrong yes, and pushing that down is exactly what is wanted. Everything in the figures is these two compositions, after squaring and squaring after , computed for more judges and averaged with different weights.
Where the question came from
The paradox was named in legal scholarship before it reached social choice. Lewis Kornhauser and Lawrence Sager described the doctrinal paradox in 1986, from the practice of multi-member appellate courts in which judges agree on a judgment for different reasons, and asked whether courts should decide issue by issue or case by case. Philip Pettit generalised it in 2001 to the discursive dilemma faced by any group that must give reasons as well as conclusions, and Christian List and Pettit turned it into the impossibility theorems of the essay that follows the court’s contradiction.
The truth-tracking question was asked soon after. Christian List in 2005 computed how probable inconsistent majorities are under a model like the one here, and Luc Bovens and Wlodek Rabinowicz in 2006 compared the two procedures’ reliability, finding the pattern drawn above: the conclusion procedure can win for small groups of modest competence, the premise procedure wins for large ones, and the outcome is sensitive to the base rate of the conclusion being true. What the figures add is the exact three-judge polynomial and its root at , the court-by-court crossover curve, and the split of each procedure’s mistakes into the two kinds that the law treats differently.
Crossovers, court by court
The average comparison depends on the court’s size in a smooth way, and the crossover can be followed from three judges up.
The curve falls fast at first — 0.707 for three judges, 0.656 for five, 0.593 for eleven — and then slowly, to 0.527 at a hundred and one. In the limit any competence above a half favours the premise vote, because the premise vote becomes certain in all four worlds and the conclusion vote stays wrong in one of them until . For a body of any real size, if its members are better than a coin on each question, it finds the truth more often by voting on the questions.
The small courts are where the choice is genuinely close, and they are the ones that matter in practice: appellate benches of three, five, seven or nine. For a bench of three with judges right seven times in ten on each issue, the two procedures are a fraction of a percentage point apart on average and very far apart in particular worlds — which says that the average is the wrong thing to look at.
It depends on the cases, not only on the judges
The average weighted the four worlds equally, which assumes that a quarter of defendants are liable. Nothing about courts makes that true, and the comparison is sensitive to it.
Because the conclusion vote is very good in the three worlds without liability and poor in the one with it, its average rises as liability becomes rarer. For eleven judges at competence 0.65 the two procedures tie when 15.3% of defendants are liable. A court whose docket is mostly meritless claims finds the truth more often by voting on the conclusion; one whose cases are mostly well founded finds it more often by voting on the premises.
That is an uncomfortable conclusion for anyone hoping for a procedural rule that is simply correct. Which procedure tracks the truth better is not a property of the procedures alone; it depends on a fact about the population of cases that the court cannot see from inside any one of them. The impossibility theorems already said no procedure is right on every profile; this says that even the question “which is right more often” has an answer that moves with the docket.
Two kinds of wrong verdict
Counting a verdict as simply right or wrong hides the most important difference between the procedures. A wrong verdict can find liability where there is none, or miss it where there is, and legal systems do not treat those alike.
With fifteen judges of competence 0.7, the conclusion vote finds liability wrongly 0.39% of the time and the premise vote 3.25% — eight times as often. In return the conclusion vote misses true liability 53.1% of the time, against 9.8%. The conclusion vote behaves like a high standard of proof, the kind criminal courts adopt deliberately — a rule that accepts many wrong acquittals to avoid a few wrong convictions — and it does so without anybody choosing the standard. The premise vote behaves like a civil standard, roughly balancing the two errors.
This reframes the old debate. Lawyers arguing for the conclusion-based procedure have often said that it protects the defendant: a defendant should be found liable only if a majority of judges, each reasoning the whole case through, find liability. The model says they are right about the effect and that the effect is large, growing with the size of the court. Whether it is desirable is the question what the agenda leaves standing ended on — what a wrong yes costs against a wrong no — and the model cannot answer it. It can say exactly how much of each kind of error each procedure buys.
What the model leaves out
Everything above rests on three assumptions, and each is false of real courts in a way that changes the numbers.
The judges are independent. Real judges deliberate, read the same briefs and share training, and their errors are correlated; a majority wiser than its members showed that correlation caps how much a majority can gain, and the premise vote’s advantage is exactly a majority’s gain, so correlation erodes it most. The judges have equal competence on both premises. A court may be expert on the law of contract and poor at assessing evidence of breach, and with unequal competences and the threshold becomes the curve rather than a number. And the judges vote sincerely. A judge who knows which procedure is in use and cares about the outcome can vote strategically on a premise to secure a verdict, and a lie that pays is the general reason such opportunities exist.
None of these changes the qualitative picture: the conclusion vote leans towards no because it multiplies before it aggregates, and the premise vote is better at finding liability for the same reason. But they change every crossover in the figures, and the figures say which model they draw.
Still open: many premises, and judges who listen
With two premises the arithmetic is small enough to do exactly, and the threshold is because a judge must be right twice. With premises conjoined the same argument gives a threshold of — for five premises, competence must exceed 0.87 before the conclusion vote finds liability reliably in a large court — and the premise vote’s advantage in the liable world grows correspondingly. But agendas in practice are not pure conjunctions, and for the mixed agendas of agendas that cannot contradict themselves, with disjunctions, chains and exceptions, the truth-tracking comparison has been worked out only case by case.
The deeper open question is about deliberation. The model has judges who form views alone and then vote; real panels talk first. Whether discussion moves judges towards the truth or merely towards one another — whether it raises competence or only correlation — is an empirical question, and the mathematical models of deliberation that exist give opposite answers under different assumptions. Until it is settled, the clean results here describe a court that votes without conferring, which no appellate court does.
A procedure is also a lens
The doctrinal paradox was first presented as a failure: a court that contradicts itself. Seen as a problem of finding the truth it is something more useful. The two procedures are two different instruments for looking at the same evidence, with different sensitivities: one combines the judges’ answers to each question and then reasons, and the other lets each judge reason and then counts. The first is sharper in general and the second has a built-in bias towards no.
For three judges the choice turns on whether a single judge gets both premises right more than half the time, . For larger courts it turns, almost always, in favour of voting on the questions. And for any court it turns also on what the court is trying to avoid — which is why the procedure a legal system chooses is best read not as an answer to a puzzle in logic, but as a decision, made once and in advance, about which of two kinds of mistake it would rather make.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A bell curve assembled out of coin flips — both name binomial distribution, independence
- A walk that always comes home, until it does not — both name binomial distribution, independence
- Four ways out, and what each costs — both name judgement aggregation, majority rule
- The nearest consistent verdict — both name judgement aggregation, majority rule
- Two weak sources make one fair bit — both name bias, independence
Named objects
A dashed tag is an object no other essay names yet.
BiasBinomial distributionCondorcet jury theoremIndependenceJudgement aggregationMajority ruleProposition