Every premise makes the verdict vote stricter
Worth reading first: Which vote finds the truth more often · A majority wiser than its members.
A court finds a defendant liable for breach of contract only if there was a contract and it was broken. When three judges disagree — one doubting the contract, one doubting the breach — a majority can accept each premise while a majority rejects the verdict, and the court must choose whether to decide premise by premise or by counting the judges’ own verdicts. The comparison of the two as ways of finding the truth showed that with two premises and three judges the verdict vote is right more often below a competence of exactly , and the premise vote above it.
Real legal tests often have more than two elements. Negligence needs a duty, a breach of it, causation and damage; many statutes list five or six conditions, all of which must hold. This essay asks what happens to the comparison as the number of premises grows, and the answer is that the verdict vote’s weakness — it is cautious about finding liability — compounds with every element, while the premise vote’s strength does not erode at all.
The competence the verdict vote needs
Each judge is independently right about each premise with chance and votes for liability exactly when voting yes on every premise. In the world where the defendant is in fact liable — every premise true — a judge’s own verdict is right only if all premise judgements are right, which happens with chance . Condorcet’s jury theorem says a large majority of independent voters follows whichever side each voter is more likely to be on: so a large court voting on the verdict finds a liable defendant liable only if .
The threshold climbs steeply. With one premise there is no difference between the procedures. With two, judges must be right on each element more than seventy times in a hundred. With five, eighty-seven. With ten, ninety-three. The premise vote’s threshold, meanwhile, is a half for every : a majority on each premise is right when each judge is more likely right than wrong about it, and if every premise majority is right, so is their conjunction.
Five elements and three judges, by hand
The mechanism is easiest to see in a single case. Suppose a test has five elements and the defendant satisfies all of them. Three judges hear it. Judge one doubts the first element and accepts the rest; judge two doubts the second; judge three doubts the third. Each judge is right about four elements in five, which is very good judging.
Voting element by element, the court accepts every element: the first by two votes to one, the second by two to one, the third by two to one, and the fourth and fifth unanimously. The conjunction holds and the court finds liability — correctly. Voting on the judges’ own verdicts, each judge has rejected one element and so votes against liability, and the court finds for the defendant by three votes to nil. Every judge was right about most of the case, the court was right about every element, and the verdict vote produced a unanimous wrong answer.
Nothing about this example is contrived except the neatness of the errors. With five elements, a judge right four times in five on each makes at least one mistake about two times in three — — and any single mistake turns that judge’s verdict against liability. The judges’ mistakes need not coincide for the verdict vote to fail; they only need each judge to have one somewhere, and with enough elements nearly every judge does.
Where a larger court is worse
The threshold describes large courts. The figure below follows real sizes, in the world where the defendant is liable, for judges right four times in five.
At four-in-five competence, three premises put the verdict vote just above its threshold — — and the court’s chance of finding liability rises only slowly, to 0.6 with ninety-nine judges. Five premises put it below: , a judge’s own verdict is more likely wrong than right, and adding judges amplifies the error. A three-judge bench finds a liable defendant liable about a quarter of the time; a ninety-nine-judge court, essentially never. The jury theorem works in both directions, and here it works against the court.
The premise vote, at the same competence, is above 0.99 by twenty judges for every number of premises drawn. The contrast is starkest at the sizes courts actually have. An appellate bench of three, at five elements and four-in-five competence, finds a liable defendant liable a quarter of the time by the verdict vote and fifty-eight times in a hundred by the premise vote. A court of fifteen does better than the bench by the premise vote, ninety-eight times in a hundred, and worse by the verdict vote, eight times in a hundred. It is not that the premise vote is cleverer. It is that it aggregates before it combines, and aggregation is where the jury theorem’s amplification happens: each premise’s majority is pushed towards certainty first, and the conjunction of near-certainties is near-certain. The verdict vote combines first — each judge multiplies their own probabilities — and only then aggregates, and a product of several probabilities below one can fall below a half however good each factor is.
How fast a large court goes wrong
Below the threshold, the verdict vote does not merely fail; it fails at an exponential rate. Each judge’s verdict is a coin that comes up “liable” with chance , and the chance that a majority of such coins come up “liable” when shrinks like , where measures how far is from a half — the relative entropy of a fair coin against the biased one. For five premises at four-in-five competence, and , so every sixteen judges added cut the chance of a correct finding of liability by a factor of about . A court of ninety-nine finds liability correctly with chance about one in five thousand.
The same exponent works for the premise vote in its favour. There each premise majority is a coin with , far above a half, and its error shrinks like with a much larger ; five such errors added together are still tiny. So the comparison is not between two procedures that are each mostly right. Past the threshold it is between one that converges to certainty and one that converges to certain error, at comparable exponential speeds, and the size of the court only sharpens the difference.
Where the threshold falls matters for actual legal tests. Negligence has four classical elements, and the verdict vote needs judges right on each more than 84 times in a hundred. A judge who is right on each element eight times in ten — which, for questions of fact in contested cases, is not a low standard — is already below it.
An average that hides the liable world
The comparison with two premises was summarised by averaging over the four combinations of true and false premises, weighted equally. Applying the same summary to more premises gives a misleading result, and the reason is worth seeing.
By this average the verdict vote looks better and better as premises are added: for three judges, the premise vote needs competence above 0.75 to win with three premises and above 0.83 with eight. But equal weighting makes the liable world one case in — one in 256 at eight premises — and in every other world the verdict vote is excellent, because its caution is exactly right when the defendant is not liable. The average is dominated by the worlds where the question is easy and buries the one where it matters.
This is a general caution about evaluating procedures by averages. The equal weighting amounts to assuming that each premise is true half the time, independently, so that liable defendants are rare in proportion to the number of elements. A court that hears mostly meritorious cases faces a different mixture, and what each procedure is good at depends on which cases it is given. The figures here do not choose the weights; they show that the choice decides the comparison.
How often the two disagree
The doctrinal paradox — the two procedures giving different answers — can be counted directly by simulating courts.
Over all cases the paradox is uncommon and becomes rarer with many premises: it peaks at about nine per cent with three premises and falls to under one per cent with eight, because with many premises most cases have several false ones and both procedures say no. Among liable defendants the picture is the opposite. The two procedures disagree on a quarter of them with two premises, three in five with three, and about three in four from four premises on. The paradox is concentrated in exactly the cases a court exists to decide correctly, and in those cases the procedures do not merely differ — one is usually right and the other usually wrong.
The premise vote never contradicts itself here
A procedure that is more truthful could still be objectionable if it produced incoherent judgements, and in general voting proposition by proposition can contradict itself. For a purely conjunctive test it cannot. The premise majorities can be any pattern of yes and no, and the court’s verdict is defined as their conjunction, so the set of things the court holds — each element’s finding and the verdict — is always consistent. The contradiction in the doctrinal paradox lives entirely in the verdict vote, which can reject a conclusion that follows from premises the court has accepted.
That is why the choice between the two procedures is usually framed as one between coherence and deference to the judges’ own reasoning, and why each of the escapes from the impossibility theorem costs something different. The truth-tracking comparison adds a third consideration, and with many elements it points firmly one way. Repairing an inconsistent set of majorities by moving to the nearest consistent verdict is a third procedure, a compromise that sometimes sides with the premise majorities and sometimes with the verdict majority depending on how lopsided each is; how it tracks the truth as elements multiply is not computed here.
When any one ground is enough
Not every legal test is a conjunction. Some are disjunctions: a contract may be void for any of several reasons, a defendant liable under any one of several heads. The two procedures can be compared there too, and their roles swap.
In the world where no ground holds, a judge correctly finds no liability only by being right about every ground — chance again — and otherwise finds liability on whichever ground they got wrong. Each judge errs on a different ground, but the verdict vote counts only their conclusions, so a majority of “liable” verdicts can be assembled from judges who each found liability for a different reason, none of which a majority accepts. At four-in-five competence and four grounds, that happens nine times in ten with fifty-one judges. The ground-by-ground vote asks whether a majority accepts any single ground, and almost never errs.
So the verdict vote is cautious with conjunctions and reckless with disjunctions, and both for the same reason: it lets each judge combine their own errors before the court aggregates. The logical form of the law, not the quality of the judges, decides which way the error runs. A procedure that is safe for one kind of legal test is dangerous for the other, and no rule escapes the dilemma in general — but in these two common forms the premise vote is the more truthful by a wide margin.
How large a premise-voting court must be
The premise vote has a cost the verdict vote does not: every premise is a separate majority that must come out right, and more premises are more chances for one of them to fail.
The cost is small. A majority’s chance of error falls exponentially with the size of the court, so keeping majorities all right at once needs the court to grow only like the logarithm of : at competence 0.7, from seventeen judges for one premise to forty-one for twelve. At competence 0.6 the numbers are larger — a judge barely better than a coin needs many colleagues — but the growth is the same slow kind. The verdict vote, past its threshold, cannot be rescued by any court size at all.
Equal judges and independent errors, assumed throughout
The probabilities in the threshold, liable-world, average, disjunction and court-size figures are exact binomial sums, and the two-premise, three-judge crossover is checked against the value found in the earlier comparison, . The disagreement figure is a simulation of twenty thousand seeded cases at each , and its shares among liable cases rest on only a few hundred liable cases at the larger — enough for the pattern, not for the second digit.
All of it rests on the model: judges of equal competence, the same on every premise, erring independently of each other and across premises. Real judges are better on some elements than others, read the same briefs, and influence each other’s views, and each of those changes the numbers. Correlated errors in particular weaken the jury theorem’s amplification for both procedures; whether they weaken one more than the other depends on whether the correlation is between judges or between premises, and the figures do not separate those cases.
Still open: deliberation, and premises that depend on each other
The model treats premises as independent questions. In real cases they are often entangled — whether a duty existed and whether it was breached may turn on the same facts — and judgements about them are correlated within each judge. Christian List and Franz Dietrich have studied aggregation over logically connected propositions, and which agendas can be voted on without contradiction is understood; how truth-tracking behaves when the premises themselves are statistically dependent is understood only in special cases.
The other open edge is deliberation. Judges talk before they vote, and a court that discusses the premises before voting on the verdict may behave like neither procedure modelled here. Whether discussion moves a court towards the premise vote’s accuracy or towards conformity — everyone adopting the most persuasive judge’s errors — is an empirical question as much as a mathematical one, and the models that treat it give answers that depend on assumptions about how judges update.
A third question is how unequal competence changes the threshold. If judges are better on some elements than others — reliable on the law, less so on contested facts — the verdict vote’s per-judge chance is the product of different numbers, and the weakest element dominates it. The premise vote suffers less, since each element gets its own majority, but it still needs every one of those majorities to be right, and a single element on which judges are barely better than chance drags its accuracy down too. Which procedure degrades more gracefully as competence becomes uneven across elements, and whether weighting the elements differently could restore the comparison, has been studied only for small numbers of elements.
Where aggregation happens
Every premise added to a conjunctive legal test raises the competence the verdict vote needs, to , and below that threshold a larger court is worse than a smaller one. The premise vote needs only better-than-even judges at every , and a court that grows only logarithmically. The difference is when each procedure aggregates: before combining the premises, or after. With disjunctions the same difference produces the opposite failure, and the verdict vote convicts defendants that no majority thinks liable on any ground.
The jury theorem is often read as a reason to trust large groups. Read carefully, it is a reason to be careful about what a group is asked: the amplification that makes a large majority reliable on a simple question makes it reliably wrong on a compound one, if each member has to answer the compound question alone.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The court that contradicts itself — both name aggregation, proposition
- What the agenda leaves standing — both name judgement aggregation, majority rule
Named objects
A dashed tag is an object no other essay names yet.
AggregationCondorcet jury theoremJudgement aggregationMajority ruleProbabilityProposition