Applied

The court that contradicts itself

Three judges each answer three questions, and each answers them consistently. Take the majority on each question separately and the answers no longer hang together — the body as a whole asserts a combination no member of it holds, and no rearrangement of the procedure removes the problem.

Worth reading first: The majority that goes in a circle · A formula is a corner of a cube.

A panel of three has to settle a case that turns on two questions. Was there a valid contract? Was it breached? A finding for the plaintiff requires a yes to both, and each judge answers the two questions and states the verdict their answers force.

3 consistent judges, and a majority that is not. A table of judges against three questions, every judge's row internally consistent, with the majority answer to each question underneath forming a combination no judge holds.
Fig. 1 Three judges, each answering three questions and each answering them consistently. The majority says yes to both premises and no to the verdict — which is a position no judge holds and no judge could hold, because the third answer is forced by the first two.

The first judge finds a contract and a breach, so finds for the plaintiff. The second finds a contract and no breach, so finds against. The third finds no contract and — being obliged to answer anyway — no breach, so also finds against.

Two of the three found a contract. Two of the three found a breach. And two of the three found against the plaintiff. Every judge is consistent, the majority is not, and nothing has gone wrong except that the votes were counted question by question.

Why this is not the voting paradox

The obvious comparison is the majority that goes in a circle, where three voters ranking three candidates can produce a majority preferring A to B, B to C and C to A. Both are failures of majority rule, and they are not the same failure.

A majority cycle over 3 candidates, and how often 3 voters produce one. The majority tournament as a directed polygon with each arc's margin, beside one cell for every profile of the stated size, filled where no Condorcet winner exists.
Fig. 2 The older paradox for comparison: a majority tournament with no top, found by counting the pairwise margins over the whole profile space. Twelve of the two hundred and sixteen profiles of three voters over three candidates have no Condorcet winner.

The voting cycle is about transitivity — the majority’s preferences fail to line up in an order. The paradox here is about logic — the majority’s answers fail to satisfy an implication that binds them. The two constraints are different in kind, and the second is the more general: a set of propositions can be connected by any logical relation at all, and the cycle is what happens when the relation is the transitivity of a ranking.

That generality is why the subject has a separate name. Judgement aggregation asks what happens when a group has to take a position on a set of propositions that logic ties together, and the answer is that almost nothing works.

How often it happens

The agenda here — two premises and the conclusion they force — admits exactly four consistent positions. A judge who says yes to both premises must say yes to the conclusion; a judge who says no to either must say no; and that is four combinations out of eight.

6 of 64 profiles leave the majority contradicting itself. A bar showing what fraction of all judgement profiles produce a majority whose answers do not hang together, with the four consistent positions listed above it.
Fig. 3 Every one of the sixty-four ways three judges could each take one of the four consistent positions. Six of them leave the majority asserting both premises and denying the conclusion, and that is the only inconsistent shape available.

With three judges there are sixty-four profiles, and six of them break. That is under a tenth, which sounds reassuring and is not: the six are exactly the profiles in which the three judges take three different positions, which is to say the profiles in which the panel is genuinely divided. The paradox happens precisely when there is a disagreement worth having.

150 of 1024 profiles leave the majority contradicting itself. A bar showing what fraction of all judgement profiles produce a majority whose answers do not hang together, with the four consistent positions listed above it.
Fig. 4 The same census with five judges: a hundred and fifty of the one thousand and twenty-four profiles break. The share does not shrink as the panel grows.

Adding judges does not help. It cannot, because the mechanism does not depend on the number: as long as the three positions “yes to both”, “yes to the first only” and “yes to the second only” are each held by a large enough minority, the two premise-majorities can be assembled from different pairs while the conclusion-majority is assembled from neither.

The shape the failure always takes

There is one detail in the census worth pulling out, because it is the whole mechanism in one sentence.

Every inconsistent majority on this agenda looks the same: yes to both premises, no to the conclusion. Never the reverse. That is not an accident of the example and the figure checks it — across all sixty-four profiles, the six that break all break in that direction.

The reason is that the conclusion is a conjunction. To get a majority for it, more than half the panel must say yes to both premises; to get majorities for both premises separately, it is enough that more than half say yes to each, and those two halves need not overlap. So the conclusion is harder to carry than the premises are, and the gap only ever opens one way.

Change the connective and the direction changes with it. If the verdict required a yes to either premise, the disjunction would be easier to carry than the premises, and the inconsistent majorities would be the ones saying no to both premises and yes to the verdict. The paradox is a fact about how majorities interact with logical connectives, not about courts.

5 consistent judges, and a majority that is not. A table of judges against three questions, every judge's row internally consistent, with the majority answer to each question underneath forming a combination no judge holds.
Fig. 5 Five judges rather than three, arranged the same way. Three of the five find a contract, three find a breach, and only one finds for the plaintiff — the two premise-majorities are made of different people, and the panel size changes nothing.

Two ways out, and they disagree

The panel has to do something, and there are two obvious things to do.

Two ways out of the paradox, and they disagree. The premise-based and conclusion-based readings of one judgement profile shown as two rows, with the verdict each produces differing.
Fig. 6 The same three judges read twice. Taking the majority on each premise and deriving the verdict finds for the plaintiff; taking the majority on the verdict itself finds against; and nothing in the profile says which is right.

Premise-based: vote on the two premises, then apply the law to whatever the votes produced. This finds a contract and a breach, and therefore finds for the plaintiff — even though two of the three judges would find against.

Conclusion-based: vote on the verdict directly and let the premises fall where they may. This finds against the plaintiff, and the panel then issues a verdict whose stated grounds are held by a minority, or issues no grounds at all.

Both are consistent. Both are defensible. They give opposite answers on the same evidence from the same three judges, and the choice between them is not a mathematical question — it is a question about whether a court is in the business of settling disputes or of stating law, and reasonable people have taken both sides.

What the mathematics does say is that this is not a defect of these two procedures in particular. Any procedure that decides each proposition by looking only at the votes on that proposition, and always returns a consistent position, must be a dictatorship — the position of one fixed judge, whatever the others say. That is List and Pettit’s theorem, from 2002, and it is the exact analogue of Arrow’s with rankings replaced by judgements.

Counting the profiles that break

The census is worth reading closely, because the number that comes out of it is smaller than the phenomenon deserves and the reason is instructive.

Six profiles of sixty-four break, which is nine per cent. But the sixty-four count every assignment of the four consistent positions to three labelled judges, including the twenty-eight in which two or more judges agree completely. A panel whose members agree is not a panel with a problem.

Restrict to the profiles in which all three judges take different positions and there are twenty-four of them — four choices for the first, three for the second, two for the third. Six of those twenty-four break, which is a quarter. And every one of the six is a rearrangement of the same three positions: yes-to-both, yes-to-the-first-only, and yes-to-the-second-only. There are 3!=63! = 6 ways to hand those three positions to three judges, and all six produce the paradox.

So the correct reading is not “nine per cent of the time” but “always, when the disagreement takes this particular shape”. Which shape is a fact about the agenda, and the shape is available whenever three positions exist whose premise-answers can be assembled into two majorities that do not overlap.

The conditions, and which one to give up

The theorem’s conditions are worth listing, because a reader’s instinct on reading the paradox is that some obvious fix has been overlooked, and the list is what rules the fixes out.

Universal domain: every combination of consistent individual positions must be handled. Restricting the panel to profiles that happen not to break is a real proposal and it means telling judges what they may think.

Consistency: the group’s position must be logically coherent. Giving this up means a body that asserts a contradiction, which is what the majority already does. What coherent means is fixed by the agenda’s own logic, in exactly the way a set of yes-or-no answers is constrained by the connectives tying them — the group is not free to pick a corner of the cube that the implication rules out.

Independence: the group’s answer on a proposition depends only on the individual answers to that proposition. This is the condition the premise-based procedure keeps and the conclusion-based procedure breaks — the conclusion is derived rather than voted on, so its group answer depends on votes about other propositions.

Anonymity or non-dictatorship: no single judge decides. Giving this up is the theorem’s conclusion rather than an option.

Independence is the one that goes, and it is worth noticing that it is the same condition that goes in Arrow’s theorem, where it is called independence of irrelevant alternatives. In both cases the condition is a demand that the procedure be local — decide each question on its own evidence — and in both cases locality is incompatible with global coherence.

A profile of 100 ranked ballots, and the majority in every pair. The voter groups as columns with the ranking down each, beside the pairwise majority matrix whose cells are the margins.
Fig. 7 The other tradition’s data for comparison: a profile of ranked ballots and the margin in every pair. A ranking is a set of judgements about pairs, subject to the constraint that they be transitive — which is why the older paradox is a special case of the newer one.

What the two procedures are answers to

The choice between the two is not a preference and it is not a coin toss. It is a question about which output the aggregation is supposed to produce, and separating the two outputs makes each procedure’s case visible.

One output is a decision about this instance: which way the verdict goes, and nothing else. If that is what is wanted, voting on it directly is the procedure that produces it, and the answers to the premises are commentary.

The other output is a rule for the next instance: a set of positions on the premises that later cases can be measured against. If that is what is wanted, voting on the premises is the procedure that produces it, and the verdict is derived.

The mathematics contributes exactly one thing here, and it is worth being precise about how small it is. It does not say which output should be wanted. It says the two cannot be produced at once — the position a panel would vote for and the position its collective answers imply are different things, and no amount of procedural ingenuity makes them the same.

Where it fails, and what it needs

The agenda has to be connected. If the propositions are logically independent — no implication among them — then majority voting on each is perfectly consistent, because there is nothing for it to be inconsistent with. The paradox needs the conclusion to be forced by the premises, and the theorem’s hypotheses include a technical condition on the agenda’s connectedness that is exactly this.

An even panel ties. Everything here assumes an odd number of judges so that no proposition ties. With an even number the majority may fail to exist rather than fail to be consistent, which is a different problem with different fixes.

Abstention changes it. A judge who declines to answer one question is not a consistent position in this framework, and allowing abstention opens up procedures the theorem does not cover — at the cost that the group’s position may now be incomplete rather than inconsistent.

And the escape routes are choices, not solutions. The premise-based and conclusion-based procedures are both consistent and both defensible, and the theorem’s content is that no third option avoids the choice. A reader looking for the procedure that is simply right will not find one here, and the field’s honest position is that there is not one.

Where it came from

The paradox was noticed by the legal scholars Kornhauser and Sager in 1986, writing about appellate courts and what they called the doctrinal paradox. Their interest was practical: American appellate panels do issue verdicts with reasoning supported by a minority of the judges, and they wanted to know whether that was a scandal or a structural feature.

The mathematical version is List and Pettit’s, from 2002, and the generalisation is theirs: replace the two premises and a conjunction with any set of propositions closed under negation and any logic connecting them, and the impossibility survives. The subject has grown considerably since, largely by classifying which agendas — which patterns of logical connection — admit a consistent aggregation rule and which do not.

There is a strategic half to all of this that the theorem does not touch, and it is worth naming so it is not read in. Nothing above concerns a judge who misreports a position in order to steer the verdict; that is a separate impossibility with its own theorem, of the shape a profitable misreport located by exhausting every misreport rather than a failure of coherence. A panel of perfectly sincere judges produces the paradox on its own.

It is worth recording that the result was found twice, from opposite ends. Guilbaud had written essentially the same observation in 1952, in a paper on the logical problem of aggregation that read Condorcet’s paradox as a special case of exactly the phenomenon above. It was forgotten for fifty years, which happens to results whose natural audience does not yet exist.

What the pictures cannot show

The agenda drawn here has three propositions because a table of three columns is legible. The theorem is about agendas of any size, and the condition it needs on them — that they contain a minimally inconsistent subset of at least three propositions — is a statement about the whole agenda that no single table displays.

The census counts profiles and treats each as equally likely, which is a modelling choice with no justification beyond making the count well defined. Real panels are correlated, and a share of six in sixty-four says nothing about how often a court runs into this.

And no figure here proves the impossibility theorem. The figures exhibit the paradox, count how often it occurs on one small agenda, and draw the two escapes; the statement that every independent consistent rule is a dictatorship is a claim about infinitely many rules, and it is quoted rather than performed.

The ladder from here

Below: the majority that goes in a circle, which is this paradox in the special case where the propositions are pairwise comparisons, and a formula is a corner of a cube, where a set of yes-or-no answers is a point and consistency is a constraint on which points are allowed. Sideways: four conditions and no rule, the same impossibility for rankings, and a lie that pays, where the failure is strategic rather than logical. Above: the agenda characterisation theorems, distance-based aggregation rules, and the connection to belief merging.

What is worth carrying away

Majority rule is a procedure for one question. Applied to several questions that logic ties together, it produces an answer that is not a position anybody could hold, and the failure is not a rounding error — it is what happens whenever different majorities are assembled from different people.

The general shape is worth carrying past this example. A rule that treats each part of a decision separately cannot be relied on to respect a constraint that binds the parts together, because the rule never looks at the constraint. Anything that must be globally coherent has to be decided globally, and deciding globally means giving up the locality that made the procedure appealing in the first place.