The rule with no favourites
Worth reading first: Five rules and one dial · The seat that vanishes when the house grows.
Turning the dial moves seats from small regions to large ones, and that much is visible on any single instance. What a single instance cannot say is whether one setting of the dial is fair in the sense of favouring neither — and that is a question about many instances, so it has to be measured.
Every seat count behind those bars is exact. What is averaged is seats minus exact quota, which for one region on one instance is a rational number between and , and over four hundred instances it settles down.
What bias is, and what it is not
The word has an ordinary meaning that has to be set aside first. A method is not biased because it produces an answer somebody dislikes, and it is not biased because large regions get more seats than small ones — they should, and every method gives them more. Bias here is a technical quantity with a definition, and the definition is a comparison against the exact share.
A method is biased toward large regions when, averaged over instances, large regions get more than their exact quotas and small regions get less. That is a statement about the average, and three things it is not are worth stating first.
It is not a statement about any single instance. On one census a method may give the largest region less than its quota and the smallest more; the bars measure a tendency.
It is not the same as breaking the quota rule. A method can be badly biased and never leave the quota band — being consistently at the top of one’s band and consistently at the bottom of somebody else’s is a bias of nearly a seat with no violation anywhere.
And it is not a statement about the number of seats a region has. Every method gives large regions more seats than small ones; the question is whether it gives them more than their share.
Why the direction is a theorem
The numbers in the figure are measurements. Their signs are not — the signs follow from the signpost, and the argument takes a paragraph.
Consider a region whose exact quota is . Under the divisor method with signpost , it receives seats where and is its rescaled share. Averaged over instances, the rescaling factor adjusts so the total is right; what matters is that the region rounds up when its fractional part exceeds and rounds down otherwise. With it always rounds down.
Now the asymmetry. A region with a large quota has more room to lose in the rescaling than a region with a small one. Under Jefferson, every region rounds down, so the divisor must be pushed below the natural one to make up the shortfall — and pushing the divisor down adds seats in proportion to population, so it adds more to the large regions than to the small ones. Under Adams, every region rounds up, the divisor is pushed the other way, and the correction is again proportional and again favours the large — but the rounding it corrects was a full seat for every region, which costs the small ones proportionally far more.
The precise statement, and it is Balinski and Young’s: among the divisor methods, Webster is the only one that is unbiased. The others are biased in the direction their signpost sits, and the size of the bias depends on the distribution of populations rather than being a fixed number.
What the numbers actually say
Reading the bars as quantities rather than as directions is worth doing, because their sizes are not what a reader expects.
Jefferson: +0.35 for the largest region, −0.32 for the smallest. A third of a seat, on average, in a house of twenty-seven. On the instance the rung below uses, that shows up as the largest region getting seventeen seats against an exact quota of 15.417 — a full seat and a half above its share, and a violation of the quota rule as well as a bias.
Adams: −0.34 and +0.32. Almost exactly the mirror image, which is what the symmetry of the dial predicts.
Webster: +0.01 and +0.01. Both averages are a hundredth of a seat and both are positive, which is a good illustration of what a measurement of an unbiased quantity looks like: not zero, but indistinguishable from zero at this sample size.
Hill: −0.06 and +0.14. Small, and pointing the same way as Adams. Hill’s signpost is a geometric mean, which sits below a half at small numbers, so it protects small regions — less than Adams does, and measurably.
Hamilton: 0.00 and +0.01. The nearest thing to zero on the chart. Hamilton’s method is not a divisor method at all and is unbiased for a different reason: it gives every region the floor or ceiling of its own quota, so its deviation is bounded by one seat by construction and averages out. Nothing about that argument mentions size, which is why the bound holds for the largest region and the smallest alike — and it is the only guarantee on this page that does not need a family of instances to state.
That last row is the one to be careful with. Hamilton is unbiased and it is the method that loses a seat when the house grows, so an unbiased average is not a recommendation. Being unbiased on average and being well-behaved are different properties, and a rule can be flawless on the criterion it was designed for and absurd on one nobody thought to state.
The quantity being averaged
One more definitional point, because the choice of quantity is doing work.
What is averaged is : seats awarded minus exact quota, in seats. An obvious alternative is the relative error , in which a region entitled to one seat and given two counts as badly served as a region entitled to ten and given twenty.
The two choices give different rankings, and they are answering different questions. Absolute deviation asks how many seats went astray; relative deviation asks how well was each region served. A method unbiased in one sense need not be unbiased in the other, and the disagreement between them is exactly the disagreement between Webster and Hill that the section below is about.
The figures here use the absolute one, and say so. Every claim on this page is a claim about it.
The instances, and why they are stated
A measurement over generated instances is only as good as the account of how they were generated, so here it is in full.
Four hundred instances. Five regions each. Populations drawn as whole numbers between one hundred and six thousand, from one seeded generator, so the four hundred instances are the same four hundred every time the figure is drawn. Instances in which two regions tie for largest or for smallest are discarded, because the largest region then names two things and the statistic stops meaning anything. House size twenty-seven throughout.
Two honest limitations follow from that description.
The distribution is uniform and populations in practice are not. Actual regions have a distribution with a long tail — a few very large, many small — and bias measured on such a distribution is larger for every method, because there is more disparity for the rounding to act on — the same sensitivity to the shape of the input that makes a fair division depend on what the pieces are worth to whom. The direction of each bias would not change; the sizes would.
Five regions is few. With more regions there are more small ones for a bias against small regions to act on, and the averages grow. The figure is a demonstration that the effect exists and points where the theory says, not an estimate of its size in any particular application.
Why Webster was not chosen
The unbiasedness of Webster’s method has been known since the 1920s and it is not the method that assemblies have generally adopted. The reason is worth knowing because it is a good example of a technical argument losing to a different technical argument.
Hill’s method — the geometric mean — was recommended in 1929 by a committee of mathematicians that framed the question differently. Rather than asking which method has no average bias, it asked which method minimises the relative difference in representation between any two regions, and answered Hill. Both are defensible criteria, they give different answers, and the committee’s choice was made on the second.
Whether the two criteria can be reconciled is the subject of the rung above but one. What can be said here is that they are genuinely different questions: unbiasedness is a property of a method averaged over many instances, and Huntington’s stability is a property of a single apportionment, and there is no reason for the answers to coincide.
The pattern is the field’s, and it recurs whenever a criterion is chosen after the candidates are known. Four conditions on a voting rule are individually reasonable and jointly impossible; which one to drop is decided by argument rather than by proof, and the argument is made by people with a view about the answer. Choosing the criterion is choosing the winner, and doing it in that order is not dishonest — it is unavoidable, because a criterion nobody has tested against candidates is a criterion nobody knows the consequences of.
There is also a blunter reading, and honesty requires including it. Hill’s method gives small regions slightly more than Webster’s, and the choice was made by an assembly containing small regions. That the two arguments pointed in the same direction as the interests does not make the mathematical argument wrong; it does mean the mathematical argument was not the only thing being weighed, and no analysis in this subject has ever been conducted anywhere else. A rule for choosing is chosen by somebody, and the somebody has a stake.
How large a bias can get
A third of a seat sounds small. Scaling the question is worth a paragraph, because the size of a bias is not a fixed property of a method.
The bias acts on the disparity between regions. Five regions all of similar size give a rounding rule almost nothing to work with, and the averages collapse toward zero for every method. A hundred regions spanning three orders of magnitude give it a great deal, and the accumulated advantage under Jefferson can run to several seats for the largest region.
The house size matters in the other direction. As the house grows, every region’s quota grows, the fractional parts become a smaller fraction of the whole, and the relative bias falls — while the absolute number of seats involved rises. A method that is a third of a seat biased at twenty-seven seats is not a third of a seat biased at four hundred and thirty-five; it is a larger number of seats and a smaller fraction of each region’s entitlement.
So there is no single number to quote, and the honest statement is the one the theorem makes: Webster is the setting at which the average is zero for every distribution and every house size, and every other setting has an average whose sign is fixed and whose size is whatever the instances make it.
What the pictures cannot show
Four hundred instances is a sample. The averages have a spread, and the figure prints two decimal places without an error bar. Webster’s 0.01 is consistent with zero; it is not a measurement of zero.
The generator is stated and it is not reality. Uniform populations between one hundred and six thousand are a family chosen for being describable, not for resembling anything. Every number here moves under a different family, and only the signs are stable.
The theorem is not measured. That Webster is the unique unbiased divisor method is a statement about all instances and all distributions of a certain kind, and four hundred instances of one distribution cannot establish it. The figure shows the theorem’s prediction being met.
The two bars per method are not independent. Seats taken from the smallest region go somewhere, and on a five-region instance they usually go to the largest, so the two averages move together by construction. Reading them as two measurements is reading one measurement twice, and the symmetry of the Jefferson and Adams rows is partly that.
And bias is one criterion among several. The bars rank the five methods on a single axis, and the next two rungs show that ranking them on other axes gives different orders. A figure that ranks is always inviting a reader to stop reading.
Where the ladder goes next
The rung above is the impossibility. Three properties — staying within quota, never taking a seat away when the house grows, never taking one from a region that grew faster — and no method with all three. The five methods here have at most two each, and Balinski and Young proved that no rule anywhere does better.
Above that is Huntington’s account, which is the one that makes the choice between five rules that give five answers a choice about the question rather than about the rule: each of the five methods is the one that no transfer of a seat can improve, for a particular measure of inequality between two regions. That turns which method is right into which measure of unfairness is meant, which is the closest the subject comes to an answer.
What is worth carrying away
A property that has to be averaged over instances is a different kind of property from one that can be checked on an instance, and the two are easy to confuse when both are called fair.
Bias is the first kind. No single apportionment is biased; a method is, and only over a family of instances. That makes it measurable rather than checkable, and it makes the family part of the claim — a point the pigeonhole principle never has to make, because a statement true of every instance needs no family at all — which is why the description of the four hundred instances is as much of this essay as the numbers are. A measurement without its family is a number with no theorem attached, and a reader who takes 0.35 away from this page without taking away five regions, twenty-seven seats, uniform populations has taken away the wrong thing.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Choosing what unfair means — both name apportionment, divisor method, geometric mean, quota
- Two out of three, and never all three — both name apportionment, divisor method, monotonicity, quota
- Sampling where the answer lives — both name expectation, sampling
Named objects
A dashed tag is an object no other essay names yet.
ApportionmentBiasDivisor methodExpectationGeometric meanMonotonicityQuotaRoundingSampling