An average that never settles
Worth reading first: The shape that averaging leaves alone · The average settles and the wobble does not.
The law of large numbers says that the average of independent draws settles on the mean, and the limit theorem says how it wobbles on the way. Both have a hypothesis that is usually left unread, and this rung is what the theorems look like when it is dropped.
The left panel is not a picture of slow convergence. Those runs do not settle at any point in their history, and running them for a million draws produces the same picture at a different scale: long plateaus interrupted by jumps, with the size of the jumps growing in proportion to the number of draws so far.
What the distribution is
The Cauchy distribution has density
which is a perfectly respectable bump: symmetric about zero, with a single peak, falling away on both sides. Nothing about its shape near the middle suggests trouble.
The trouble is in the tails, and it is entirely a matter of the rate. The density falls off like , and therefore falls off like , whose integral diverges. So the mean does not exist — not is infinite, but does not exist, since the integral over the positive half and the integral over the negative half are separately infinite and their difference is not defined. With no mean there is no variance either, and the hypothesis of both theorems above is gone.
There is a construction that makes the distribution feel less exotic. Stand at distance one from an infinite straight wall and point a beam at a uniformly random angle; where it lands is Cauchy distributed. The tangent of a uniform angle is exactly this density, and the heavy tail is the fact that angles near a right angle throw the beam enormously far along the wall, while occupying a perfectly ordinary share of the angles.
That construction also explains why no amount of averaging helps. Averaging the landing points of many beams is not averaging anything the geometry cares about; the occasional near-right-angle beam lands so far away that it dominates every measurement taken so far.
Stable under its own average
The identity that makes the right-hand panel flat is worth stating precisely, because it is stronger than the failure of convergence.
If are independent Cauchy quantities, then
has exactly the Cauchy distribution — not approximately, and not in the limit. The average of a thousand draws is distributed like one draw. The rescaling that would be needed for a bell curve is ; here the sum needs dividing by to stay put, and after that division nothing has changed at all.
The figure checks this in two ways. It samples twenty thousand averages at each of three sizes and compares their cumulative distributions with the exact one, requiring the largest gap to be under two hundredths. And it does the convolution integral directly at nine points: the density of is computed by integrating the product of two Cauchy densities, and compared against a single Cauchy density.
The mechanism, in the language of the previous rung, is that the Cauchy is a fixed point of convolution-and-rescaling — with the rescaling exponent rather than . It is one of the stable laws, the family of fixed points that appears once finite variance is not required, and it is the only one besides the bell curve with a density anybody writes down.
What a light tail looks like, for contrast
Set beside the ordinary case, the failure becomes legible.
The die’s average settles because its variance is finite and the variance of the average is that number divided by . Chebyshev’s inequality then bounds the chance of being far from the mean, and the bound goes to zero.
Every step of that reasoning fails for the Cauchy, and it fails at the first line rather than the last. There is no mean to settle on, no variance to divide, and no inequality to apply. It is not that the convergence is slow; there is nothing for it to converge to.
The largest term carries the sum
The most useful intuition about heavy tails is a statement about the maximum, and it is what makes the plateaus in the first figure inevitable.
For a Cauchy sample of size , the largest draw is typically of order — because the chance of exceeding is about , so the largest of is around . The sum is also of order . So a single term is a constant fraction of the whole sum, at every sample size, and the average is dominated by whichever draw happened to be biggest.
That is exactly what the running averages show. A run sits on a plateau while nothing extreme arrives; then a single draw of order arrives and moves the average by an amount comparable to itself divided by , which is order one. The jump does not shrink as the run gets longer, because both the size of the extremes and the amount of averaging grow together.
Under a light tail the opposite holds: the largest of draws grows like for a normal quantity, which is negligible beside the sum’s , and no individual term matters. The difference between those two regimes is the whole practical content of a tail condition, and it is visible without any theorem: ask whether the biggest of a sample is comparable to the total.
A heavy tail already in this collection
The Cauchy can look like an invented counterexample, and it is worth pointing at a heavy tail that arose here without anybody choosing it.
The first return of a random walk happens with certainty and takes, on average, forever. The tail exponent is , which puts it below the Cauchy in the stable family, and the consequences are correspondingly worse: the total time for successive returns grows like rather than like , so the average of return times grows without bound instead of settling anywhere.
That example is instructive because nothing about the setup is unusual. A fair coin, a step each way, and the wait for a first return — and the resulting distribution breaks the law of large numbers more thoroughly than the Cauchy does. Heavy tails are not a curiosity imported from elsewhere; they arise from the plainest possible mechanism, whenever a quantity is a waiting time for something with a chance of taking a very long while.
The median, which does settle
The failure is specific to the average rather than general to the distribution, and the cleanest way to see that is to take a different summary of the same sample.
The median of independent Cauchy draws concentrates on zero as grows, with a spread of about — the ordinary square-root rate. Nothing about the heavy tail interferes, because the median depends only on the ordering of the sample and not on the sizes, and a draw of enormous magnitude moves the ordering by one place regardless of how enormous it is.
So the sample carries perfectly good information about where the distribution is centred, and the average is simply the wrong function of it. The average has the property that a single term can move it arbitrarily far; the median has the property that no single term can move it more than one position. Under a light tail the first property costs nothing and the average is the better-behaved of the two; under a heavy tail it is fatal.
That is a fact about the two functions rather than about statistics, and it is worth separating from any question of what should be done with data. A function of many variables that is insensitive to any one of them is a different kind of object from one that is not, and heavy tails are the setting where the difference becomes visible.
Where it came from
Poisson wrote the density down in 1824 and noticed that its mean does not exist. The name attached to Cauchy after an exchange in 1853 with Bienaymé, in which Cauchy produced the distribution precisely as a counterexample — Bienaymé had argued for the method of least squares on grounds that presupposed finite variances, and Cauchy’s reply was a distribution for which the argument’s conclusion is false.
The episode is a good illustration of what counterexamples are for. Neither man was wrong about the mathematics: Bienaymé’s argument is correct under its hypotheses, and Cauchy’s example shows that the hypotheses are load-bearing rather than decorative. What the exchange established is the boundary of a method, which is a more durable contribution than another theorem inside it.
The distribution’s later career has been almost entirely as a counterexample, and it is the standard first port of call when checking whether a result in probability quietly assumes a moment. If a claim survives the Cauchy, it survives most things.
Where the boundary lies
The stable laws are indexed by an exponent between and , with tails falling off like , and the exponent decides everything.
- is the bell curve: finite variance, sums rescale by , the limit theorem holds.
- : the mean exists, the variance does not. Averages converge to the mean, so the law of large numbers holds — and the fluctuation around it is not a bell curve, and is of size rather than .
- is the Cauchy: neither theorem holds.
- : not even the average converges; the sum grows faster than .
The middle band is the one that catches people out, because a sample from it produces a well-behaved-looking average with wildly misbehaved fluctuations, and the misbehaviour appears only in the rare large draws. A histogram of such a sample looks entirely ordinary until it does not.
What survives the failure
It would be wrong to conclude that nothing can be said about sums of Cauchy quantities. A great deal can, and all of it is exact rather than asymptotic.
The sum of independent Cauchy quantities has, precisely, the Cauchy distribution scaled by . So every question about the sum has an exact answer at every : the chance that the sum of a thousand draws exceeds ten thousand is the chance that a single draw exceeds ten, which is about . No limit theorem is needed, and none would help.
That is the compensation the stable laws offer. A distribution that is a fixed point of its own convolution is one for which sums require no approximation at all, and the exactness holds at every sample size rather than in a limit. The bell curve has the same property and it is usually invoked the other way round — as the limit of things that are not normal, rather than as the shape whose own sums are trivial.
So the situation is not that the Cauchy is intractable. It is that the tool everybody reaches for is the wrong one, and the right one — the exact stability relation — is simpler than the tool it replaces.
What the picture cannot show
Every drawn axis is finite, and the whole content of a heavy tail is what happens outside any finite window. The cumulative distributions in the first figure are drawn from to , which contains about of the mass of a Cauchy; the remaining is spread over the entire rest of the line, and it is where the mean fails to exist.
The running averages are worse. Five runs of fifteen hundred draws each contain a handful of genuinely large values, and the plateaus and jumps are a fair picture of what happened; but a run that showed the phenomenon at its full strength would need a vertical axis a thousand times taller and would then show nothing at all in the middle. The figure clamps its runs to a visible band, and the clamping is a choice that hides exactly the events the essay is about.
Nor can any picture distinguish no mean from a mean that has not been found yet. That is a statement about an integral, and the only honest way to establish it is to do the integral.
The ladder from here
Below: the bell curve from coin flips and the two scalings. Sideways: the shape averaging leaves alone, of which this is the other famous fixed point, and Chebyshev’s inequality, whose hypothesis is the one that fails here. Above: how fast the bell arrives when it does arrive, and the walk that becomes a curve, whose limit is continuous precisely because its steps have a finite variance — with heavy-tailed steps the limiting path jumps instead.
How to notice it in advance
Two checks distinguish a heavy tail from a light one, and neither requires knowing the distribution.
The first is the one described above: compare the largest value in a sample with the total. If the biggest of a thousand draws is a noticeable fraction of their sum, no average of that quantity will settle, and taking more draws will not repair it.
The second is to watch the running variance rather than the running mean. For a light tail it converges like everything else; for a tail heavy enough to break the limit theorem it grows without bound, and it does so visibly within a few thousand draws even when the mean looks well behaved.
Both are properties of a sample rather than theorems, and neither settles the question — a distribution with a finite but enormous variance behaves exactly like one with none over any finite sample. What they give is a warning, which is what is wanted.
A hypothesis worth reading
The lasting point is that the failure is not exotic. The Cauchy is a symmetric bump with a peak in the middle and no strange features anywhere in the range a picture can show; it fails the theorems because of what its density does far away, and that is invisible in every sample of ordinary size.
Theorems in probability are unusually rich in hypotheses of this kind — conditions on a tail, a moment or an integral, none of which can be checked by looking at a picture of the data. Independence is the other one, and it fails just as invisibly. The habit worth having is to treat those conditions as the substance of the theorem rather than its small print, because the conclusion is not approximately true when they are approximately satisfied — it is simply false.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A constant that does not care which map — both name convergence, scaling
- A curve with a corner at every point — both name convergence, counterexample
- Getting pi by dropping needles on the floor — both name convergence, independence
- The staircase that is not the diagonal — both name convergence, counterexample
Named objects
A dashed tag is an object no other essay names yet.
ConvergenceCounterexampleExpectationHeavy tailsIndependenceNormal distributionScalingVariance