Concept

Heavy tails

A distribution whose extreme values are common enough to dominate averages, and whose variance or mean may not exist at all. The largest of a sample is then comparable to the whole sum, which is why averaging such quantities buys so little.

Named by 7 essays across one field — each of them below, with the objects they name alongside it.

A Galton board after 600 balls. 600 balls fall through 12 rows of pegs, each bouncing left or right at random, and pile up in a bell-shaped heap.

A bell curve assembled out of coin flips

Drop six hundred balls through a board of pegs, each bouncing left or right at random, and they pile up in a shape that can be predicted precisely. Nothing coordinated them.

probability · Central limit
five values, unevenly weighted, and the mass outside 3 standard deviations. A distribution drawn as bars, with the windows one and a half, two and three standard deviations wide marked. The probability outside each window is summed and compared with the bound that knows only the variance.

How far from the average a thing can be

Knowing only an average and a spread — nothing about the shape, nothing about the number of outcomes, nothing about symmetry — the chance of landing three standard deviations out is at most one in nine. And there is a distribution that lands there exactly that often, so the bound cannot be improved.

probability · Concentration
Averages of a heavy-tailed quantity, which never settle. Running averages of draws from a Cauchy distribution, which jump rather than converge, beside the cumulative distributions of averages of 1, 4 and 16 draws, which lie on top of one another.

An average that never settles

The average of many independent quantities is supposed to steady as their number grows. For one famous distribution it does not steady at all — the average of a thousand draws has exactly the same distribution as a single draw, and no amount of further averaging changes it.

probability · Central limit
How fast a sum becomes a bell curve. The largest gap between the distribution of a standardised sum and the bell curve, against the number of terms, on logarithmic axes. Both summands fall along a line of slope about minus a half.

How fast the bell arrives

The limit theorem says a standardised sum approaches the bell curve and says nothing about when. The rate is one over the square root of the number of terms, the constant in front is made of the third moment, and both are visible.

probability · Central limit
The chance the average clears 0.75, against the number of draws. The exact probability that the average of n draws exceeds a fixed level, on a logarithmic scale, falling along a straight line whose slope is the rate function, with the normal approximation drawn beside it and diverging.

The tail is not a bell

The limit theorem describes a window of width one over the root of n around the mean; ask instead for the chance that an average lands a fixed distance away and the answer falls exponentially, at a rate computed from the summand before any n is chosen.

probability · Central limit
How far each estimate can stray, when rare jumps are possible. The error that the sample mean and the median of 12 block means stay within, at confidence levels from 80% to 99.9%, over 60000 repeated samples of 120 draws from normal draws with rare large jumps, with Chebyshev's guarantee for the mean.

The median of many small averages

Knowing only that a quantity has a finite spread, the plain average of n samples can be promised to within σ/√(nδ) with confidence 1 − δ, and no better — Chebyshev's bound is tight, and rare large jumps achieve it. Cut the same samples into a dozen blocks, average each block, and take the median of the averages, and the promise improves to within about σ√(log(1/δ)/n). Nothing about the data has been assumed beyond the spread; only the way of combining it has changed.

probability · Concentration
What a rule collects when every value comes from the same distribution. uniform on [0, 1]: best rule 0.995, one threshold 0.635 of E[max] at n = 200; exponential: best rule 0.905, one threshold 0.678 of E[max] at n = 200; Pareto, tail exponent 2: best rule 0.802, one threshold 0.714 of E[max] at n = 200; Pareto, tail exponent 1.2: best rule 0.802, one threshold 0.682 of E[max] at n = 200.

When every value comes from the same hat

A rule that sees values one at a time and must keep or discard each on the spot can guarantee half of what a prophet collects, and no more, when the values come from different distributions. When they all come from the same one, the guarantee rises to 0.745 — and a single fixed threshold, set so that each value crosses it with chance 1/n, already secures 1 − 1/e. For bounded values the best rule collects nearly everything; only a heavy tail, where one enormous value carries the prize, keeps the gap open.

probability · Optimal stopping

Named alongside it

The objects these essays reach for when they reach for this one.

ExpectationNormal distributionVarianceIndependenceTail boundConcentration inequalityConvergenceConvergence rateCounterexampleScalingApproximationBackward induction

All concepts