Probability

The shape that averaging leaves alone

Adding independent quantities blurs their distributions together, and rescaling restores the width. Almost every shape is changed by that operation. Exactly one is returned unaltered, and that is why sums of unrelated things keep arriving at it.

Worth reading first: A bell curve assembled out of coin flips · The average settles and the wobble does not.

That sums of many small independent quantities look like a bell curve is the first rung of this ladder and is usually where the matter is left. It leaves the obvious question unasked: why that curve? Nothing in the statement of the problem mentions it, and a hundred other humped shapes exist.

A lopsided distribution added to itself, and the shape that returns. On the left, the exact distribution of a sum of copies of one lopsided distribution, standardised, for several counts: the shapes converge. On the right, the bell curve convolved with itself, which is the bell curve again.
Fig. 1 On the left, a badly lopsided distribution added to itself repeatedly and standardised each time: the shapes walk in on one curve. On the right, that curve fed to the same operation — convolved with itself on a grid and compared with itself widened by the square root of two. It comes back unchanged.

The right-hand panel is the answer. Adding two independent copies of a quantity and rescaling is an operation on shapes; almost every shape is altered by it; the bell curve is not. A limit of repeated application must be something the operation leaves alone, and once that is the question, the curve stops being a coincidence and becomes the solution of an equation.

What adding does to a shape

If XX and YY are independent, the distribution of X+YX + Y is the convolution of theirs: the chance of the sum landing at zz is the total, over all ways of splitting zz into two parts, of the chance of the parts.

A Galton board after 600 balls. 600 balls fall through 12 rows of pegs, each bouncing left or right at random, and pile up in a bell-shaped heap.
Fig. 2 A board of pegs, where each ball’s position is a sum of twelve independent left-or-right decisions. Every extra row convolves the distribution with one more copy of a single step, and the pile is what twelve convolutions of a two-point distribution look like.

Two things happen at once. The distribution spreads, because variances add: the variance of a sum of independent quantities is the sum of their variances, exactly and with no approximation anywhere. And it smooths, because a convolution averages its input over the shape of what it is convolved with; corners round off, and the result is at least as smooth as either ingredient.

Left alone, spreading wins and the distribution flattens toward nothing. To see what the smoothing does, the width has to be put back — which means dividing by the square root of the number of terms, since that is what variances adding entails. Convolution followed by that rescaling is the operation this rung is about, and it maps shapes of variance one to shapes of variance one.

Why the limit has to be a fixed point

The argument that the limit, if it exists, is a fixed point is one line and it is worth having before any calculation.

Let TT be the operation: convolve a distribution with an independent copy of itself and rescale to keep the variance at one. Then applying TT to the standardised sum of 2n2n terms gives the standardised sum of 4n4n terms — adding two independent halves is the same as adding all the terms. So if the standardised sums converge to some shape FF, then TT applied to FF is the limit of the same sequence taken along a subsequence, which is FF again.

T(F)=F.T(F) = F.

Any limiting shape is therefore a fixed point of TT, and finding the possible limits becomes the algebraic problem of solving that equation. This is exactly the manoeuvre that makes an iterated map tractable: asking what an operation leaves alone before asking what it does.

Solving the equation

The equation is hard in the language of shapes and easy in another language, and the translation is Fourier’s.

The spectrum of the square wave. One bar per harmonic, its height the size of that harmonic's coefficient. The 11 non-zero coefficients fall away like 1 over m to the power 1.00.
Fig. 3 A wave and its spectrum. The transform turns a function into its content at each frequency, and — the fact that matters here — it turns convolution into ordinary multiplication.

The characteristic function of a distribution is the Fourier transform of its density, φ(t)=E[eitX]\varphi(t) = \mathbb{E}[e^{itX}], and it converts the awkward operation into an easy one: the transform of a convolution is the product of the transforms. So TT becomes

φ(t)↦φ ⁣(t2)2,\varphi(t) \mapsto \varphi\!\left(\frac{t}{\sqrt{2}}\right)^{2},

and the fixed-point equation is φ(t)=φ(t/2)2\varphi(t) = \varphi(t/\sqrt{2})^2. Iterating that relation gives φ(t)=φ(t/2 k)2k\varphi(t) = \varphi(t/\sqrt{2^{\,k}})^{2^{k}} for every kk, and expanding the inner factor near zero — where φ(0)=1\varphi(0) = 1, φ′(0)=0\varphi'(0) = 0 for a centred distribution, and φ′′(0)=−1\varphi''(0) = -1 for variance one — the limit as kk grows is

φ(t)=e−t2/2,\varphi(t) = e^{-t^2/2},

which is the transform of the bell curve. The fixed point is unique among distributions with finite variance, and the derivation shows exactly where each hypothesis is used: centring kills the first derivative, finite variance supplies the second, and the expansion needs nothing beyond it.

The self-consistency is visible in the shape itself. The bell is the exponential of a quadratic; multiplying two of them adds the quadratics, which is again a quadratic — and no other family of functions has that property. Convolution turning into multiplication is what makes sums of independent quantities behave like products of ordinary functions, and the exponential is the function that converts between the two.

The check, done rather than quoted

The right-hand panel of the first figure does not cite this argument; it performs the convolution.

The standard bell curve is sampled on a grid, convolved with itself by summing the product over the grid, and the result is compared with the bell curve of standard deviation 2\sqrt{2} at nine test points. The two agree to within 10−1510^{-15}, which is the arithmetic precision available rather than a property of the approximation — a Gaussian integrand on a fine grid is one of the few things a trapezoid sum computes almost exactly.

That is the difference between a figure and an illustration. Nothing in the drawing depends on knowing the answer: had the fixed-point property been false, the two curves would have separated and the generator would have refused to draw. It is the same discipline as checking a dissection tiles its target, applied to an identity between functions.

A lopsided distribution added to itself, and the shape that returns. On the left, the exact distribution of a sum of copies of one lopsided distribution, standardised, for several counts: the shapes converge. On the right, the bell curve convolved with itself, which is the bell curve again.
Fig. 4 A different lopsided starting distribution, put through the same machinery. The intermediate shapes are different and the destination is the same, which is what a fixed point that attracts looks like.

Two other ways of singling it out

A shape characterised in one way is a curiosity; a shape characterised in three unrelated ways is a fact about the subject. The bell has at least three, and the other two involve no sums at all.

Averages of a heavy-tailed quantity, which never settle. Running averages of draws from a Cauchy distribution, which jump rather than converge, beside the cumulative distributions of averages of 1, 4 and 16 draws, which lie on top of one another.
Fig. 5 And what the hypothesis is holding up. The same standardising applied to a distribution with no variance at all: on the left, five running averages that jump to new levels and stay there rather than settling; on the right, twenty thousand averages of 11, 44 and 1616 draws whose cumulative distributions lie exactly on top of one another and on top of a single draw’s. Averaging changes nothing. The fixed point is still a fixed point — it is simply a different one, and the theorem above is a statement about which distributions are pulled toward which.

Rotational symmetry with independence. Suppose a random point in the plane has independent coordinates, and suppose its distribution depends only on the distance from the origin. Then the density factorises as f(x)f(y)f(x)f(y) and simultaneously depends only on x2+y2x^2 + y^2, so f(x)f(y)f(x)f(y) must be a function of x2+y2x^2 + y^2 alone. The only functions turning a sum into a product are exponentials, and the conclusion is f(x)=c e−kx2f(x) = c\,e^{-kx^2}. That is Herschel’s argument of 1850, recovered by Maxwell for velocities in a gas, and it produces the bell curve out of a symmetry with no limit and no sum anywhere.

Notice which distance is being used. The argument needs rotations to preserve the quantity the density depends on, and only the ordinary distance is unchanged by a rotation — so the appearance of squares in the exponent traces directly back to the appearance of squares in the theorem this collection began with.

Maximum spread for a given width. Among all distributions with a specified variance, the bell curve is the one with the largest entropy — the most spread-out, in the precise sense that counts how many outcomes are effectively possible. Any other shape of the same variance encodes some additional structure, and adding independent quantities destroys structure; so the limit of repeated addition ought to be the shape with none left, which is what this characterisation says.

The three descriptions — fixed under convolution, invariant under rotation with independent coordinates, and maximally spread at fixed width — pick out the same curve for reasons that look unrelated. That is the usual sign that an object is not a construction but a fact.

A fourth description, in the tool itself

There is a fourth characterisation and it is the one closest to the machinery above, because it is a statement about the transform rather than about distributions.

The bell curve is its own Fourier transform. The derivation ended at φ(t)=e−t2/2\varphi(t) = e^{-t^2/2}, which is the same function of tt that the density is of xx — and that is not a convenience of notation. Transforming a Gaussian gives a Gaussian, and the standard one is returned unchanged, eigenvalue and all.

The Fourier transform applied four times is the identity, so its eigenvalues can only be the four fourth roots of one, and every function decomposes into four pieces on which the transform acts as 11, −i-i, −1-1 and ii. The bell curve is the simplest member of the first family — the ground state, in the vocabulary the physicists use — and the rest of that family are the Hermite functions, the bell curve multiplied by polynomials. So the object this essay is about is not merely fixed by convolution; it is fixed by the transform that turns convolution into multiplication, which is why the two facts kept collapsing into each other in the derivation.

That also gives the shortest route to a property with no probability in it at all. A function and its transform cannot both be concentrated: making one narrow makes the other wide, and the product of the two spreads has a floor. The bell curve is the shape that sits exactly on that floor — the unique minimiser, up to shifting and scaling — which is the mathematical content of the uncertainty principle, arrived at without mentioning a particle. A shape that is its own transform can hardly do otherwise, since narrowing it would have to widen it.

Five characterisations now, and it is worth noting what they share, because the list would otherwise look like a coincidence pile. Fixed under convolution and rescaling; the only rotation-invariant density with independent coordinates; maximum entropy at fixed variance; its own Fourier transform; and minimal in the spread product. Every one of them is an extremal or invariance property — something the curve is alone in doing, or alone in not being changed by — and none of them constructs it. That is the general situation for objects that turn up everywhere: they are not built, they are cornered, and the many descriptions are many ways of closing off the alternatives.

Watching the convolutions happen

Convolution is easier to believe when the individual contributions can be seen, and a board of pegs is the standard apparatus for that.

How fast a sum becomes a bell curve. The largest gap between the distribution of a standardised sum and the bell curve, against the number of terms, on logarithmic axes. Both summands fall along a line of slope about minus a half.
Fig. 6 How fast the shape arrives, since the theorem says only that it does. The largest gap between the exact distribution of a standardised sum and the bell curve, against the number of terms, on logarithmic axes: a line of slope about minus a half for both summands drawn, with the lopsided one sitting above the symmetric one at every count. Halving the gap costs four times as many terms, which is the slowest rate anybody calls convergence — and the vertical offset between the two lines is the summands’ skewness, which is what the constant in the bound is made of.

Two features of the pile are worth separating, because they are the two effects the operation has.

The pile is symmetric and humped because the routes are counted by binomial coefficients, which peak in the middle. That is combinatorics, and it holds at every row count.

The pile becomes smooth as the rows increase, and that is convolution’s doing: each row averages the previous row’s counts with a two-point kernel, and repeated averaging removes fine structure. A jagged distribution put through the same machinery loses its jaggedness in a handful of steps.

One set of sums, two scalings, two different limits. The exact distribution of a sum of n independent copies, scaled two ways. Divided by n it collapses onto the mean; divided by the square root of n it holds a fixed width and settles into a shape.
Fig. 7 The same sums under two rescalings. Divided by the number of terms they collapse to a point; divided by the square root they hold a fixed width and settle into a shape. Only one of the two scalings has a fixed point worth having, and the exponent is decided by variances adding.

The rescaling has to be right or nothing is fixed, and the exponent is not a matter of taste. Divide by too much and every shape flattens to a spike; too little and every shape spreads to nothing. The operation only has a fixed point at all when the rescaling matches the rate at which variances accumulate, which is why the square root is not a convention but the answer to the question of which power leaves anything behind.

Attracting, and not merely fixed

A fixed point can be repelling — the difference decides everything about an iteration — so the equation alone does not explain why sums arrive at the bell.

The left panel is the evidence that it attracts: two quite different starting distributions, standardised, converge on it, and the largest gap between the cumulative distributions falls at every doubling. The theorem behind that is the central limit theorem itself, and the fixed-point argument does not replace it. What the fixed-point argument supplies is the identity of the limit; what the theorem supplies is that there is one.

The two halves are worth keeping separate, because they fail separately. There are situations where sums converge to something that is not the bell — the next rung — and situations where they converge to nothing at all. The fixed point of TT under finite variance is unique; the question of which starting shapes are drawn to it is a separate one with a separate answer.

Where it fails

Everything above used finite variance twice: to make the rescaling n\sqrt{n} and to supply the second derivative at zero.

Drop it, and the fixed-point equation still has solutions — a whole family of them, the stable laws, one for each exponent α\alpha between zero and two, with the bell as the case α=2\alpha = 2. Each is a fixed point of convolution with rescaling by n1/αn^{1/\alpha} rather than n1/2n^{1/2}, and each attracts the sums of quantities whose tails decay at the matching rate.

The best-known of the others is the Cauchy distribution, at α=1\alpha = 1, and its fixed-point property is the reason averaging Cauchy quantities achieves nothing at all: the average of nn of them is fixed with no rescaling whatever.

So the bell’s monopoly is conditional, and the condition is exactly the finiteness of the second moment. Under that condition it is unique; without it, it is one of a continuum.

What the fixed point does not explain

Two things are commonly read off this argument that it does not support, and both matter in practice.

It does not say that a sum of a few terms is bell-shaped. The fixed point is where the operation ends up, and the left panel shows the journey taking many doublings from a badly lopsided start. Four terms of a skewed quantity are not approximately normal in any useful sense, and treating them as such is the standard misuse of the theorem.

It does not say that anything with a hump is a bell curve. Many distributions look like the drawn shape at the resolution of a picture and differ from it exactly where the difference matters, in the tails. The fixed-point property is about the whole function; two shapes agreeing to within the thickness of the line near the middle can disagree by factors of thousands four standard deviations out, and it is the far tail that decides how often a rare event happens.

The honest summary is that the argument identifies the limit and says nothing about how far along the journey any particular sum has got. That question is a different rung, it has a quantitative answer, and the answer is slower than most people expect.

What the picture cannot show

The left panel draws four distributions from a family that is discrete, so each is a set of spikes drawn as a density by spreading each atom over the lattice spacing. That is a convention, and the convergence being drawn is convergence of the cumulative distributions rather than of the drawn curves — a lattice distribution never becomes a smooth density, however many terms are added, and the local shape at the scale of the spacing does not converge at all.

The right panel cannot show uniqueness. It shows that the bell is fixed; the argument that nothing else with finite variance is fixed lives in the transform, and the transform is not on the page. A figure could plausibly be drawn showing some other shape failing to be fixed, and it would establish nothing about the shapes not drawn.

And neither panel shows the rate. The gaps quoted fall from 0.4910.491 to 0.1050.105 over five doublings, which is a slow business, and how slow it is is the next rung.

The ladder from here

Below: the bell curve from coin flips and the two scalings that separate the law of large numbers from the limit theorem. Above: the average that never settles, where finite variance fails, and how fast the bell arrives, which is the rate this rung is silent about. Sideways: the contraction principle, which is the same fixed-point argument in a metric space, and the walk that becomes a curve, which is this limit taken in a space of paths rather than of numbers.

Asking what an operation leaves alone

The lasting point is the change of question. Why do sums of unrelated things look like a bell curve is a question about a limit and is hard. What shapes does convolution-and-rescaling leave alone is a question about an equation, and it has one answer.

The habit generalises. A quantity conserved in expectation forces a ruin probability to be linear; a shape unchanged by a set of contractions is a fractal; a direction unmoved by a linear map is an eigenvector. In every case the object that seemed to require a limit turns out to be pinned down by an invariance, and the limit merely explains why it is arrived at rather than what it is.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

ConvergenceFixed pointFourier analysisIndependenceInvariantNormal distributionScalingVariance