Probability

The shape that averaging leaves alone

Adding independent quantities blurs their distributions together, and rescaling restores the width. Almost every shape is changed by that operation. Exactly one is returned unaltered, and that is why sums of unrelated things keep arriving at it.

Worth reading first: A bell curve assembled out of coin flips · The average settles and the wobble does not.

That sums of many small independent quantities look like a bell curve is the first rung of this ladder and is usually where the matter is left. It leaves the obvious question unasked: why that curve? Nothing in the statement of the problem mentions it, and a hundred other humped shapes exist.

A lopsided distribution added to itself, and the shape that returnsOn the left, the exact distribution of a sum of copies of one lopsided distribution, standardised, for several counts: the shapes converge. On the right, the bell curve convolved with itself, which is the bell curve again.-3-2-101230.00.20.40.6standardised valuen = 2n = 4n = 16n = 64the sums, standardised-4-20240.00.10.20.30.4valuethe bell, convolved with itselfone bell curvewidened by √2the convolution, computed0 nine times in ten, 10 otherwise, added to itself 4, 16, 64 times and standardised each time: the largest gapto the bell curve falls 0.491 → 0.404 → 0.206 → 0.105on the right, the bell curve convolved with itself on a grid, against the bell curve widened by √2 — the twoagree to 6.1e-16, which is what being the fixed point of the operation means
Fig. 1 On the left, a badly lopsided distribution added to itself repeatedly and standardised each time: the shapes walk in on one curve. On the right, that curve fed to the same operation — convolved with itself on a grid and compared with itself widened by the square root of two. It comes back unchanged.

The right-hand panel is the answer. Adding two independent copies of a quantity and rescaling is an operation on shapes; almost every shape is altered by it; the bell curve is not. A limit of repeated application must be something the operation leaves alone, and once that is the question, the curve stops being a coincidence and becomes the solution of an equation.

What adding does to a shape

If XX and YY are independent, the distribution of X+YX + Y is the convolution of theirs: the chance of the sum landing at zz is the total, over all ways of splitting zz into two parts, of the chance of the parts.

A Galton board after 600 balls600 balls fall through 12 rows of pegs, each bouncing left or right at random, and pile up in a bell-shaped heap.1173760121153109683661left or right, 12 times, 600 times over
Fig. 2 A board of pegs, where each ball’s position is a sum of twelve independent left-or-right decisions. Every extra row convolves the distribution with one more copy of a single step, and the pile is what twelve convolutions of a two-point distribution look like.

Two things happen at once. The distribution spreads, because variances add: the variance of a sum of independent quantities is the sum of their variances, exactly and with no approximation anywhere. And it smooths, because a convolution averages its input over the shape of what it is convolved with; corners round off, and the result is at least as smooth as either ingredient.

Left alone, spreading wins and the distribution flattens toward nothing. To see what the smoothing does, the width has to be put back — which means dividing by the square root of the number of terms, since that is what variances adding entails. Convolution followed by that rescaling is the operation this rung is about, and it maps shapes of variance one to shapes of variance one.

Why the limit has to be a fixed point

The argument that the limit, if it exists, is a fixed point is one line and it is worth having before any calculation.

Let TT be the operation: convolve a distribution with an independent copy of itself and rescale to keep the variance at one. Then applying TT to the standardised sum of 2n2n terms gives the standardised sum of 4n4n terms — adding two independent halves is the same as adding all the terms. So if the standardised sums converge to some shape FF, then TT applied to FF is the limit of the same sequence taken along a subsequence, which is FF again.

T(F)=F.T(F) = F.

Any limiting shape is therefore a fixed point of TT, and finding the possible limits becomes the algebraic problem of solving that equation. This is exactly the manoeuvre that makes an iterated map tractable: asking what an operation leaves alone before asking what it does.

Solving the equation

The equation is hard in the language of shapes and easy in another language, and the translation is Fourier’s.

The spectrum of the square waveOne bar per harmonic, its height the size of that harmonic's coefficient. The 11 non-zero coefficients fall away like 1 over m to the power 1.00.0.320.640.951.27123456789131721harmoniccoefficientthe square wave, harmonic by harmonic: 11 of the first 21 are non-zero, and their sizes fall like 1 / m^1.00the exponent is fitted to the bars rather than taken from the formula, and it is what says how smooth the target is:a jump gives 1, a corner gives 2, and a smooth function falls faster than any power
Fig. 3 A wave and its spectrum. The transform turns a function into its content at each frequency, and — the fact that matters here — it turns convolution into ordinary multiplication.

The characteristic function of a distribution is the Fourier transform of its density, φ(t)=E[eitX]\varphi(t) = \mathbb{E}[e^{itX}], and it converts the awkward operation into an easy one: the transform of a convolution is the product of the transforms. So TT becomes

φ(t)φ ⁣(t2)2,\varphi(t) \mapsto \varphi\!\left(\frac{t}{\sqrt{2}}\right)^{2},

and the fixed-point equation is φ(t)=φ(t/2)2\varphi(t) = \varphi(t/\sqrt{2})^2. Iterating that relation gives φ(t)=φ(t/2k)2k\varphi(t) = \varphi(t/\sqrt{2^{\,k}})^{2^{k}} for every kk, and expanding the inner factor near zero — where φ(0)=1\varphi(0) = 1, φ(0)=0\varphi'(0) = 0 for a centred distribution, and φ(0)=1\varphi''(0) = -1 for variance one — the limit as kk grows is

φ(t)=et2/2,\varphi(t) = e^{-t^2/2},

which is the transform of the bell curve. The fixed point is unique among distributions with finite variance, and the derivation shows exactly where each hypothesis is used: centring kills the first derivative, finite variance supplies the second, and the expansion needs nothing beyond it.

The self-consistency is visible in the shape itself. The bell is the exponential of a quadratic; multiplying two of them adds the quadratics, which is again a quadratic — and no other family of functions has that property. Convolution turning into multiplication is what makes sums of independent quantities behave like products of ordinary functions, and the exponential is the function that converts between the two.

The check, done rather than quoted

The right-hand panel of the first figure does not cite this argument; it performs the convolution.

The standard bell curve is sampled on a grid, convolved with itself by summing the product over the grid, and the result is compared with the bell curve of standard deviation 2\sqrt{2} at nine test points. The two agree to within 101510^{-15}, which is the arithmetic precision available rather than a property of the approximation — a Gaussian integrand on a fine grid is one of the few things a trapezoid sum computes almost exactly.

That is the difference between a figure and an illustration. Nothing in the drawing depends on knowing the answer: had the fixed-point property been false, the two curves would have separated and the generator would have refused to draw. It is the same discipline as checking a dissection tiles its target, applied to an identity between functions.

A lopsided distribution added to itself, and the shape that returnsOn the left, the exact distribution of a sum of copies of one lopsided distribution, standardised, for several counts: the shapes converge. On the right, the bell curve convolved with itself, which is the bell curve again.-3-2-101230.00.20.40.6standardised valuen = 2n = 4n = 16n = 64the sums, standardised-4-20240.00.10.20.30.4valuethe bell, convolved with itselfone bell curvewidened by √2the convolution, computedfive values, unevenly weighted, added to itself 4, 16, 64 times and standardised each time: the largest gap tothe bell curve falls 0.121 → 0.072 → 0.032 → 0.016on the right, the bell curve convolved with itself on a grid, against the bell curve widened by √2 — the twoagree to 6.1e-16, which is what being the fixed point of the operation means
Fig. 4 A different lopsided starting distribution, put through the same machinery. The intermediate shapes are different and the destination is the same, which is what a fixed point that attracts looks like.

Two other ways of singling it out

A shape characterised in one way is a curiosity; a shape characterised in three unrelated ways is a fact about the subject. The bell has at least three, and the other two involve no sums at all.

The unit ball at p = 2.00The set of points one unit from the origin, when distance is measured by the p-th power sum. At p = 1 it is a diamond, at p = 2 a circle, and as p grows it fills out a square.11p = 2the diagonal point is 0.707convex — so this is a distancedashed: p = 1dashed: p → ∞the set of points at distance one from the origin when distance is (|x|^p + |y|^p)^(1/p), at p = 2.00p = 1 is the diamond of city blocks, p = 2 the circle, and large p a square; the ball meets the diagonal at0.707 in each coordinate, which is where the exponent shows most
Fig. 5 The unit ball of the ordinary distance, with the diamond and the square behind it. Only the round one survives a rotation — which is the hypothesis of the argument below, and the reason the bell curve appears in it.

Rotational symmetry with independence. Suppose a random point in the plane has independent coordinates, and suppose its distribution depends only on the distance from the origin. Then the density factorises as f(x)f(y)f(x)f(y) and simultaneously depends only on x2+y2x^2 + y^2, so f(x)f(y)f(x)f(y) must be a function of x2+y2x^2 + y^2 alone. The only functions turning a sum into a product are exponentials, and the conclusion is f(x)=cekx2f(x) = c\,e^{-kx^2}. That is Herschel’s argument of 1850, recovered by Maxwell for velocities in a gas, and it produces the bell curve out of a symmetry with no limit and no sum anywhere.

Notice which distance is being used. The argument needs rotations to preserve the quantity the density depends on, and only the ordinary distance is unchanged by a rotation — so the appearance of squares in the exponent traces directly back to the appearance of squares in the theorem this collection began with.

Maximum spread for a given width. Among all distributions with a specified variance, the bell curve is the one with the largest entropy — the most spread-out, in the precise sense that counts how many outcomes are effectively possible. Any other shape of the same variance encodes some additional structure, and adding independent quantities destroys structure; so the limit of repeated addition ought to be the shape with none left, which is what this characterisation says.

The three descriptions — fixed under convolution, invariant under rotation with independent coordinates, and maximally spread at fixed width — pick out the same curve for reasons that look unrelated. That is the usual sign that an object is not a construction but a fact.

Watching the convolutions happen

Convolution is easier to believe when the individual contributions can be seen, and a board of pegs is the standard apparatus for that.

5 paths through a Galton boardA few individual routes down through the pegs, each a sequence of coin flips.
Fig. 6 The paths through the board, rather than the pile at the bottom. Each ball’s final bin is decided by how many of its decisions went right, and the number of routes to a bin is the binomial coefficient — so the pile is a picture of counting, and the smoothing is a consequence of there being many routes to the middle and one to each edge.

Two features of the pile are worth separating, because they are the two effects the operation has.

The pile is symmetric and humped because the routes are counted by binomial coefficients, which peak in the middle. That is combinatorics, and it holds at every row count.

The pile becomes smooth as the rows increase, and that is convolution’s doing: each row averages the previous row’s counts with a two-point kernel, and repeated averaging removes fine structure. A jagged distribution put through the same machinery loses its jaggedness in a handful of steps.

One set of sums, two scalings, two different limitsThe exact distribution of a sum of n independent copies, scaled two ways. Divided by n it collapses onto the mean; divided by the square root of n it holds a fixed width and settles into a shape.-2-1120.00.51.0the average, minus the meann = 2n = 8n = 32-6-4-22460.00.10.20.3the sum, minus n means, over √nn = 2n = 8n = 32the same exact distributions, drawn twice as densities: divided by n the spread falls 1.306 → 0.653 →0.326, divided by √n it is 1.847 every timeso the left curves climb — 0.40 → 0.61 → 1.22 at the peak — and the right ones settle on 0.216, which is theheight of the bell curve
Fig. 7 The same sums under two rescalings. Divided by the number of terms they collapse to a point; divided by the square root they hold a fixed width and settle into a shape. Only one of the two scalings has a fixed point worth having, and the exponent is decided by variances adding.

The rescaling has to be right or nothing is fixed, and the exponent is not a matter of taste. Divide by too much and every shape flattens to a spike; too little and every shape spreads to nothing. The operation only has a fixed point at all when the rescaling matches the rate at which variances accumulate, which is why the square root is not a convention but the answer to the question of which power leaves anything behind.

Attracting, and not merely fixed

A fixed point can be repelling — the difference decides everything about an iteration — so the equation alone does not explain why sums arrive at the bell.

The left panel is the evidence that it attracts: two quite different starting distributions, standardised, converge on it, and the largest gap between the cumulative distributions falls at every doubling. The theorem behind that is the central limit theorem itself, and the fixed-point argument does not replace it. What the fixed-point argument supplies is the identity of the limit; what the theorem supplies is that there is one.

The two halves are worth keeping separate, because they fail separately. There are situations where sums converge to something that is not the bell — the next rung — and situations where they converge to nothing at all. The fixed point of TT under finite variance is unique; the question of which starting shapes are drawn to it is a separate one with a separate answer.

Where it fails

Everything above used finite variance twice: to make the rescaling n\sqrt{n} and to supply the second derivative at zero.

Drop it, and the fixed-point equation still has solutions — a whole family of them, the stable laws, one for each exponent α\alpha between zero and two, with the bell as the case α=2\alpha = 2. Each is a fixed point of convolution with rescaling by n1/αn^{1/\alpha} rather than n1/2n^{1/2}, and each attracts the sums of quantities whose tails decay at the matching rate.

The best-known of the others is the Cauchy distribution, at α=1\alpha = 1, and its fixed-point property is the reason averaging Cauchy quantities achieves nothing at all: the average of nn of them is fixed with no rescaling whatever.

So the bell’s monopoly is conditional, and the condition is exactly the finiteness of the second moment. Under that condition it is unique; without it, it is one of a continuum.

What the fixed point does not explain

Two things are commonly read off this argument that it does not support, and both matter in practice.

It does not say that a sum of a few terms is bell-shaped. The fixed point is where the operation ends up, and the left panel shows the journey taking many doublings from a badly lopsided start. Four terms of a skewed quantity are not approximately normal in any useful sense, and treating them as such is the standard misuse of the theorem.

It does not say that anything with a hump is a bell curve. Many distributions look like the drawn shape at the resolution of a picture and differ from it exactly where the difference matters, in the tails. The fixed-point property is about the whole function; two shapes agreeing to within the thickness of the line near the middle can disagree by factors of thousands four standard deviations out, and it is the far tail that decides how often a rare event happens.

The honest summary is that the argument identifies the limit and says nothing about how far along the journey any particular sum has got. That question is a different rung, it has a quantitative answer, and the answer is slower than most people expect.

What the picture cannot show

The left panel draws four distributions from a family that is discrete, so each is a set of spikes drawn as a density by spreading each atom over the lattice spacing. That is a convention, and the convergence being drawn is convergence of the cumulative distributions rather than of the drawn curves — a lattice distribution never becomes a smooth density, however many terms are added, and the local shape at the scale of the spacing does not converge at all.

The right panel cannot show uniqueness. It shows that the bell is fixed; the argument that nothing else with finite variance is fixed lives in the transform, and the transform is not on the page. A figure could plausibly be drawn showing some other shape failing to be fixed, and it would establish nothing about the shapes not drawn.

And neither panel shows the rate. The gaps quoted fall from 0.4910.491 to 0.1050.105 over five doublings, which is a slow business, and how slow it is is the next rung.

The ladder from here

Below: the bell curve from coin flips and the two scalings that separate the law of large numbers from the limit theorem. Above: the average that never settles, where finite variance fails, and how fast the bell arrives, which is the rate this rung is silent about. Sideways: the contraction principle, which is the same fixed-point argument in a metric space, and the walk that becomes a curve, which is this limit taken in a space of paths rather than of numbers.

Asking what an operation leaves alone

The lasting point is the change of question. Why do sums of unrelated things look like a bell curve is a question about a limit and is hard. What shapes does convolution-and-rescaling leave alone is a question about an equation, and it has one answer.

The habit generalises. A quantity conserved in expectation forces a ruin probability to be linear; a shape unchanged by a set of contractions is a fractal; a direction unmoved by a linear map is an eigenvector. In every case the object that seemed to require a limit turns out to be pinned down by an invariance, and the limit merely explains why it is arrived at rather than what it is.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

ConvergenceFixed pointFourier analysisIndependenceInvariantNormal distributionScalingVariance