How fast the bell arrives
Worth reading first: The shape that averaging leaves alone · A bell curve assembled out of coin flips.
A theorem that says a sequence converges is only half an answer. The other half is how fast, and for the limit theorem the answer is unwelcome: the error falls like one over the square root of the number of terms, which is the slowest rate that anybody calls convergence.
Both quantities in that picture are computed rather than sampled: the distribution of the sum is built by convolving the one-step distribution with itself, and the gap is the largest difference between its cumulative distribution and the bell curve’s. The straight lines on logarithmic axes are what a power law looks like, and their slope is the exponent.
The rate, stated
The theorem is Berry’s and Esseen’s, from 1941 and 1942 independently. For independent copies of a quantity with mean zero, standard deviation and third absolute moment ,
where is the cumulative distribution of the standardised sum, is the bell curve’s, and is an absolute constant — currently known to be below and known to be at least .
Three features of that statement are worth pulling out.
The rate is and no better in general. Quadrupling the number of terms halves the error. To get two more decimal places, multiply the terms by ten thousand.
The constant carries the third moment. The quantity is essentially the skewness — how lopsided the summand is — and it is the only feature of the summand that appears. A symmetric summand has a smaller one; a badly lopsided one has a large one, and its sums are correspondingly further from the bell at every .
The bound is uniform in . It controls the worst point of the whole distribution at once, which is what makes it usable and also what makes it pessimistic: the worst point is not where most of the mass is.
Where the square root comes from
The rate is not an artefact of the proof. It comes from the first correction term to the limit, and that term can be written down.
Expanding the distribution of a standardised sum in powers of — the Edgeworth expansion — gives
where is the skewness and the bell curve’s density. The leading error is exactly the skewness divided by , multiplied by a fixed shape.
That formula explains the whole of the first figure. The slope is because the leading term carries . The lopsided summand sits above the symmetric one because is larger. And the shape vanishes at and peaks near the shoulders, so the error is concentrated where the distribution is bending rather than at its peak or in its tails.
What a symmetric summand buys, and what it does not
If the summand is symmetric, the skewness is zero and the leading term vanishes — so the error should fall like rather than . The first figure shows a symmetric summand, a fair coin, converging at all the same. The discrepancy is instructive.
The reason is lattice spacing. A sum of coin flips lives on a lattice of spacing , and after dividing by the spacing is . Its cumulative distribution is a staircase with steps of that height, and no smooth curve can be closer to a staircase than half a step. So the discreteness alone forces a gap of order , whatever the skewness does.
This is a real effect and not a technicality: the improvement from symmetry is available for summands with a density and unavailable for lattice ones. For a lattice summand there is a correction — a continuity adjustment, which shifts the comparison by half a step — and with it the error does fall like for symmetric cases. Without it, the rate is what the picture shows.
The general lesson is that a rate can be limited by something entirely unrelated to the mechanism it is supposed to measure. Here the mechanism is smoothing by convolution and the limit is set by arithmetic — the sum of whole numbers is a whole number, and no amount of averaging changes that.
The middle and the tails converge at different speeds
The Berry–Esseen bound is uniform, and uniformity conceals the most important practical fact about the approximation: it is far better in the middle than at the edges.
Consider a sum of a hundred independent draws and ask for the chance of exceeding four standard deviations. The bell curve says about . The true answer depends heavily on the summand and can easily be a factor of two or ten away, in either direction — while the Berry–Esseen bound at is about , which is a thousand times larger than the quantity being estimated and therefore says nothing at all.
The bound is not weak; it is answering a different question. It controls the absolute error uniformly, and in the tail the absolute error being small is compatible with the relative error being enormous. Getting the tail right needs a different theory altogether, in which the error is multiplicative and the rate depends on how far out the question is asked.
That distinction — absolute error uniform, relative error unbounded in the tail — is the single most useful thing to carry away from this rung, because the bell curve is used most often precisely where it is worst: to estimate the chance of something rare. The same trap in another subject is a ratio tending to one while the difference grows without bound.
The extremal case, and where the constant comes from
A bound with a constant in it invites the question of what makes the constant what it is, and here the answer is a specific distribution.
The lower bound of on the constant comes from a two-point summand taking one value with a small probability and another with a large one, and it is a lattice distribution — so the mechanism forcing the constant up is the same staircase effect that stops a symmetric coin from converging faster. The worst case for the rate is not a badly skewed continuous distribution but a discrete one, whose sums can never be smooth at all.
Shevtsova brought the upper constant down to in 2011, from Esseen’s original value of about , and the gap between the two numbers has been the state of the subject for over a decade. That gap is worth noticing: the theorem is eighty years old, the constant is known to within about fifteen per cent, and closing the remaining gap has resisted everybody who has tried.
The exact answer, when it is available
For sums of independent copies of a known discrete quantity, none of this machinery is needed, and it is worth saying plainly because it is often forgotten.
Convolving a distribution with itself times is a finite computation whose cost grows with the size of the support, and for the sums drawn in this rung it takes a fraction of a second. The result is not an approximation at all: it is the distribution, and every probability the sum has can be read off it directly.
The approximation earns its place in three situations, and only in those. When the summands have different distributions, so there is nothing to convolve repeatedly. When the summands are not known individually and only their moments are available. And when the number of terms is large enough that convolution is impractical — which, for a summand on a lattice, means very large indeed.
The habit worth having is to check whether the exact computation is available before reaching for the limit. In this collection it usually is, which is why every figure in this ladder is exact and the theorem is discussed rather than used.
How the theorem grew a hypothesis, and then lost one
The rate arrived late in a long process of finding out what the limit theorem actually requires.
De Moivre had the coin-flipping case in 1733 and Laplace generalised it in 1812, both for sums of identically distributed terms. The question of what happens when the terms differ took another century. Lyapunov gave a workable condition in 1901 — a bound on a moment slightly higher than the second — and it is his method, the characteristic function, that every subsequent proof uses.
Lindeberg replaced it in 1922 with a condition that turned out to be the right one: no single term may contribute a significant share of the total variance. Feller showed in 1935 that, in the presence of a mild extra requirement, Lindeberg’s condition is not merely sufficient but necessary, so the question of when sums are asymptotically normal was closed.
Berry and Esseen then asked the quantitative version and found that a bound needs slightly more than Lindeberg’s condition: a third moment, which the condition does not require. That is the shape of the whole story — the qualitative theorem needs less than anybody expected and the quantitative one needs more, and the extra moment is exactly the price of a number instead of a limit.
How many terms are enough
The practical question has a practical answer, and it depends on what is being asked for rather than on a number of terms.
For the centre of the distribution — the chance of being within one or two standard deviations — a handful of terms is often enough, and the folklore figure of thirty comes from that. For the shoulders the requirement is larger. For a tail probability of one in a thousand, no plausible number of terms makes the bell curve reliable if the summand is lopsided, because the relative error there is governed by the tail of the summand, which never disappears.
The honest procedure is the one this rung’s figures use: for a sum of independent copies of a known discrete quantity, the exact distribution is computable by convolution, and the approximation is unnecessary. The limit theorem earns its place where the summands are not identically distributed, are not known individually, or are too numerous to convolve — and in exactly those cases the rate is the only thing available.
What the rate is not
Two readings of the bound are common and both are wrong.
It is not a statement that the sum is nearly normal after enough terms in every sense. It bounds one particular distance — the largest vertical gap between two cumulative distributions — and a small value of that distance is compatible with the two densities looking quite different, since a density is a derivative and derivatives are not controlled by a bound on the functions.
It is also not a statement about a particular sum. Everything here concerns distributions, and a single realised sum is a number; asking whether that number is close to normal is not a question. The same confusion between a sample and a distribution is what makes the convergence of a rescaled walk hard to state, and it is worth guarding against wherever a limit theorem is quoted about data rather than about a law.
What the bound does say is exactly what it says: at terms, no probability computed from the bell curve is wrong by more than a stated amount. That is a modest claim and a genuinely useful one, and it is the strongest thing available without knowing more about the summand than its first three moments — which is the same trade that Chebyshev’s inequality makes at the level of two.
What the picture cannot show
The plot stops at a hundred and twenty-eight terms, and the interesting range for the tail question is unreachable: the quantity that misbehaves is a probability of order , and a plot of the gap in absolute terms cannot show a relative error in something that small.
The lines are also computed for one summand each. The Berry–Esseen bound is over all summands with a given third moment, and the worst case is not the die or the coin — it is a two-point distribution chosen to be as lopsided as the constraint allows, and the constant mentioned above comes from a family of exactly such examples.
And the whole picture is about sums of independent, identically distributed quantities. Both hypotheses can be relaxed and the rate survives in a modified form, but a figure of this kind would have to choose a particular way of relaxing them, and the choice would carry more of the answer than the drawing would.
The ladder from here
Below: the bell curve from coin flips, the two scalings and the fixed point that identifies the limit. Sideways: the average that never settles, where the third moment is infinite and the rate is not merely slow but absent, and Chebyshev’s inequality, which gives a bound with no convergence at all and needs only two moments. Above: large deviations, where the tail is measured multiplicatively and the answer is an exponential rate rather than a power.
A theorem, and the question it does not answer
The lasting point is that an existence statement and a rate are different results with different proofs, and that the first is nearly useless without the second.
The sums converge to the bell curve is compatible with the approximation being worthless at every sample size anybody will ever have. What makes it usable is a bound with a number in it, and the bound turns out to depend on the summand through exactly one quantity — its skewness — which is a stronger and more surprising statement than the convergence itself.
Asking how fast is the standard follow-up to any limit in this collection, and it is nearly always harder than the limit. The count of primes converges to its estimate and the rate is the deepest open problem in the subject; an iteration converges to its fixed point and the rate is what decides whether the iteration is worth running.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- How long until every one turns up — both name approximation, convergence rate, expectation
- The staircase that is not the diagonal — both name approximation, counterexample
Named objects
A dashed tag is an object no other essay names yet.
ApproximationConvergence rateCounterexampleExpectationHeavy tailsNormal distributionScalingVariance