The tail is not a bell
Worth reading first: A bell curve assembled out of coin flips · How far from the average a thing can be.
Toss a fair coin scoring sixty-four times and average. The central limit theorem says the average is approximately normal with standard deviation , and that approximation is excellent — for questions about the region within a few eighths of zero.
Now ask a different question: what is the chance the average is at least ? That is six standard deviations out, the normal approximation puts it near , and the true answer is about four times smaller. Push the number of tosses up and the discrepancy grows, because the two answers are not close approximations to each other at all: they are governed by different exponents.
The subject that describes the second question is called large deviations, and its central fact is that the probability falls like for a number that can be computed from the single-draw distribution before any is chosen.
Two regimes, and where the limit theorem lives
The reason for the discrepancy is visible in what the limit theorem actually claims.
The central limit theorem says the distribution of converges to the standard normal. Fix a number and the theorem describes the event that the average is within of the mean — a window that closes as grows. A fixed level sits at , which runs off to infinity, and the theorem says nothing about a region that is escaping from it.
Worse, the error in the approximation is of the same order as the answer out there. Berry and Esseen’s bound says the largest gap between the exact distribution function and the normal one falls like — which is tiny compared with probabilities near the middle and vast compared with a probability of . An approximation whose absolute error dwarfs the quantity being estimated carries no information about it.
So the two regimes are genuinely separate. Within of the mean: the bell curve, error controlled. At a fixed distance from the mean: a different question, needing a different theorem, and the answer is exponentially small rather than merely small.
The rate, and where it comes from
Cramér’s theorem, from 1938, gives the answer. For independent copies of a distribution with mean , and any level ,
where the rate function is
is the moment generating function, and is its logarithm’s Legendre transform — the same construction that turns a Lagrangian into a Hamiltonian in mechanics and energy into free energy in thermodynamics.
The transform is hard to believe as a formula and obvious as a picture. For each slope , slide a line of that slope down until it just touches the curve ; the intercept where it meets the vertical axis is . So the rate function is a re-parametrisation of the curve by its slopes, and the properties it inherits are exactly the ones a reader would want: it is convex, it vanishes at the mean and only there, and it increases away from the mean in both directions.
The upper bound half of Cramér’s theorem is three lines and worth having, because it is exact at every rather than a limit. For any , Markov’s inequality applied to gives
using independence to factor the expectation. Optimising over gives . That is Chernoff’s bound, and the figures assert it at every drawn : the exact probability never rises above the straight line.
The lower bound — that the exponential rate is not merely an upper bound but the truth — is the harder half, and the mechanism behind it is worth a picture of its own.
Tilting: making the rare event ordinary
Multiply each probability by and renormalise. The result is a genuine distribution — the exponential tilt — and at the that the maximisation picks out, its mean is exactly .
That is the whole idea. Under the tilted law, an average of is the typical outcome, and the law of large numbers says it happens with probability approaching one. So the rare event under the original law is the ordinary event under a nearby law, and the probability of the rare event is the cost of changing laws — which is the total reweighting, .
The cost per draw is the information distance between the two distributions, , and it equals exactly. The figure computes both and requires them to agree. That identity — rate function equals information distance to the nearest law with the required mean — is the form of the theorem that generalises furthest, and it is why the subject and information theory share a vocabulary.
What the numbers do
Two examples make the difference between the regimes concrete, and both are computed rather than quoted.
For the coin at level , the rate is
while the normal approximation’s exponent is . At these are and — close, which is why moderate deviations behave nearly normally. At they are and ; the ratio of the two answers is , which at sixty-four draws is a factor of about ten. And as approaches the rate goes to while the normal exponent goes to : the true probability of all heads is , and the normal approximation is out by an exponentially large factor.
The general shape of the discrepancy is worth stating in one sentence. The normal approximation is a quadratic approximation to the rate function at the mean, so it is accurate for deviations that shrink with and increasingly wrong for those that do not — the curvature of at is , which is precisely what makes the two agree to second order and disagree beyond it.
Where the theorem needs its hypothesis
Everything above assumed that is finite for some positive : the distribution has an exponentially decaying tail. Without that, there is no transform, no rate function, and — more importantly — no exponential decay to describe.
For a heavy-tailed distribution the mechanism of a large deviation changes completely, and the change has a slogan: the whole sum is large because one term is large. For a distribution whose tail falls like a power, the chance that is asymptotically times the chance that a single draw exceeds — the cheapest way to get a big sum is one big summand, rather than a conspiracy of many slightly enlarged ones. In the light-tailed case the opposite is true: every summand contributes a little, which is exactly what the tilt describes.
That contrast is the practical content of the theory. It says which explanation of a surprising outcome is the likely one, and the answer depends on the tail: many small pushes, or one shove.
The shape of the whole picture
Putting the two regimes side by side gives a single statement about where the probability of an average sits, and it is worth writing out because it explains why the subject has three names for three parts of one curve.
Write the average’s departure from the mean as . For of order the probability is described by the bell curve and the answer is of order one. For of order one the probability is and the answer is exponentially small. In between — departures growing with but slower than — is the moderate deviations regime, where the answer is and may be replaced by its quadratic approximation , because the departure is still small enough that the curvature at the mean is all that matters.
So the three regimes are one formula seen at three magnifications, and the normal approximation is the parabola that osculates the rate function at its minimum. That is the cleanest way to hold the relationship in mind, and it makes the failure mode obvious: using the parabola where the true curve has bent away from it.
There is a second reason the two halves cannot be merged into one theorem. The central limit theorem needs only a finite variance; Cramér’s theorem needs an exponential moment, which is a far stronger hypothesis. Between them lie distributions — those with finite variance but heavy tails, such as one whose tail falls like — for which the bell curve is the right answer in the middle and no exponential rate exists at all. The two theorems are not two views of one result; they are two results with different hypotheses that happen to describe the same random quantity.
Where it came from, and where it goes
Cramér worked on insurance risk, and the question was the probability of ruin: a reserve depleted by claims, and the chance that an unlikely run of them exhausts it. His 1938 paper gave the theorem for sums of independent variables; Chernoff’s bound came in 1952 in a statistical context; and the general theory — of large deviations for measures rather than for sums — was built by Donsker and Varadhan in the 1970s, with Sanov’s theorem as its combinatorial heart.
The physics connection is not an analogy. In statistical mechanics the number of configurations with a given energy is exponentially large, the entropy is the exponent, and the free energy is the Legendre transform of the entropy. Cramér’s theorem has the same three objects in the same relation, with as the free energy and as the entropy, and the tilted distribution is the Boltzmann distribution at the corresponding temperature — the same reweighting that turns a count of arrangements into a probability whenever a system is described by its energy. Which of the two subjects is a special case of the other is a matter of taste.
The applications outside both are large. The exponential rate is what makes hypothesis testing work — the error probabilities of a test fall exponentially in the sample size, with a rate given by an information distance, which is Stein’s lemma. It is what governs the failure probability of a procedure repeated many times, in the way that a random walk’s chance of straying far is governed by the same exponent. And it is the reason a concentration inequality is worth having at all: Chebyshev’s bound gives a probability falling like , and Chernoff’s gives one falling like , which is the difference between a useless guarantee and a usable one.
Where it fails, and what it costs
The rate is only the exponent. Cramér’s theorem gives , which determines the probability only up to factors growing slower than exponentially. The truth is typically , and the polynomial prefactor is why a measured rate at finite always sits above the limiting one — which the figures show and their assertions require.
Independence matters, but can be weakened. Sums of dependent variables have their own theory, and for Markov chains the rate function is an eigenvalue problem rather than a Legendre transform.
The level must be reachable. For a bounded summand the rate becomes infinite past the largest value the summand takes, correctly: the probability is exactly zero there. The figures guard against being asked for a level outside the support.
What the pictures cannot show
The logarithmic axis compresses a range of ten decades into a plot, and that is the only way the exponential can be drawn at all — but it also hides how small these numbers are. A probability of at sixty-four tosses and one of at a hundred and twenty-eight look like neighbouring points, and they differ by a factor of a billion.
The tangent construction is drawn for one distribution and three levels. What it cannot show is that the supremum is attained — that the tangent touches — which fails exactly when the level exceeds what the summand can produce, the case the guard refuses.
And no picture here shows the lower bound. Chernoff’s inequality is visible as the drawn probabilities staying under the line; the claim that they stay near it is supported by measuring the slope over the last doubling and comparing, which is evidence at the sizes drawn and not the theorem.
The ladder from here
Below: the bell curve from coin flips, which is the regime this essay is about leaving, and Chebyshev’s bound, which is the same question answered with only a variance and answered far more weakly. Sideways: how fast the bell arrives, which measures the error of the normal approximation in the region where it is valid, and the average that never settles, where no rate exists because no mean does. Above: Sanov’s theorem, where the object deviating is a whole distribution rather than a number, and the rate is again an information distance.
What is worth carrying away
The lesson is about the reach of a limit theorem. The central limit theorem is one of the most useful statements in mathematics and it describes a window of width , and the habit of treating it as a description of the whole distribution is the single commonest error in applied probability — it produces tail estimates that are wrong by exponential factors, always in the direction of underestimating the rare event when the tail is heavy and overestimating it when the tail is light.
The repair is to know which question is being asked. A question about typical fluctuations is a question and the bell curve answers it. A question about a fixed departure is an exponential question and needs a rate. And a question about a heavy-tailed quantity is neither, and is usually answered by asking which single term was large.
Named objects
A dashed tag is an object no other essay names yet.
Central limit theoremConcentration inequalityHeavy tailsLarge deviationsLegendre transformMoment generating functionRate functionTail bound