Concept

Tail bound

A statement of how much probability can lie beyond a given distance from the centre of a distribution. Its strength follows from its hypothesis: a variance gives a bound falling like one over the square of the distance, an exponential moment gives one falling exponentially.

Named by 8 essays across one field — each of them below, with the objects they name alongside it.

five values, unevenly weighted, and the mass outside 3 standard deviations. A distribution drawn as bars, with the windows one and a half, two and three standard deviations wide marked. The probability outside each window is summed and compared with the bound that knows only the variance.

How far from the average a thing can be

Knowing only an average and a spread — nothing about the shape, nothing about the number of outcomes, nothing about symmetry — the chance of landing three standard deviations out is at most one in nine. And there is a distribution that lands there exactly that often, so the bound cannot be improved.

probability · Concentration
The chance the average clears 0.75, against the number of draws. The exact probability that the average of n draws exceeds a fixed level, on a logarithmic scale, falling along a straight line whose slope is the rate function, with the normal approximation drawn beside it and diverging.

The tail is not a bell

The limit theorem describes a window of width one over the root of n around the mean; ask instead for the chance that an average lands a fixed distance away and the answer falls exponentially, at a rate computed from the summand before any n is chosen.

probability · Central limit
The 91 histograms 12 draws can produce. A triangle whose points are the possible histograms of a fixed number of draws over three faces, each drawn as a dot shaded by how far it is from the true distribution.

When the whole histogram deviates

A rare average has a price, an exponent that grows with the number of trials. Ask instead for the chance that the whole tally of outcomes comes out wrong, and the exponent is no longer a function of one number — it is a distance between two distributions, and every rare-average rate is a shadow of it.

probability · Central limit
The most mass 2 deviations out, with only a mean and a variance. The distribution putting as much probability as possible outside a window 2 standard deviations wide, found by searching every triple of support points, with the quadratic certificate that bounds it drawn over.

The bound is the answer to a search

Chebyshev's inequality is not a clever estimate that happens to be sharp. It is the exact answer to a maximisation over all distributions with a stated mean and variance, and the polynomial that proves nothing beats it is the certificate a search of that kind always produces.

probability · Concentration
What one input can do, over 1024 cases. A table of functions of several inputs with the largest effect any single input has on each, the bound that effect implies, and the true tail probability — computed by enumerating every input.

No single input can move it far

Independence was never the hypothesis doing the work. A quantity built from many separately drawn inputs concentrates whenever changing one of them moves it only a little — and that covers quantities which are not sums of anything and have no formula at all.

probability · Concentration
How far each estimate can stray, when rare jumps are possible. The error that the sample mean and the median of 12 block means stay within, at confidence levels from 80% to 99.9%, over 60000 repeated samples of 120 draws from normal draws with rare large jumps, with Chebyshev's guarantee for the mean.

The median of many small averages

Knowing only that a quantity has a finite spread, the plain average of n samples can be promised to within σ/√(nδ) with confidence 1 − δ, and no better — Chebyshev's bound is tight, and rare large jumps achieve it. Cut the same samples into a dozen blocks, average each block, and take the median of the averages, and the promise improves to within about σ√(log(1/δ)/n). Nothing about the data has been assumed beyond the spread; only the way of combining it has changed.

probability · Concentration
How much of a sphere lies near its equator. Curves of the share of the sphere within ε of the equator against ε, for spheres in 3, 10, 100, 1000 dimensions: a straight line in three dimensions, a near step in a thousand.

A sphere that is nearly all equator

On an ordinary globe, the band within a tenth of the radius of the equator holds a tenth of the surface. On a sphere in a thousand dimensions the same band holds 99.85% of it, and the band of a fifth holds all but about two parts in ten billion. Almost every point of a high-dimensional sphere is near every equator at once — and so any function that cannot change quickly is, over almost all of the sphere, almost constant.

probability · Concentration
Tails of a sum of rare events: exact, and three bounds. 10: exact 5.42e-1, Bennett 1.00e+0, Bernstein 1.00e+0, Hoeffding 1.00e+0; 15: exact 8.34e-2, Bennett 3.39e-1, Bernstein 3.42e-1, Hoeffding 9.95e-1; 20: exact 3.44e-3, Bennett 2.09e-2, Bernstein 2.35e-2, Hoeffding 9.80e-1; 25: exact 4.64e-5, Bennett 3.66e-4, Bernstein 5.50e-4, Hoeffding 9.56e-1; 30: exact 2.46e-7, Bennett 2.34e-6, Bernstein 6.10e-6, Hoeffding 9.23e-1; 35: exact 5.88e-10, Bennett 6.45e-9, Bernstein 3.92e-8, Hoeffding 8.82e-1; 40: exact 7.03e-13, Bennett 8.70e-12, Bernstein 1.67e-10, Hoeffding 8.35e-1; 45: exact 4.56e-16, Bennett 6.27e-15, Bernstein 5.21e-13, Hoeffding 7.83e-1; 50: exact 1.71e-19, Bennett 2.59e-18, Bernstein 1.27e-15, Hoeffding 7.26e-1.

Charged for the variance, not the range

Hoeffding's inequality knows one thing about each term of a sum: the interval it lies in. For ten thousand coins that each land heads once in a thousand, that makes it promise almost nothing — a 92% chance of thirty heads, when the truth is two in ten million. Tell the bound each term's variance as well and it changes character: Bernstein's and Bennett's inequalities decay like the normal curve while the deviation is small and like a Poisson tail beyond, and for rare events they are millions of times sharper.

probability · Concentration

Named alongside it

The objects these essays reach for when they reach for this one.

Concentration inequalityVarianceExpectationExtremal exampleHeavy tailsCentral limit theoremConvergence rateExhaustive searchHoeffding inequalityLarge deviationsMoment generating functionNormal distribution

All concepts