A sphere that is nearly all equator
Worth reading first: No single input can move it far · The median of many small averages.
The concentration inequalities met so far have been about sums and functions of many independent inputs: an average lands near its mean, and a function no single input can move far lands near its median. There is a geometric version of the same phenomenon, and it is the most striking of them because nothing in it is random at the start. It is a fact about the shape of a sphere in many dimensions, and it says that almost all of that shape is somewhere it seems it has no business being.
Take the unit sphere in dimensions — all points at distance one from the origin — and any equator: the points whose first coordinate is zero. Ask what share of the sphere’s surface lies within a distance of that equator. On the ordinary sphere, in three dimensions, the answer is exactly . In a thousand dimensions, with , it is . The equator is not special — every equator gets the same share — so almost every point of a high-dimensional sphere is close to every equator at once. This is called concentration of measure, and it was first noticed, in essentially this form, by Paul Lévy around 1920.
The share of a sphere near its equator
In three dimensions the curve is a straight line, and that is a famous theorem. Archimedes proved that a band cut from a sphere by two parallel planes has the same area as the band they cut from the cylinder wrapped round the sphere — the hat-box theorem — so the area of a band depends only on its width, and a band of width holds a share of the sphere. The figure checks the three-dimensional curve against the straight line at every point it draws.
In ten dimensions the curve already bends: the band of half-width holds of the sphere, not . In a hundred dimensions it holds , and in a thousand, . The band of half-width holds in ten dimensions, in a hundred, and all but in a thousand. As the dimension grows, the curve turns into a step: nothing, then everything, with the jump at a width of order .
One coordinate of a random direction
The band’s share is the chance that the first coordinate of a point chosen uniformly on the sphere is between and . So the figure is really a picture of how one coordinate of a random unit vector is distributed.
In three dimensions it is flat: every value from to equally likely, which is the hat-box theorem in the language of probability. In dimensions its density is proportional to — the factor records how much room the remaining coordinates have when the first is fixed at — and as grows that factor falls off steeply away from zero. By a hundred dimensions the distribution is within of a normal curve with standard deviation .
The reason is a counting argument that needs no calculus. The squares of the coordinates of a unit vector add up to one. By symmetry they have the same average, so each has average square , and each coordinate is typically of size . A random unit vector in a thousand dimensions has coordinates around each — every coordinate is small, although together they make a vector of length one. That is the whole of the concentration near the equator: the first coordinate is small because every coordinate is.
The normal curve is not an accident either. A uniform random point on the sphere can be produced by drawing independent normal numbers and dividing by the length of the resulting vector. In high dimension that length is almost exactly — it is the square root of a sum of many independent squares, which concentrates — so each coordinate is almost exactly a normal number divided by . The sphere in high dimension is, to a very good approximation, a cloud of independent normal coordinates scaled down.
The same computation, read in physics, is the origin of the bell curve of molecular speeds. A gas of molecules with a fixed total energy has its velocity components constrained to a sphere: the sum of their squares is fixed by the energy. If every point of that sphere is equally likely — the assumption of statistical mechanics — then any one velocity component is distributed like one coordinate of a random point on a sphere in dimensions, and with in the billions of billions that is a normal distribution to any accuracy anyone could measure. The Maxwell distribution of velocities, which Maxwell derived in 1860 from assumptions about collisions, is the shadow of a very high-dimensional sphere on one of its axes; Henri Poincaré and Émile Borel made the observation precise, and it is sometimes called the Poincaré–Borel lemma. The bell curve that coin flips build and the bell curve a sphere casts are the same curve for the same reason: many small independent contributions, constrained or summed.
Two random directions are nearly perpendicular
A consequence that sounds more surprising than it is: pick two directions at random in high dimension, and they are almost certainly almost at right angles.
The cosine of the angle between two unit vectors is their dot product, and by symmetry it is distributed exactly like the first coordinate of a random unit vector: fix the first vector as the north pole, and the second one’s first coordinate is the cosine. So the cosine is of size , and the angle is within a few multiples of radians of . In three dimensions only of random pairs are within ten degrees of perpendicular; in thirty, ; in three hundred, essentially every pair.
That is why high-dimensional space has room for exponentially many nearly perpendicular directions — far more than exactly perpendicular ones — and it is the geometric fact behind the Johnson–Lindenstrauss lemma, which says that any set of points in high dimension can be projected onto a random subspace of dimension only about with all their distances nearly preserved. Random directions do not interfere with each other, because they are almost orthogonal, and that is the resource random projections spend.
Functions that cannot change quickly are nearly constant
The equator is only one set, and the first coordinate only one function. The real content of concentration of measure is that the same thing happens for every function that cannot change quickly.
Call a function on the sphere 1-Lipschitz if it never changes by more than the distance moved: . The first coordinate is such a function. So is the distance to any fixed point, the largest coordinate, and the sum of the sizes of the coordinates divided by . The theorem is that every such function, evaluated at a random point of a high-dimensional sphere, is within a few multiples of of its median with overwhelming probability — the same scale as for the first coordinate, whatever the function.
The function in the figure has nothing to do with any equator. Its values in ten dimensions spread over a range of several tenths; in a thousand dimensions they are within about of , the value the normal-coordinate picture predicts, since each coordinate’s size averages . The spread shrinks by a factor of about three each time the dimension grows tenfold — like , as the theorem says. A function that is free to vary across the whole sphere in principle takes almost exactly one value on almost all of it.
This is the geometric twin of the bounded-differences inequality. There, a function of many independent inputs that no single input could move far was nearly constant; here, a function of a point on a sphere that no small movement can change much is nearly constant. Both are consequences of one phenomenon — many small independent influences, whether coordinates or inputs, averaging each other out — and in high dimensions a single point is already many small influences. It is also the reason the median of many small averages works: a median is a Lipschitz summary of its inputs, and the concentration of a count of failed blocks is the same averaging again.
Lévy’s inequality: fatten half and get everything
The sharpest form of the theorem is an isoperimetric statement, and it explains why the equator is the right thing to have looked at. On the ordinary plane, the circle encloses the most area for a given fence, and equivalently, among all sets of a given area, the disc grows least when fattened by a margin . Lévy proved the analogue on the sphere: among all subsets of the sphere of a given measure, a spherical cap grows least when fattened by . A hemisphere is a cap of measure one half, so every set holding half the sphere, fattened by , holds at least as much as a fattened hemisphere does — which is the band computation from the first figure, extended to one side.
The leftover is bounded by , and the figure computes the exact leftover and confirms the bound at every point drawn. In a thousand dimensions, fattening a hemisphere by leaves out less than of the sphere. Combined with isoperimetry this gives the Lipschitz theorem at once: for a 1-Lipschitz function with median , the set where holds half the sphere, its -fattening is contained in the set where , and so exceeds on a share at most . Every half of a high-dimensional sphere, however it is shaped, is within a short distance of almost everything.
A ball that is almost all skin
The sphere’s surface concentrates near every equator; the ball it bounds concentrates in a second way, near its surface. Shrinking the ball’s radius by a factor shrinks its volume by , so the shell between radius and holds a share of the volume. In three dimensions a shell one hundredth thick holds about of the ball. In a thousand dimensions it holds , which is . A high-dimensional orange is almost all peel.
The two concentrations together describe where a uniformly random point of a high-dimensional ball is: almost certainly within a hair of the surface, and almost certainly within a hair of every equator through the centre. There is no contradiction — the region near the surface and near a given equator is still most of the surface — but it means that the ball’s “interior”, in any intuitive sense, is empty. Nearly all of its volume, and nearly all of its surface, lives in a thin band that no picture of a ball in three dimensions suggests.
The cube, which is secretly round
The cube shows the same phenomenon in a form that looks paradoxical. The cube has corners at distance from its centre and faces at distance , so it seems very far from round: in a thousand dimensions the corners are thirty-two times further out than the faces. Yet a point chosen uniformly in the cube is at distance very nearly from the centre — its squared distance is a sum of independent squares, each averaging , and a sum of many independent pieces concentrates. So almost all of the cube’s volume lies in a thin shell at radius , far from both the faces and the corners.
The high-dimensional cube is best pictured as a ball of radius with extremely thin spikes reaching out to the corners — spikes that are long but hold almost no volume. It is one reason that regular polytopes in high dimension are so few and so strange: the cube, the cross-polytope and the simplex are the only ones, and all three look, from the point of view of their volume, like spheres with a few peculiar extremities.
Why the curse of dimensionality is a blessing too
Concentration of measure is one face of what is usually called the curse of dimensionality. Most of a high-dimensional ball’s volume lies near its surface, most of a cube’s volume lies near its corners, and grids of points become hopelessly sparse — which is why grids fail and random points succeed at computing integrals in many dimensions. Random sampling works there precisely because of concentration: the average of a function over random points concentrates on the function’s mean, and it concentrates at a rate that does not depend on the dimension.
The same fact makes high-dimensional statistics both hard and possible. It is hard because distances lose contrast — when every pair of random points is at nearly the same distance, as they are on a high-dimensional sphere, “nearest neighbour” means little. It is possible because averages and Lipschitz summaries are extraordinarily stable, so that quantities computed from high-dimensional data can be trusted even when the data themselves are too spread out to picture. Vitali Milman, who turned Lévy’s observation into a central tool of geometry in the 1970s, called it the phenomenon that makes high-dimensional objects look simple from far enough away.
What the curves cannot show
They cannot show the sphere. A sphere in a thousand dimensions cannot be drawn; the figures draw distributions of numbers computed from it — a coordinate, an angle, a function value — and those distributions are what the theorems are about. The band shares are computed exactly from the density of one coordinate; the angles and function values are sampled, with the sample sizes stated.
They cannot show every Lipschitz function. The theorem is about all of them at once, with the same ; the figure shows one function. That the same bound holds for every set of half the sphere is Lévy’s isoperimetric inequality, quoted, and its proof — by symmetrisation, pushing mass towards a cap without increasing the fattened measure — is not drawn.
And they cannot show what happens away from the sphere. Concentration holds on the sphere, on the cube with the uniform measure, for Gaussian space, and for many other spaces, with different constants; it fails on spaces that are “thin” in some direction, and the figures show only the sphere, where the constants are exact.
Still open: the right constant for every convex body
On the sphere the concentration is completely understood. For a general convex body in high dimension — a cube, a simplex, the ball of some unusual norm — the question of how concentrated its volume is near its “equators” has a precise conjecture attached. The Kannan–Lovász–Simonovits conjecture, from 1995, says that every convex body, suitably normalised, is at least as concentrated as a Gaussian up to a universal constant, in the sense that its thinnest cut is controlled by its variance. After two decades of slow progress, Yuansi Chen proved in 2021 that the constant grows more slowly than any power of the dimension, and further work has since reduced it to a polylogarithm; whether it is bounded by a constant independent of the dimension, as conjectured, remains open.
Near every equator at once
On the unit sphere in dimensions, the share of the surface within of any equator tends to one as grows, and at a width of order . In three dimensions the share is exactly by Archimedes’ hat-box theorem; in a thousand, a band of half-width holds . The reason is that every coordinate of a random unit vector is of size , and each is nearly normal.
The same concentration makes two random directions nearly perpendicular, makes every 1-Lipschitz function of a random point nearly constant, and, by Lévy’s isoperimetric inequality, makes every set holding half the sphere within a small distance of almost all of it, with the leftover bounded by . It is why random sampling works in high dimension and why high-dimensional data can be summarised stably.
In high dimension, the typical is overwhelmingly typical — which is why a single random sample of a high-dimensional thing is often enough to know the whole.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A ball whose outside is not one — both name dimension, sphere
- Two pieces, in every dimension — both name dimension, sphere
- Two right angles and the diagonal of a box — both name dimension, orthogonality
- Zero in four dimensions — both name dimension, sphere
Named objects
A dashed tag is an object no other essay names yet.
Concentration of measureCurse of dimensionalityDimensionIsoperimetric inequalityLipschitz functionNormal distributionOrthogonalitySphereTail bound