Multiplying makes the digit one common
Worth reading first: The average settles and the wobble does not · How fast the bell arrives.
In 1881 the astronomer Simon Newcomb noticed that the books of logarithm tables in his library were dirtier at the front than at the back. The pages for numbers beginning with 1 were worn; the pages for numbers beginning with 9 were nearly clean. He concluded that the numbers people looked up began with small digits more often than large ones, and wrote down the law that would fit: the leading digit is with probability
which is 30.1% for 1, 17.6% for 2, and falls to 4.6% for 9. The physicist Frank Benford rediscovered it in 1938, checked it against twenty thousand numbers from river areas to street addresses, and his name stuck.
The law is strange because it seems to need a reason in every dataset where it holds, and no single dataset supplies one. This essay gives the reason for one large class of numbers: those that come from multiplying. It is the central limit theorem, applied to the logarithm.
A single number drawn uniformly between 0 and 1 begins with each digit equally often, 11.1% each: nothing about it favours 1. The product of two such numbers already begins with 1 about 24% of the time and with 9 only 3.4%. The product of three is at 30.1% ones and 4.2% nines. The product of six is within a tenth of a percentage point of the law across all nine digits. None of the factors knows anything about the law, and every product converges to it. The figures compute these probabilities exactly, not by simulation, so the convergence they show is the true one.
A digit is a fractional part
The leading digit of a positive number is decided by one thing: where falls between two consecutive whole numbers. The number 347 has ; the whole part, 2, says how many digits it has, and the fractional part, 0.540, says what they are — . The leading digit is 1 exactly when the fractional part lies between 0 and , is 2 when it lies between 0.301 and , and so on up to 9, between and 1.
So Benford’s law is a statement about the fractional part of the logarithm: it holds exactly when that fractional part is uniform between 0 and 1. If it is uniform, the chance of the digit 1 is the length of the interval from 0 to 0.301, which is , and the chance of is . The law is nothing more than a uniform distribution, seen through a logarithm.
That reformulation explains the old observation that Benford’s law is the only digit law that does not depend on units. Measuring a quantity in feet instead of metres multiplies every value by 3.28, which shifts every logarithm by the same amount, which rotates the fractional parts round the interval. A uniform distribution on a circle is unchanged by rotation; no other distribution is. A law of leading digits that is the same in every system of units must be Benford’s.
The logarithm of a product spreads out
For a product, the logarithm is a sum: . If the factors are independent, the logarithm of the product is a sum of independent terms, and the average settles while the wobble does not: its centre moves in proportion to , and its spread grows in proportion to . For uniform factors, the logarithm of one factor has a standard deviation of 0.43 powers of ten; the product of 4 has 0.87, of 16 has 1.74, and of 64 has 3.47.
The left of the figure shows the densities of those logarithms, each settling into the bell shape as it spreads. The vertical lines mark the whole numbers, the boundaries between powers of ten. A single factor’s logarithm is crammed into the last power of ten below 1. By 16 factors the bell covers about ten powers of ten, and by 64 more than twenty.
The right of the figure takes each density and wraps it onto one power of ten: it cuts the density at every whole number and stacks the pieces, which gives the distribution of the fractional part. For one factor that wrapped density is far from flat, which is why one uniform factor has no preference for 1. For four it is already close to the flat line. By sixteen it is within of flat, and by sixty-four the difference is below what the arithmetic can resolve. A smooth bell that spans many units, cut into unit lengths and stacked, is flat: each slice contributes a nearly straight piece, and the pieces from the rising and falling sides cancel each other’s slopes.
That is the whole mechanism. The central limit theorem guarantees the spreading, the spreading guarantees the flattening, and the flattening is Benford’s law.
How wide is wide enough
The flattening can be put as a single rule of thumb. If the logarithm of a quantity, measured in powers of ten, has a bell-shaped distribution with standard deviation , then its fractional part differs from uniform by a ripple of height about . The in the exponent makes the ripple collapse astonishingly fast. At — a quantity whose values mostly lie within a factor of two either side of a typical value — the ripple is about 0.34, and the leading digits miss Benford’s law by about 9 percentage points in total. At the ripple is 0.014 and the miss is 0.4 points. At , a spread of one power of ten either way, the ripple is .
So the practical criterion is not that data span many orders of magnitude, as is often said, but that their logarithms are spread over about one: a bell for that is half a unit wide is already indistinguishable from Benford’s law in any dataset of ordinary size. The heights of adult humans fail badly, varying by perhaps 10% either way. Lengths of rivers, populations of towns and the masses of stars pass easily.
The same exponent explains why the products in the first figure arrive so quickly. Three uniform factors give a logarithm with standard deviation 0.75 powers of ten, at which the ripple is already down to ; the remaining distance comes from the bell not yet being a bell, the skew that how fast the bell arrives measured, and that disappears too as factors are added.
The approach is geometric
How quickly the flattening happens can be computed exactly, and the answer is simpler than the theorem it comes from.
A distribution on a circle is uniform exactly when all of its Fourier coefficients but the constant vanish, and the deviation from uniform is controlled by how large those coefficients are. For the fractional part of , the first coefficient is with — the characteristic function of , evaluated at one point. For a product of independent factors, the expectation of a product is the product of expectations, so the coefficient for factors is the coefficient for one factor raised to the -th power.
For a uniform factor, , whose size is . Each extra factor multiplies the remaining ripple by 0.344, and the figure confirms it: the exact distance from Benford’s law falls along a straight line on the logarithmic scale, by a factor of 0.344 per factor, reaching at 24 factors. The central limit theorem only says the logarithm spreads; the coefficient says precisely how fast the spreading erases the digits’ memory of the factors.
The dice do something different. Their first coefficient is 0.317, smaller than a uniform factor’s, and the first few dice close in on the law as fast. Then they stall near a distance of 0.01 and stay there for twenty more dice. The reason is arithmetic. A die’s faces are made from the primes 2, 3 and 5, so the logarithm of a product of dice lies on a lattice built from , and . And is nearly a whole number, so the tenth Fourier coefficient — a fine ripple, ten waves to each power of ten — is nearly unaffected by multiplying by 2 or 4 or 5, and survives. It does decay, but at 0.782 per die instead of 0.317, and its effect on the leading digit, though small, is what holds the distance at 0.01. The central limit theorem’s spreading is necessary, but whether the fractional part becomes uniform depends on whether the factors’ logarithms can avoid a lattice.
One number per kind of factor
Because the coefficient for factors is the coefficient for one factor to the -th power, each kind of factor has a single number that says how fast it leads to Benford’s law. The figure compares four. A number between 1 and 2 has a logarithm confined to the first third of one power of ten; it smooths almost nothing per factor, 0.861, and needs 31 factors to bring the first ripple below 1%. A uniform number between 0 and 1 and a die are near 0.33, and need 5. An exponential waiting time, whose logarithm already spreads over about half a power of ten, has a coefficient of 0.057 and needs only 2.
The number is computed from a single factor before any product is formed, as the exponential rate in the tail is not a bell was computed from a single summand before any sum. The rule hidden in these numbers is that a factor smooths the digits in proportion to how widely its own logarithm is spread, measured in powers of ten — which is the same thing the central limit theorem cares about, seen one factor at a time. It also identifies the one way to fail. If is confined to a lattice of points apart for some whole number , then the -th coefficient has size exactly one and never decays. The extreme case is multiplying by 10, which never changes a leading digit at all. Short of that, every factor leads to Benford’s law, and the only question is how fast.
Sums do not get there
The same theorem that makes products Benford makes sums the opposite. The sum of ten dice has mean 35 and standard deviation about 5.4. The bell that the sum settles into is narrow compared with the sum’s size: almost all of it lies between 25 and 45, and 64% of sums begin with 3. The leading digit of a sum is decided by where its centre is, and the central limit theorem, which narrows a sum relative to its size, makes that decision sharper as more terms are added. Forty dice sum to about 140, and the leading digit is 1 almost every time — not because of Benford, but because every sum lands between 100 and 199.
The product of the same ten dice has a logarithm that is also a sum, with a standard deviation of 0.83 powers of ten — wide on the scale that matters for digits — and its leading digits are within 1.2 percentage points of the law. The distinction is between adding things and multiplying them, and it is visible in data. Quantities that grow by percentages — prices, populations, the sizes of cities, physical quantities built from products of others — tend to follow Benford’s law. Quantities that accumulate by adding similar amounts — heights of adults, scores in a test, the sum of a shopping basket of similar items — do not.
This is also why Benford’s law is used to look for fabricated accounts. Genuine financial figures arise from growth, interest, quantities times prices: multiplicative processes, with logarithms spread over several powers of ten. Numbers invented by people tend to have their leading digits spread too evenly, or clustered near a few favoured values. A set of expense claims whose first digits are uniform is not proof of fraud, but it is a set whose numbers did not come from the process they claim to.
Powers of two, with no randomness
The last figure removes randomness entirely. Of the first thousand powers of two, 301 begin with 1, 176 with 2, and 45 with 9 — within 0.31 percentage points of Benford’s law in total. There is no product of random factors here, only the same factor, 2, used again and again.
The mechanism is the circle again. The fractional part of is the fractional part of , a point stepping round a circle by a fixed irrational amount each time. A point stepping round a circle by an irrational angle never returns to where it started, and Hermann Weyl proved in 1916 that it visits every arc in proportion to the arc’s length: the fractional parts become uniform. That is exactly the condition for Benford’s law, reached by a rotation where products reached it by spreading. Powers of 3, of 7, of any number that is not a rational power of 10 behave the same way. Powers of 10 do not, since their fractional part is always 0, which is the lattice failure from the last section in its most complete form.
The same rotation is behind three gaps and no more, which counted the gaps that such points leave on the circle. A rotation’s even spreading is quick in that sense, and the leading digits of inherit it: their distance from the law shrinks roughly like one over the number of powers counted, up to a logarithmic factor, much faster than the one-over-square-root rate a random sample would show. The powers of two are, in that precise sense, more even than random.
Two routes to one circle
Benford’s law arises when the fractional part of a logarithm is uniform, and the figures show two unrelated ways that happens. Multiplying independent random factors spreads the logarithm by the central limit theorem, and a widely spread distribution wraps to a flat one at a geometric rate set by one coefficient per factor. Multiplying by the same factor over and over rotates the logarithm by an irrational step, and an irrational rotation fills the circle evenly. In both cases the law appears because something has made the fractional part forget where it started, and in both cases it fails exactly when the logarithms are confined to a lattice.
That is why the law appears in so many unrelated tables. Any quantity produced by a long enough chain of multiplications — interest compounding, a population growing, a physical constant built from a dozen others — has had its logarithm spread or rotated, and its first digits follow the law whatever the details. Newcomb’s worn pages were the tables of the numbers his colleagues needed, and those numbers had come, mostly, from multiplying.
Still open: which sequences are Benford
For sequences built from both addition and multiplication, whether they follow Benford’s law is mostly unknown. Sequences of pure multiplication are well understood. The powers of 2 follow the law because is irrational, which is easy. The Fibonacci numbers follow it because they grow like powers of the golden ratio, and the factorials because their logarithms grow smoothly enough for a theorem about smooth sequences to apply. Even the powers at prime exponents follow it, though the proof needs Ivan Vinogradov’s 1937 theorem on how multiples of an irrational number fall along the primes.
Mixing addition in breaks the analysis. The iteration that sends to when is even and to when it is odd multiplies and adds in turn, and Jeffrey Lagarias and Kannan Soundararajan proved in 2006 that for almost every starting value its iterates’ leading digits approach Benford’s law. For any particular starting value nobody can prove it, and the question is tangled with the unsolved question of whether every such iteration reaches 1.
The probabilistic side has its own open edge. Theodore Hill proved in 1995 that if one draws numbers from a randomly chosen distribution, in a way that does not favour any scale, the result follows Benford’s law — which explains why tables mixing many sources tend to obey it. What is not settled is a usable criterion for when a given real dataset should be expected to follow the law and to what accuracy, which is what anyone using it to detect fabricated figures actually needs. The convergence rates in the figures here are exact for the factors drawn; for data whose generating process is unknown, the rate is the thing that is missing.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A bounce is a fold of the table — both name equidistribution, irrational rotation
- When two circular motions come home — both name equidistribution, irrational rotation
Named objects
A dashed tag is an object no other essay names yet.
Central limit theoremCharacteristic functionEquidistributionFourier seriesIrrational rotationLogarithm