Number

The race that makes ζ(3) irrational

Roger Apéry's 1978 proof that the sum of the reciprocal cubes is not a fraction comes down to a race between two numbers. A whole-number multiplier grows by a factor of ten every 0.79 steps; the gap it multiplies shrinks by a factor of ten every 0.65. The gap wins, by a margin of about seventeen per cent, and that margin is the whole proof.

Worth reading first: An integral that cannot be a whole number · A tail too small to be a whole number.

A tail too small to be a whole number proved that ee is not a fraction by producing a quantity that would have to be a positive whole number if ee were a fraction and showing that it is less than one. An integral that cannot be a whole number did the same for π\pi with an integral, and ended on the number that held out longest against the method: ζ(3)=1+1/8+1/27+1/64+⋯\zeta(3) = 1 + 1/8 + 1/27 + 1/64 + \cdots, the sum of the reciprocal cubes. Euler had found the sum of the reciprocal squares in 1735, π2/6\pi^2/6, and the reciprocal cubes resisted everything for two hundred and forty years.

In June 1978 Roger Apéry announced a proof that ζ(3)\zeta(3) is irrational, and the audience did not believe it. The proof rested on two sequences defined by a recurrence nobody had seen, on claims about them that Apéry stated without explanation, and on arithmetic that looked miraculous. Within a few months Henri Cohen, Hendrik Lenstra and Alf van der Poorten had checked every step, and van der Poorten’s account of the episode, “A proof that Euler missed”, made the argument famous.

This essay runs the argument as a computation, in exact rational arithmetic, to see what the miracle consists of. It turns out to be a race between two numbers — one growing and one shrinking — and the first figure is the race.

Apéry's proof for ζ(3), as three numbers. log₁₀ of 2·lcm(1, …, n)³, of |a(n)ζ(3) − b(n)| and of their product for n up to 60: slopes 1.262, -1.546, -0.284 per step.
Fig. 1 Apéry’s two sequences for ζ(3) to n = 60, on a logarithmic scale: the whole-number multiplier 2·lcm(1, …, n)³ that clears every denominator, the gap a(n)ζ(3) − b(n), and their product, which would have to be a non-zero whole number divided by a fixed denominator if ζ(3) were a fraction. The product falls to nought.

The shape of every such proof

Every proof of this kind has the same skeleton. Suppose the number xx is a fraction p/qp/q. Find whole numbers AnA_n and BnB_n such that Anx−BnA_n x - B_n is never zero. If x=p/qx = p/q, then q(Anx−Bn)=Anp−Bnqq(A_n x - B_n) = A_n p - B_n q is a whole number, and since it is not zero it is at least one in size. So ∣Anx−Bn∣≥1/q|A_n x - B_n| \ge 1/q for every nn. If one can also show that Anx−BnA_n x - B_n tends to nought, the assumption is contradicted, and xx is irrational.

The difficulty is entirely in finding AnA_n and BnB_n. They must be whole numbers, and the combination Anx−BnA_n x - B_n must be small — smaller than any fixed 1/q1/q — while not being exactly zero. For ee the series supplies them: An=n!A_n = n! and BnB_n the partial sum times n!n!. For π\pi an integral against a polynomial does. For ζ(3)\zeta(3) nothing obvious does, because its series has denominators k3k^3 that do not cancel in any helpful way, and the partial sums, multiplied out, have denominators far too large.

Apéry’s contribution was a pair of sequences in which the smallness and the wholeness are both available, though neither is obvious. The smallness costs nothing: the gap between the two sequences is tiny. The wholeness costs a multiplier, and the race is between how fast the multiplier grows and how fast the gap shrinks.

Two sequences and one recurrence

Apéry’s sequences both satisfy the recurrence

n3un=(34n3−51n2+27n−5) un−1−(n−1)3un−2.n^3 u_n = (34n^3 - 51n^2 + 27n - 5)\,u_{n-1} - (n-1)^3 u_{n-2}.

Started from a0=1a_0 = 1, a1=5a_1 = 5, it produces 1, 5, 73, 1445, 33001, 819005, … — and a first surprise is that these are all whole numbers, although the recurrence divides by n3n^3 at every step. They are the sums an=∑k(nk)2(n+kk)2a_n = \sum_k \binom{n}{k}^2 \binom{n+k}{k}^2, and the computation behind the figures checks that identity for every nn up to sixty. Started from b0=0b_0 = 0, b1=6b_1 = 6, the same recurrence produces fractions: 0, 6, 351/4, 62531/36, 11424695/288, …

The first terms of Apéry's sequences. 0 | 1 | 0 | — | 0; 1 | 5 | 6 | 1.200000000000 | 12; 2 | 73 | 351/4 | 1.202054794521 | 1404; 3 | 1445 | 62531/36 | 1.202056901192 | 750372; 4 | 33001 | 11424695/288 | 1.202056903158 | 137096340; 5 | 819005 | 35441662103/36000 | 1.202056903160 | 425299945236; 6 | 21460825 | 20637706271/800 | 1.202056903160 | 11144361386340
Fig. 2 The first terms of Apéry’s sequences, computed exactly from the recurrence. The a(n) are whole numbers, the b(n) are fractions, their ratio approaches ζ(3), and multiplying b(n) by twice the cube of lcm(1, …, n) always gives a whole number.

The ratio bn/anb_n/a_n approaches ζ(3)=1.2020569031…\zeta(3) = 1.2020569031\ldots astonishingly quickly: 1.2, then 1.20205, then 1.2020569, gaining more than one and a half decimal places at every step, so that by n=6n = 6 it is right to eleven places. The reason is that ana_n grows like (1+2)4n(1 + \sqrt 2)^{4n}, about 34 times per step, and the gap anζ(3)−bna_n \zeta(3) - b_n shrinks like (2−1)4n(\sqrt 2 - 1)^{4n}, about 34 times per step the other way, so the ratio’s error shrinks like (2−1)8n(\sqrt2 - 1)^{8n}.

The second surprise is in the denominators of bnb_n. Multiply bnb_n by 2 lcm(1,2,…,n)32\,\mathrm{lcm}(1, 2, \ldots, n)^3 — twice the cube of the least common multiple of the numbers up to nn — and the result is always a whole number. That is the statement the audience in 1978 found hardest to accept, and it is checked here for every nn to sixty. With it, An=2 lcm(1,…,n)3anA_n = 2\,\mathrm{lcm}(1, \ldots, n)^3 a_n and Bn=2 lcm(1,…,n)3bnB_n = 2\,\mathrm{lcm}(1, \ldots, n)^3 b_n are whole numbers, and the skeleton applies.

The race

Now the skeleton needs Anζ(3)−BnA_n \zeta(3) - B_n to tend to nought, and that is a product: the multiplier 2 lcm(1,…,n)32\,\mathrm{lcm}(1, \ldots, n)^3 times the gap anζ(3)−bna_n \zeta(3) - b_n. The first figure plots both and their product on a logarithmic scale. The multiplier grows: the least common multiple of the numbers up to nn is about ene^n by the prime number theorem, so its cube is about e3ne^{3n}, a factor of ten every 0.79 steps on the measured stretch. The gap shrinks like (2−1)4n=e−3.5255n(\sqrt 2 - 1)^{4n} = e^{-3.5255n}, a factor of ten every 0.65 steps. Since 3.5255>33.5255 > 3, the gap shrinks faster than the multiplier grows, and the product falls — by a factor of ten every three and a half steps, to below 10−1910^{-19} at n=60n = 60.

So for every fraction p/qp/q, the product would have to stay above 1/q1/q, and it does not: whatever qq is, the product eventually falls below it. The gap is also never exactly zero: it equals a positive integral, as the account of where the sequences came from explains below. And that is the entire proof. All the difficulty lies in the two surprising facts — the whole-number ana_n and the bounded denominators of bnb_n — and in the exponent 3.52553.5255 beating 33.

The margin is about seventeen and a half per cent, and it is worth seeing how narrow that is. If the least common multiple grew like e1.18ne^{1.18n} instead of ene^n, the multiplier would grow faster than the gap shrinks and the argument would prove nothing. The race is won, but not by much, and nobody has found a way to widen it for ζ(3)\zeta(3) or to run an equivalent race for its neighbours.

The same route for ζ(2)

Apéry gave the same kind of argument for ζ(2)=π2/6\zeta(2) = \pi^2/6, which was already known to be irrational because π\pi is. The recurrence is shorter,

n2un=(11n2−11n+3) un−1+(n−1)2un−2,n^2 u_n = (11n^2 - 11n + 3)\,u_{n-1} + (n-1)^2 u_{n-2},

with whole numbers an=∑k(nk)2(n+kk)a_n = \sum_k \binom{n}{k}^2 \binom{n+k}{k} and fractions bnb_n whose denominators divide lcm(1,…,n)2\mathrm{lcm}(1, \ldots, n)^2.

Apéry's proof for ζ(2), as three numbers. log₁₀ of lcm(1, …, n)², of |a(n)ζ(2) − b(n)| and of their product for n up to 60: slopes 0.841, -1.055, -0.214 per step.
Fig. 3 The same three numbers for ζ(2): the multiplier lcm(1, …, n)², the gap a(n)ζ(2) − b(n), and their product, to n = 60. The gap shrinks like the fifth power of the golden ratio’s reciprocal, and the product falls to nought again.

Here the multiplier grows like e2ne^{2n} and the gap shrinks like φ−5n=e−2.406n\varphi^{-5n} = e^{-2.406n}, where φ\varphi is the golden ratio — a second, independent appearance of a famous algebraic number in what is a statement about a sum of reciprocals. The product falls, by a factor of ten every four and a half steps, and the margin is about twenty per cent. The figure confirms that both the growth and the decay match their predicted rates: the gap’s measured slope for ζ(3)\zeta(3) is within one per cent of 4log⁡10(1+2)4\log_{10}(1 + \sqrt 2), and the same check holds for ζ(2)\zeta(2).

That ζ(2)\zeta(2) and ζ(3)\zeta(3) yield to the same construction, with a square and a cube in the recurrence and in the multiplier, suggested to everybody that ζ(5)\zeta(5) would follow with fifth powers. It has not. Every attempt at a similar pair of sequences for ζ(5)\zeta(5) produces a gap that shrinks too slowly to beat the multiplier lcm(1,…,n)5≈e5n\mathrm{lcm}(1, \ldots, n)^5 \approx e^{5n}, and the race is lost.

Why the even values are easy and the odd ones are not

The reciprocal squares, fourth powers and all even powers have closed forms: ζ(2)=π2/6\zeta(2) = \pi^2/6, ζ(4)=π4/90\zeta(4) = \pi^4/90, and in general ζ(2k)\zeta(2k) is a rational multiple of π2k\pi^{2k}. Where the coefficients come from found the first of these from the Fourier series of a simple function, and the others follow the same way. Each even value is therefore irrational, indeed transcendental, as soon as π\pi is, and nothing like Apéry’s race is needed for any of them.

For the odd values no closed form is known, and none is expected. The Fourier argument produces even powers because squaring a coefficient is what Parseval’s identity does, and odd powers have no such mechanism behind them. So each odd value has to be attacked on its own, and the attack has to manufacture its whole-number sequences from scratch. That ζ(3)\zeta(3) yielded at all is the anomaly; that ζ(5)\zeta(5) has not is what anyone would have expected before 1978.

The contrast also explains why Apéry’s ζ(2)\zeta(2) proof was a curiosity rather than a result: it proved something already known, by a method that was new. Its value is the one shown in the second figure — the same structure, a square where the cube was, with its own algebraic growth rate — which is the evidence that the ζ(3)\zeta(3) proof is not a one-off coincidence but an instance of something.

The skeleton, number by number

Set beside the other irrationality proofs in this collection, the skeleton is always the same and the materials are always different. The square that cannot shrink proved 2\sqrt2 irrational by descent, where the whole numbers are the sides of ever smaller squares; which roots refuse to be fractions generalised it to every root that is not whole. For ee the whole numbers came from factorials, and the pattern in e’s continued fraction found the same smallness written as integrals. For π\pi, the fraction Lambert built for the tangent used a continued fraction whose partial quotients grow.

In each case a sequence of whole-number combinations of the number is made small, and the only question is how much the wholeness costs. For square roots it costs nothing; for ee it costs a factorial, which the series repays at once; for ζ(3)\zeta(3) it costs 2 lcm(1,…,n)32\,\mathrm{lcm}(1, \ldots, n)^3, and the repayment is the exponent 3.52553.5255 against 33. The proofs differ in how much margin they have, and Apéry’s has less than any of the classical ones — which is why it came last.

The least common multiple is the opponent

How fast lcm(1, …, n) grows, against what Apéry's proofs can afford. ln lcm(1..n)/n for n up to 1000, ending at 0.9967, against thresholds 1.1752 (ζ(3)) and 1.2030 (ζ(2)).
Fig. 4 The logarithm of lcm(1, …, n), divided by n, for n up to a thousand: it tends to one by the prime number theorem. The proofs for ζ(3) and ζ(2) work as long as it stays below the dashed lines, which it does with margins of about seventeen and twenty per cent.

The growth of lcm(1,…,n)\mathrm{lcm}(1, \ldots, n) is the growth of Chebyshev’s function ψ(n)\psi(n), the sum of log⁡p\log p over all prime powers pk≤np^k \le n, since the least common multiple contains each prime to the highest power not exceeding nn. The prime number theorem says ψ(n)/n→1\psi(n)/n \to 1, and the figure shows it creeping up from below to 0.997 at n=1000n = 1000. Apéry’s argument for ζ(3)\zeta(3) needs this ratio to stay below 43log⁡(1+2)=1.1752\tfrac{4}{3}\log(1 + \sqrt2) = 1.1752, and for ζ(2)\zeta(2) below 52log⁡φ=1.2030\tfrac52 \log\varphi = 1.2030.

So the proof uses the distribution of the primes, and in a weak form: an upper bound ψ(n)≤(1+ε)n\psi(n) \le (1 + \varepsilon) n for a modest ε\varepsilon suffices, and that much was known to Chebyshev in the 1850s, long before the prime number theorem. The primes enter because the denominators of bnb_n are built from them, and the least common multiple is the smallest number divisible by everything that can appear. A sharper analysis of exactly which primes divide the denominators of bnb_n — some do not, to the full power — gives a slightly smaller multiplier and a slightly wider margin, which is how the best later versions of the argument improve the constants.

Bad approximations, good enough

The proof produces fractions Bn/AnB_n/A_n that approach ζ(3)\zeta(3), and it is natural to expect that they must be unusually good approximations. They are not.

Apéry's fractions against the best fractions for ζ(3). Error against denominator, log–log, for continued-fraction convergents of ζ(3) and for Apéry's fractions; the latter approximate with exponent about 1.14.
Fig. 5 Correct decimal digits against the number of digits in the denominator, for the continued-fraction convergents of ζ(3) (dots), which as for every irrational number are right to at least twice as many digits as their denominators have, and for Apéry’s fractions written over the whole-number denominators the proof uses (solid).

The best rational approximations to any irrational number are the convergents of its continued fraction, and they are correct to about twice as many digits as their denominators have. Apéry’s fractions, written over the denominators AnA_n that make the proof work, are correct to only about 1.14 times as many digits as their denominators have — far worse than the convergents. A fraction with a 100-digit denominator from Apéry’s sequence is right to about 114 digits; a convergent with a 100-digit denominator is right to about 200.

The proof does not need good approximations, and this is the conceptual point that separates it from the approximation arguments of approached too fast to be algebraic. What it needs is a family of fractions whose denominators are known — so that their size can be bounded in advance — and whose errors fall faster than one over the denominator. The convergents are better approximations but nobody can predict their denominators, so they prove nothing; Apéry’s fractions are worse but completely explicit, and that is what makes them useful. The same quantity measures how far from rational the number is shown to be: the irrationality measure that the proof yields for ζ(3)\zeta(3) is about 13.4, meaning that ∣ζ(3)−p/q∣>q−13.4|\zeta(3) - p/q| > q^{-13.4} for large qq, a bound later improved by Georges Rhin and Carlo Viola to about 5.5.

Where the sequences came from

Apéry did not explain how he found his recurrence, and the explanations came afterwards. Frits Beukers found in 1979 that the gap is an integral: anζ(3)−bna_n\zeta(3) - b_n equals, up to a constant factor, a triple integral over the unit cube of a polynomial in three variables raised to the nn-th power, divided by a fixed expression. The integral is visibly small, because the integrand is at most (2−1)4n(\sqrt2 - 1)^{4n} everywhere, and visibly of the form whole number times ζ(3)\zeta(3) minus a fraction with controlled denominator, because integrating powers over the cube produces exactly such numbers. In that form the proof looks like the integral proofs for π\pi and for ere^r, and the miracle becomes a choice of polynomial.

Others traced the recurrence to modular forms: the generating function of the ana_n is related to a modular form for a congruence subgroup, which explains why the ana_n satisfy congruences and why their growth rate is the algebraic number (1+2)4=17+122(1 + \sqrt2)^4 = 17 + 12\sqrt2. Neither explanation shows how to find the analogous objects for ζ(5)\zeta(5), because neither identifies which feature of the cubes made the integral or the modular form available.

What the computation does not show

The figures compute everything they draw exactly, up to sixty terms, and check every identity the proof uses at every one of those terms. They do not prove the identities for all nn. That ana_n is always the binomial sum, that 2 lcm(1,…,n)3bn2\,\mathrm{lcm}(1, \ldots, n)^3 b_n is always whole, and that the gap is never exactly zero are theorems whose proofs are algebraic — the integrality of bnb_n in particular needs an argument about the denominators of partial sums of 1/k31/k^3 that the original audience reasonably wanted to see.

The rates measured in the figures are also only rates on a finite stretch. That the gap shrinks like (2−1)4n(\sqrt2 - 1)^{4n} in the limit follows from the recurrence, whose characteristic roots are (1±2)4(1 \pm \sqrt 2)^4; the computation confirms the rate to within a per cent over nn from 30 to 60, which is evidence that the asymptotics have set in, not a substitute for them.

One more limitation is easy to overlook. The exact arithmetic makes the figures trustworthy as computations, but the value of ζ(3)\zeta(3) they compare against is itself computed — to four hundred digits, from a rapidly converging series of Apéry’s own — and the comparison would be meaningless if those digits were wrong. They are checked against the known decimal expansion, and four hundred digits is far more than the sixty-term gaps need, which reach about 10−9510^{-95}; but the chain of trust runs through that series, and a reader who wants the proof rather than the picture needs the identities, not the digits.

Still open: the odd values beyond three

Whether ζ(5)\zeta(5) is irrational is not known. Nor is the irrationality of any individual ζ(2k+1)\zeta(2k + 1) for k≥2k \ge 2, or of Catalan’s constant, or of ζ(3)/π3\zeta(3)/\pi^3. The progress since Apéry has come in a different form. Tanguy Rivoal and Keith Ball proved in 2000 that infinitely many of the values ζ(3),ζ(5),ζ(7),…\zeta(3), \zeta(5), \zeta(7), \ldots are irrational, and Wadim Zudilin proved in 2001 that at least one of ζ(5)\zeta(5), ζ(7)\zeta(7), ζ(9)\zeta(9) and ζ(11)\zeta(11) is — results that run races like Apéry’s but with many constants at once, so that the gap can be made small enough by spending freedom across several numbers rather than one.

Those theorems say that irrational values exist without saying which, and the obstruction is the same race. For a single odd value beyond 3, every known construction of whole-number sequences produces gaps that shrink more slowly than the least common multiple’s power grows. Whether some better sequence exists for ζ(5)\zeta(5), or whether the method itself has reached its limit, is not known; nor is whether ζ(3)\zeta(3) is transcendental, which Apéry’s argument, measuring only how far ζ(3)\zeta(3) is from fractions, cannot touch.