Number

A cube root that looks like chance

The continued fraction of √61 repeats after eleven terms, and every square root's must. The cube root of 2 obeys no such law. Computed exactly, by a method of Lagrange's that never leaves the whole numbers, its terms run 1, 3, 1, 5, 1, 1, 4 … and throw up a 534 at the thirty-fifth place and a 7,451 at the 571st — and in every statistic anyone has measured, they look like the terms of a number picked at random.
19 min read 6 figures Order out of noiseSmall cases lie

Worth reading first: Why the expansion has to repeat · A fraction that never closes.

Why the expansion has to repeat proved that the continued fraction of every square root is periodic: 61=[7;1,4,3,1,2,2,1,3,4,1,14‾]\sqrt{61} = [7; \overline{1, 4, 3, 1, 2, 2, 1, 3, 4, 1, 14}], and it has no choice, because each step’s state is a pair of whole numbers trapped in a small box. Every equation in the preceding essays on Pell’s equation rested on that periodicity. The fundamental solution is the product of one period’s complete quotients, the sign of −1-1 is the period’s parity, and the whole theory of the equation is a theory of one repeating block.

That essay ended on the next simplest kind of number: the root of a cubic, such as 23\sqrt[3]{2}. Its expansion begins [1;3,1,5,1,1,4,1,1,8,1,14,…][1; 3, 1, 5, 1, 1, 4, 1, 1, 8, 1, 14, \ldots], and it is not known whether its terms stay bounded. This essay computes the expansion properly, to six hundred terms and beyond, looks at what comes out, and measures it against the one theory that predicts anything about it — which is a theory not of cube roots but of numbers chosen at random.

The contrast is the point. For a square root there is a box, and so a period, and so everything. For a cube root the same argument has nowhere to start: the state of the expansion is a cubic polynomial whose coefficients grow without limit, so no pigeonhole ever closes. What replaces the law is a set of statistics, and whether they are laws or only habits is unknown.

Seventeen terms from a calculator, and then noise

The obvious way to compute a continued fraction is the definition. Take x=23=1.2599…x = \sqrt[3]{2} = 1.2599\ldots, record its integer part 1, replace xx by 1/(x−1)=3.8473…1/(x - 1) = 3.8473\ldots, record 3, and repeat. The trouble is that each step subtracts two nearly equal numbers and then divides by the small difference, which multiplies the error already present by roughly the square of the new complete quotient.

Starting from the best value an ordinary computer can hold — about sixteen correct digits — this procedure gives exactly seventeen correct terms of 23\sqrt[3]{2}, and then diverges from the truth. The eighteenth term it produces is wrong, and every term after that is noise dressed as a continued fraction. Nothing in the output warns of the moment it happens: a wrong term looks exactly like a right one.

More digits postpone the failure without removing it. Each term consumes, on average, a little over one decimal digit of precision — the reason is the growth rate of the denominators, measured below — so six hundred terms need about six hundred correct digits at the start, and there is no way to know in advance how many the large terms will eat. A term of 7,451 costs about eight digits on its own.

A cubic for every step

Lagrange found a way out in the 1760s, and it avoids decimals entirely. Instead of carrying an approximation of the current complete quotient, carry the polynomial it is a root of.

Lagrange's method: a cubic for each complete quotient of ∛2. x³ − 2: integer part 1; x³ − 3x² − 3x − 1: integer part 3; 10x³ − 6x² − 6x − 1: integer part 1; 3x³ − 12x² − 24x − 10: integer part 5; 55x³ − 81x² − 33x − 3: integer part 1.
Fig. 1 Lagrange’s method on the cube root of 2. Each row is a cubic with whole-number coefficients whose one root above 1 is the next complete quotient. The integer part aa of that root is found by where the cubic changes sign, and substituting x=a+1/yx = a + 1/y and clearing fractions gives the next row. The integer parts read off are 1, 3, 1, 5, 1 — the continued fraction.

23\sqrt[3]{2} is the root of x3−2x^3 - 2. That cubic is negative at 11 and positive at 22, so the root lies between them and its integer part is 11. Now write x=1+1/yx = 1 + 1/y, where yy is the next complete quotient, and substitute: (1+1/y)3−2=0(1 + 1/y)^3 - 2 = 0 becomes, after multiplying through by y3y^3, the cubic y3−3y2−3y−1=0y^3 - 3y^2 - 3y - 1 = 0. That cubic has one root above 1, it changes sign between 3 and 4, so the next term is 3. Substitute y=3+1/zy = 3 + 1/z, clear fractions, and the cubic for zz is 10z3−6z2−6z−110z^3 - 6z^2 - 6z - 1, whose root lies between 1 and 2.

Every step is whole-number arithmetic: a shift of the polynomial by the integer part, which is a table of binomial coefficients, and a reversal of the coefficient list, which is the 1/y1/y. Nothing is rounded, so every term the method produces is exactly right, however far it runs. The cost has moved from precision to size. The coefficients grow: after three thousand terms they have about 1,526 digits each. But whole numbers of that size are cheap, and three thousand terms take a fraction of a second.

What does not happen is the thing that happens for square roots. For 61\sqrt{61} the analogous state was a pair (m,d)(m, d) that could take only fourteen values, so it had to repeat. For 23\sqrt[3]{2} the state is a cubic whose coefficients grow roughly in step with the denominators of the convergents. It never repeats, and it never could: a repeating expansion would make 23\sqrt[3]{2} the root of a quadratic, which it is not.

The invariant that survives, and the box that does not

The failure of periodicity can be pinned down more exactly than “the coefficients grow”, and doing so shows what the square-root argument really used.

Each of Lagrange’s cubics can be read as a binary cubic form, ax3+bx2y+cxy2+dy3ax^3 + bx^2y + cxy^2 + dy^3, and the step x=a+1/yx = a + 1/y is a change of variables with whole-number coefficients and determinant ±1\pm 1. Such a change preserves the form’s discriminant, 18abcd−4b3d+b2c2−4ac3−27a2d218abcd - 4b^3d + b^2c^2 - 4ac^3 - 27a^2d^2. For x3−2x^3 - 2 it is −108-108. For x3−3x2−3x−1x^3 - 3x^2 - 3x - 1 it is −162−108+81+108−27=−108-162 - 108 + 81 + 108 - 27 = -108 again. Computed exactly for each of the first six hundred cubics in the expansion, it is −108-108 every time, while the leading coefficient grows from one digit to four by the tenth step and to fifty-seven by the hundredth.

The square roots had exactly this invariant. The quadratic whose root is each complete quotient of 61\sqrt{61} has the same discriminant at every step, 4⋅614 \cdot 61. What made the expansion repeat was a second fact: the quadratic’s other root always lies between −1-1 and 00, and a quadratic with a fixed discriminant and its two roots so placed has bounded coefficients. Finitely many quadratics qualify, so one recurs. That is the box why the expansion has to repeat drew, and it is also why sixty needs two digits and sixty-one needs ten could read the fundamental solution off one period’s complete quotients.

A cubic’s other two roots are a complex pair, and nothing holds them apart. As the expansion runs they crowd together, and a cubic can keep its discriminant fixed while its roots crowd and its coefficients grow without limit. The invariant survives; the box does not. The field Q(23)\mathbb{Q}(\sqrt[3]{2}) still has infinitely many units — units that form a lattice found them all as powers of 23−1\sqrt[3]{2} - 1 — but in a quadratic field the period and the unit are the same object, and here there is a unit with no period attached to it.

Six hundred terms, and a 7,451

The first 600 terms of the continued fraction of ∛2. Partial quotients of ∛2 computed exactly: 1, 3, 1, 5, 1, 1, 4, 1, 1, 8, 1, 14, 1, 10, 2, 1, 4, 12, 2, 3 …; the largest in the first 600 is 7451.
Fig. 2 The first 600 partial quotients of 23\sqrt[3]{2}, computed exactly, each drawn as a bar on a logarithmic scale. Most are 1, 2 or 3. Every so often a large one appears: 534 at place 35, 121 at 41, 186 at 91, 372 at 114, and a 7,451 at place 571. No pattern connects their positions.

Most of the terms are small. In the first six hundred, about four in ten are 1, and more than half are 1 or 2. Then, at the thirty-fifth place, comes a 534, and at the 571st a 7,451. Running the method further, to three thousand terms, finds a 4,941 at place 619 and a 12,737 at place 1,990, and above a hundred about fifty times in all.

There is no visible rhythm in where they fall. The gaps between the terms above 100 run 6, 50, 23, 261, 45, 90, 61 and on, with no two alike and no drift. A sequence built by drawing each term at random from a fixed distribution would look like this — and a sequence built by a rule with a hidden period of a few thousand would also look like this for its first few thousand terms, which is why no amount of looking settles anything.

Lang and Trotter published tables of this expansion in 1972, and computations since have pushed it to millions of terms. Very large terms keep appearing, at intervals that stretch as the expansion lengthens, as they do for a number chosen at random.

Four numbers, three kinds of expansion

Set beside other familiar numbers, the cube root’s expansion falls into a clear place.

Continued fractions of √2, e, π and ∛2 side by side. √2: 1,2,2,2,2,2,2,2,2,2,2,2 …; e: 2,1,2,1,1,4,1,1,6,1,1,8 …; π: 3,7,15,1,292,1,1,1,2,1,3,1 …; ∛2: 1,3,1,5,1,1,4,1,1,8,1,14 ….
Fig. 3 The first 40 partial quotients of four numbers, each bar scaled by its logarithm, with the large ones labelled. 2\sqrt 2 repeats 2 forever, as every quadratic irrational eventually repeats. ee follows Euler’s pattern 1, 2k, 1. π\pi and 23\sqrt[3]{2} show no pattern at all: π\pi has 292 at its fourth place, 23\sqrt[3]{2} has 534 at its thirty-fifth.

2\sqrt 2 is [1;2,2,2,…][1; 2, 2, 2, \ldots]: the periodic case, which a fraction that never closes first drew. ee is [2;1,2,1,1,4,1,1,6,…][2; 1, 2, 1, 1, 4, 1, 1, 6, \ldots], which Euler found in 1737 — not periodic, but perfectly regular, with the pattern 1,2k,11, 2k, 1 running forever. π\pi is [3;7,15,1,292,1,1,1,2,…][3; 7, 15, 1, 292, 1, 1, 1, 2, \ldots], and 23\sqrt[3]{2} is the fourth row. Their terms were computed differently — π\pi’s and ee’s from four hundred correct digits, keeping only the terms that the error interval fixes; 23\sqrt[3]{2}'s from cubics — and they look alike.

That kinship is strange, because the two numbers could hardly be more different. π\pi is transcendental: no polynomial with whole-number coefficients has it as a root. 23\sqrt[3]{2} is the root of x3−2x^3 - 2, about as simple a polynomial as exists. Yet for every quantity anyone has measured on their continued fractions, the two behave the same, and both behave like a random number. ee, which is transcendental like π\pi, is the one that stands out, because its terms follow a formula.

What a large term buys

A large partial quotient is not merely a large number in a list. It is a record of an unusually good fraction.

How close each of the first 600 convergents of ∛2 comes, for its size. q² |∛2 − p/q| for 600 convergents; smallest 1.34e-4 at place 570 before the term 7451; (ln q)/n = 1.2109, Lévy 1.1866.
Fig. 4 For each of the first 600 convergents p/qp/q of 23\sqrt[3]{2}, the error ∣23−p/q∣|\sqrt[3]{2} - p/q| multiplied by q2q^2, on a logarithmic scale. It always lies between 1/(a+2)1/(a+2) and 1/a1/a, where aa is the next partial quotient, so every large term appears as a deep notch — the deepest just before the 7,451.

If the next partial quotient is aa, the convergent just before it satisfies ∣23−p/q∣≈1/(aq2)|\sqrt[3]{2} - p/q| \approx 1/(a q^2). A term of 7,451 means the convergent before it is about 7,451 times better than the typical approximation with a denominator of that size. These convergents are the best approximations there are — each closer than every fraction with a smaller denominator, as the fractions that beat every smaller one showed by walking down the tree of fractions — and a large term is a best approximation that is also unusually good. The notches in the figure are those convergents: q2q^2 times the error drops below 0.0010.001 a few times and to about 0.000130.00013 once.

So asking whether the terms are bounded is asking whether 23\sqrt[3]{2} can be approximated unusually well infinitely often. Here the answer is partly known. Approached too fast to be algebraic followed the story from Liouville to Roth: in 1955 Klaus Roth proved that for any algebraic irrational α\alpha and any ε>0\varepsilon > 0, the inequality ∣α−p/q∣<q−2−ε|\alpha - p/q| < q^{-2-\varepsilon} has only finitely many solutions. In the language of the figure, apart from finitely many notches, none goes below q−εq^{-\varepsilon}. For the terms, that means a<qεa < q^{\varepsilon} eventually, for every ε\varepsilon.

That is a real constraint and a weak one. It allows the terms to grow — like log⁡q\log q, or like q1/1000q^{1/1000} — without limit. Roth’s theorem forbids 23\sqrt[3]{2} to be approached as fast as a Liouville number; it does not forbid it the occasional very good fraction forever. And Roth’s theorem is ineffective: it says the exceptions are finite and gives no way to find the last one.

The denominators themselves grow at a rate that, again, is the random rate. Paul Lévy proved in 1936 that for almost every number (ln⁡qn)/n(\ln q_n)/n tends to π2/(12ln⁡2)≈1.1866\pi^2/(12 \ln 2) \approx 1.1866. For 23\sqrt[3]{2}, after 1,200 terms the denominator has 614 digits and (ln⁡qn)/n=1.179(\ln q_n)/n = 1.179. That rate is also why each term costs about one decimal digit of precision. The $n$th convergent is accurate to about 1/qn21/q_n^2, and qnq_n has about 1.1866 n/ln⁡10≈0.515 n1.1866\,n / \ln 10 \approx 0.515\,n digits, so pinning down nn terms takes about 1.03 n1.03\,n correct digits at the start.

The statistics of a number chosen at random

Almost every real number — every number except a set of total length zero — has a continued fraction with the same statistics. Gauss found the first of them, and Rodion Kuzmin proved it in 1928: the proportion of terms equal to kk tends to log⁡2 ⁣(1+1k(k+2))\log_2\!\left(1 + \frac{1}{k(k+2)}\right), which is about 41.5% for 1, 17.0% for 2, 9.3% for 3, and falls off like 1/k21/k^2.

How often each term appears in ∛2, against the Gauss–Kuzmin law. 1: 42.5% (law 41.5%); 2: 15.8% (law 17.0%); 3: 8.8% (law 9.3%); 4: 6.5% (law 5.9%); 5: 4.2% (law 4.1%); 6: 3.8% (law 3.0%); 7: 2.0% (law 2.3%); 8: 1.6% (law 1.8%); 9: 1.5% (law 1.4%); 10: 1.7% (law 1.2%).
Fig. 5 How often each value from 1 to 10 occurs among the first 1,200 partial quotients of 23\sqrt[3]{2} (bars), against the Gauss–Kuzmin law (dots). 42.5% of the terms are 1 against the law’s 41.5%, and 15.8% are 2 against 17.0%. Nothing about being a root of x3−2x^3 - 2 shows in the frequencies.

The cube root follows it closely. Among the first 1,200 terms, 42.5% are 1, 15.8% are 2, 8.8% are 3 and 6.5% are 4, against the law’s 41.5%, 17.0%, 9.3% and 5.9%. The differences are the size that sampling 1,200 terms from the law would produce.

A second statistic is the geometric mean of the terms. Alexander Khinchin proved in 1935 that for almost every number it converges to the same constant, about 2.68552.6855, whatever the number. A square root fails this: 2\sqrt 2’s mean is 2 forever. ee fails it the other way, because the terms 2k2k grow and its geometric mean grows with them.

The geometric mean of ∛2's terms, against Khinchin's constant. Running geometric mean of ∛2's partial quotients reaches 2.6561 at 1200 terms; Khinchin's constant 2.685452001.
Fig. 6 The geometric mean of the first nn partial quotients of 23\sqrt[3]{2}, for nn up to 1,200. After swinging through the first few hundred terms — the 534 pushes it up early, and a run of small terms pulls it back — it settles near 2.656, against Khinchin’s constant 2.6855.

For 23\sqrt[3]{2} the running mean swings early — the 534 at place 35 and the 372 at place 114 lift it above 3 — and then settles. At 600 terms it is 2.773, at 1,200 it is 2.656, at 3,000 it is 2.631. Khinchin’s constant is 2.685. The mean of a random sequence with the Gauss–Kuzmin distribution converges slowly, because the distribution has a heavy tail — a single term of 12,737 moves the logarithm’s running sum by more than nine — and these readings sit inside the spread such a sequence would show.

No proof connects any of this to 23\sqrt[3]{2}. The theorems of Gauss–Kuzmin, Khinchin and Lévy are about almost every number, and a set of measure zero can contain anything: every rational, every quadratic irrational, ee, and possibly every algebraic number of degree three. How close a fraction can get made the same point about Khinchin’s constant from the side of approximation. For no number that can be written down independently of its expansion has anyone proved that it obeys any of these laws.

Still open: whether the terms are bounded

For 23\sqrt[3]{2} — and for every other real algebraic number of degree three or more — it is unknown whether the partial quotients are bounded. It is not even known whether any one such number has unbounded terms, or whether any one has bounded terms; both classes could be empty for all anyone can prove.

The expectation is that they are all unbounded, with the Gauss–Kuzmin statistics, because the terms of a random number are: a term above NN appears with probability roughly 1/(Nln⁡2)1/(N \ln 2), and those probabilities add up to infinity, so large terms keep coming. The figures agree with the expectation as far as they go. A bounded expansion would have to stop producing 534s and 7,451s at some point, and within three thousand terms it has shown no sign of doing so — the largest so far is the 12,737 at place 1,990.

The difficulty is that every tool that sees continued fractions works either for quadratic irrationals, where periodicity gives everything, or for almost every number, where measure theory gives everything. 23\sqrt[3]{2} is in neither world. Roth’s theorem is the deepest result that applies to it, and it limits how fast the terms can grow without saying whether they grow. The coefficients of Lagrange’s cubics hold, in principle, the whole answer — each term is decided by where a cubic with known integer coefficients changes sign — and still nobody knows how to read a long-run pattern out of them.

What the computation shows and what it cannot

Every term drawn here is exact. The method never rounds, and its first dozen terms agree with a decimal expansion wherever the decimals are still reliable. Six hundred terms in one figure, 1,200 in the statistics, and three thousand in the prose, all come from the same whole-number recursion.

What no computation can show is a limit. The Gauss–Kuzmin frequencies, Khinchin’s mean and Lévy’s rate are statements about infinitely many terms, and the agreement after 1,200 is evidence of the kind that when minus one can be reached warned about: a quantity approaching its limit in the data is not a proof, and there the data sat twenty points from the limit that was eventually proved. Here the data sits almost on the predicted values, which is either a sign that the prediction is right or a sign that whatever makes 23\sqrt[3]{2} special shows only on scales nobody can compute.

The contrast with the square roots is the thing to keep. For 61\sqrt{61} a finite computation proved everything, because the period closed. For 23\sqrt[3]{2} an unlimited computation proves nothing about the pattern, because there is no period to close — only a cubic that keeps growing, and a sequence of terms that looks, from every angle anybody has tried, like chance.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

Algebraic numberAlgorithmContinued fractionsCubicPeriodicityQuadratic irrationalRandomness