A series that converges nowhere
Worth reading first: An error with an unknown in it · A denominator that reaches past the radius.
Take the integral
which is a perfectly ordinary function for every : it is at nought, it decreases, and at it is Now do the obvious thing. Expand as a geometric series, , and integrate term by term. Each is , so
Every step looks legitimate and the result is a series that converges for no except nought: the ratio of consecutive terms is , which passes one as soon as passes , whatever is. By every test in the subject, this series does not represent anything.
And yet twenty terms compute to eight decimal places. The series is useless if the question is what its partial sums converge to, and excellent if the question is how close a partial sum can get. Those are different questions, and this essay is about the second.
Where the term-by-term integration went wrong
The geometric series equals only for . Inside the integral , and runs to infinity, so for every there is a stretch of the integral — everything beyond — on which the expansion used was simply false. Swapping the sum and the integral pretended that stretch was not there.
That diagnosis already contains the whole story, because it says how much was ignored. The weight on the stretch beyond totals . At that is about ; at it is about . The damage done by the illegal step is exponentially small in , and so for small the series can be very wrong about the limit of its partial sums while being very nearly right about the function.
A series that converges and a series that is useful are therefore not the same idea, and the difference lives exactly where the ordinary Taylor story never looks. There, the coefficients were derivatives at a point and the only question was the disc on which the series works. Here the coefficients are also derivatives at a point — is smooth at nought from the right, and is its -th derivative divided by — and the disc has radius nought, and the series works anyway.
A remainder that can be written down
What makes Euler’s series tractable rather than merely suggestive is that its error has an exact formula.
The finite geometric identity is true for every , with no condition on its size. Put , multiply by and integrate, and nothing illegal happens:
Since , the integral on the right is at most . So the error after terms is at most , the first term left out — and its sign alternates with , so consecutive partial sums fall on opposite sides of the function and bracket it.
That is a remainder with no unknown in it, which is rarer than it sounds. Lagrange’s remainder names a point nobody can locate and is bounded by replacing it with a worst case; this one is an explicit integral whose size is bounded in one line. The hero figure requires the bound at every number of terms it draws, and every error sits below the first term left out.
Stop at the smallest term
With the bound in hand the right number of terms is a calculation. The terms have ratio from one to the next, so they shrink while and grow after. The smallest term — and therefore the best available guaranteed bound — comes at about terms.
How small is it? At the term is , and Stirling’s formula turns that into about , which is
The measured best errors in the hero figure sit at about half of that: at against a smallest term of , and at against . The factor of a half is the bracketing at work — the function lies between two consecutive partial sums, roughly in the middle.
The numbers at show the mechanism exactly. The function there is The sum of nine terms is , above it by ; ten terms give , below it by ; eleven give again. The last two terms added are and , and those are the same number, , because at the ratio between them is exactly one. After that the terms grow: twelve terms are out by , thirteen by , and by thirty the error is in the hundreds. The smallest term is reached twice and the sums swing across the function between its two appearances, and no number of terms does better than the middle of that swing.
The rule is: keep adding terms while they shrink, and stop at the smallest. It is called optimal truncation, and it inverts every instinct about series. For a convergent series more terms can only help; here the best number of terms is a function of , and the twenty-first term at undoes accuracy that the first twenty paid for.
What more terms do to the curve
The same facts look different drawn against rather than against the number of terms, and the second picture is the one that explains why the series was trusted for so long.
Near nought the picture is indistinguishable from a convergent Taylor series: each extra term improves the agreement. What gives the game away is how far the agreement reaches. Ask for the sum to stay within of the curve. Two terms manage it until about , four until about — so far, the familiar story — but eight manage it only until about , twenty until about , and a hundred until about .
The interval of agreement widens for the first few terms and then shrinks for ever. A convergent series can only widen it. The turning point is optimal truncation seen from the other side: for a fixed tolerance there is a best number of terms, and past it every term buys accuracy closer to nought by giving up accuracy everywhere else. As the number of terms grows the intervals collapse towards nought, and that collapse is what a radius of nought looks like drawn against .
That is the geometric meaning of a radius of nought. The partial sums do converge, in a sense, at every — they converge to as approaches nought, faster and faster the more terms there are. What they never do is converge at a fixed as terms are added.
A radius of nought, read off the coefficients
The coefficients announce the radius before any sum is taken. The size of the coefficients decides it: the radius is one over the limit superior of , and for Euler’s series that quantity is , which grows like and has no limit at all.
So every method that treats a function as the thing its series converges to is silent here. There is no disc, no analytic continuation from a disc, no singularity at a positive distance to locate. The information in the coefficients is real — they are the derivatives of a real function — and the machinery of convergence cannot see any of it.
Convergence was the wrong question
Poincaré and Stieltjes, independently and in the same year, 1886, gave the definition that makes sense of this. A series is asymptotic to as if, for every fixed , the error after terms is smaller than by a factor that goes to nought as does.
Read that carefully, because it swaps the order of two limits. Convergence fixes and lets the number of terms grow. Asymptoticity fixes the number of terms and lets shrink. Euler’s series fails the first and passes the second — the bound is, for fixed , a constant times — and the two properties are simply about different limits.
The price of the weaker definition is uniqueness, and it is worth being exact about. A function determines its asymptotic series, but the series does not determine the function: is smaller than every power of as , so its asymptotic series is nought, and has exactly the same series as . The coefficients cannot tell those two functions apart, and the amount by which they differ is the size of the best error optimal truncation can reach. That is not a coincidence: no method using only the coefficients can promise more accuracy than the ambiguity the coefficients leave.
What a denominator recovers
That ambiguity makes the next fact startling. A quotient of polynomials built from the coefficients needs no convergence at all — it needs only as many coefficients as it has unknowns — and for Euler’s series the Padé approximants converge, at every , to and not to anything else.
The approximants find because has the same form as the logarithm: an integral of against a positive weight, here on the whole half-line. The nodes of the quadrature that weight calls for are the roots of the Laguerre polynomials, which are what straightening the powers of one at a time produces when the inner product is — the same construction that produced Legendre’s polynomials for the logarithm, under a different weight. Such a function has a cut along the negative axis, and the approximants imitate the cut with poles.
For the logarithm the cut began at and the poles crowded towards . Here the weight reaches all the way to infinity, the cut reaches all the way to nought, and the poles crowd towards the origin — which is the same picture as the radius of nought, drawn as a place rather than as a number.
The resolution of the uniqueness problem is that the approximants are not using the coefficients alone. They are implicitly assuming the function is of this integral form, and among functions of that form the coefficients do determine the answer — a statement about when a sequence of moments determines its weight, which it does here, and which is the content of a theorem of Carleman’s. Borel’s method of summing divergent series makes the same assumption explicitly: replace by the integral it came from and sum inside, and comes back.
Spending more coefficients, slowly
The convergence is real and it is not fast, and the reason says something about where singularities are.
For the logarithm at each order divided the error by nine. Here the last step divides it by about two, and for functions of this kind the theory says no fixed ratio is ever reached: the error shrinks faster than any power of and slower than any geometric rate. The difference is the cut. The logarithm’s began at a positive distance from the point being evaluated, measured in the geometry the cut defines; Euler’s function’s cut touches nought, and the evaluation point is, in that geometry, as close to the singular set as a point on the positive axis can be.
Coefficients that grow like factorials are the signature of a singularity at distance nought, and every method working from them pays for that distance — optimal truncation with an irreducible error, Padé approximation with slow convergence. What neither pays is failure.
Where divergent series are the tool that works
Euler’s series is the tidiest case of something that is everywhere.
Stirling’s formula for the factorial extends to a series of corrections for — — and that series diverges for every . Its first two corrections nonetheless give to about nine digits, and it is what practical computations of the gamma function use.
In physics the pattern is the rule rather than the exception. Dyson argued in 1952 that the perturbation series of quantum electrodynamics in powers of the fine-structure constant must diverge, because a negative coupling constant would make the vacuum unstable and no function analytic at nought can have that property on one side only. The series is nonetheless the most accurately confirmed calculation in science, because the constant is about and optimal truncation would occur after roughly a hundred and thirty-seven orders — far more than anyone computes.
And the functions that solve the most common differential equations near their irregular singular points — Airy’s function, Bessel’s functions at large argument — have asymptotic series of exactly this kind. Stokes noticed in 1857 that the exponentially small terms the coefficients cannot see switch on and off across certain lines in the complex plane, which is the ambiguity above turning into a phenomenon.
Where the account needs care
The bound by the first omitted term is special. It holds for Euler’s series because of the integral form and the positivity of the weight. A general asymptotic series comes with no such guarantee, and optimal truncation is then a heuristic that usually works and occasionally does not — terms of irregular sign can make the smallest term a poor guide to the error.
Asymptotic means as goes to nought. At the best available accuracy is two terms and an error of , which is hardly an approximation at all. The series is a statement about small , and how small is small is decided by , not by the size of the first few coefficients.
Formal manipulation of divergent series can produce nonsense. Summing as a geometric series gives , which is meaningful in some settings and absurd in others, and a divergent sum of positive terms has no value that respects the ordinary rules. Euler’s series has a meaningful sum because it came from a specific integral; a series with no such origin need not have one.
And the integral form does the work in the recovery. Padé approximation and Borel summation both recover because is an integral of a particular kind. For a series produced by some other process, a different function with the same coefficients may be the one wanted, and nothing in the coefficients says which.
Euler, Abel’s devil, and 1886
Euler wrote about exactly this series in 1760, in a paper on divergent series. He assigned the sum the value , by several methods including the integral and a continued fraction, and was untroubled that the partial sums wander off.
The next generation was troubled. Abel wrote in 1826 that divergent series are an invention of the devil and that it is shameful to base any demonstration on them, and the rigorous analysis of Cauchy and Weierstrass banished them from proofs. Astronomers kept using them, because they worked.
The rehabilitation came in 1886, when Poincaré and Stieltjes independently defined asymptotic series — Poincaré motivated by the series of celestial mechanics, Stieltjes by exactly this integral. The mathematics was not wrong for a century; it was answering a question nobody had yet posed properly, and the constant Euler computed, , is now called the Euler–Gompertz constant.
What a U-shaped curve cannot show
The hero figure draws thirty partial sums at three points. The claims being made — that the error has a single minimum near terms, that it grows without bound after — are statements about all and all small , and they are supported by the exact remainder rather than by the curves.
The pole figure shows the approximants’ poles moving towards nought as the order grows and cannot show the limit, which is the whole negative axis filled.
And no picture here shows the ambiguity. and differ by less than a pixel for every below about on these axes, so the fact that the series cannot tell them apart is exactly the fact the drawings cannot display.
Still open: what a series knows beyond its sum
Optimal truncation leaves an error of about and says it cannot do better from the coefficients alone. The subject that takes that residue seriously — resurgence, which reads the exponentially small corrections off the large-order behaviour of the coefficients themselves — is where Stokes’ switching becomes a calculation. Borel summation in general, and the question of which divergent series have a unique natural sum, is the other thread.
On the approximation side, the node sets in the points that ruin the fit lead towards minimax approximation, where a polynomial is judged by its worst error and the best one is recognised by an error curve that touches its extremes with alternating signs. That criterion, like optimal truncation, judges an approximation by how good it is at a fixed level of effort rather than by where a sequence of them goes.
Judged at a fixed effort, not at infinity
The habit worth carrying is about which limit a claim is about.
“This series converges” is a statement about infinitely many terms at a fixed point, and it says nothing about how well any finite number of terms does. “This series is asymptotic” is a statement about a fixed number of terms at points approaching nought, and it says nothing about what happens as terms are added. Euler’s series fails the first and passes the second, and every useful fact about it came from the second.
Before asking whether an approximation converges, ask what effort it will actually be given. A method that converges slowly can lose to one that does not converge at all, at every level of effort anybody will spend, and the sum that fits in one square and this series are the two ends of that comparison.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The size of a number with no formula — both name approximation, convergence
- The staircase that is not the diagonal — both name approximation, convergence
Named objects
A dashed tag is an object no other essay names yet.
ApproximationAsymptotic seriesBoundConvergenceError analysisPade approximantRadius of convergenceRemainderTaylor series