A denominator that reaches past the radius
Worth reading first: One point's worth of information · The centre is a choice.
The series is the logarithm , and it is only the logarithm for . Beyond that its terms grow, the partial sums swing further from the curve with every term added, and no number of terms computes .
Moving the centre was one answer: expand about a point nearer to where the value is wanted, and walk the series out one overlapping disc at a time. That answer changes where the information is taken from. There is a second answer that takes exactly the same information — the coefficients at nought — and changes only what is built from it.
The quotients are Padé approximants, and the reason they work is not that a quotient is more flexible in some vague sense. It is a specific thing a quotient can do and a polynomial cannot: vanish in the denominator.
As many conditions as unknowns
A Padé approximant of order is a quotient , with of degree at most , of degree at most , and , whose own series agrees with the function’s series for as many terms as possible. There are coefficients in and free ones in , so conditions can be imposed: agreement through the term in .
The conditions are linear, which is the whole reason the construction is practical. Multiply through by and the requirement becomes . The coefficients of to in that product involve only , since has no terms that high, so they are equations for the unknowns of . Then is read off the lower terms.
The smallest non-trivial case can be done in a line. The series of starts . For , , and the coefficient of in is , which must vanish; so , and . The approximant is
It uses the three numbers , and , which are exactly the numbers the Taylor sum of degree two uses. At the Taylor sum gives and the quotient gives , against . At the Taylor sum gives and the quotient gives , against . Same information, arranged differently, and one arrangement is already usable where the other has stopped meaning anything.
Why a polynomial cannot do it
A polynomial is finite everywhere and has no singularities, and the Taylor sums of are polynomials trying to approximate a function that has one, at .
The radius is the distance to that singularity, measured in the plane, and it is the same in every direction. So the singularity on the left at decides what happens on the right at : a polynomial approximation built at nought is constrained to a disc, and a disc does not know that the trouble is on only one side. The sums fail at not because anything is wrong at but because is further from nought than is.
A quotient has a way out. Its denominator can vanish, and where it vanishes the quotient blows up — which is precisely what the function does at a pole. So a quotient can place a pole on or near the function’s singularity and absorb it there, and once the singularity is accounted for nothing ties the approximation to a disc any more.
The clearest demonstration is a function that is a quotient.
That approximant is itself: the linear system finds and , and the poles land exactly on . The Taylor sums were never going to get there, because no polynomial has a pole; the quotient needed only to find the denominator, and five coefficients were enough to determine it.
Five coefficients that know where the pole is
How five numbers located a pole at is worth seeing, because it says what the denominator’s equations really are. The coefficients of are , and they obey a rule: each is minus the one two places before. Written as , that rule is the identity , read off one power of at a time. The linear system for asks exactly this question — which short list of multipliers makes the coefficients satisfy a fixed recurrence beyond some point — and a recurrence of that kind is a denominator written sideways.
The general version is a theorem. A function whose only singularities inside some disc are poles has coefficients that are, far enough out, dominated by those poles: a pole at contributes a term like to the -th coefficient, times a polynomial in if the pole is repeated. A sequence assembled from such geometric pieces satisfies a linear recurrence with terms, and the roots of that recurrence’s characteristic polynomial are the reciprocals . The approximants of order , with equal to the number of poles, solve for that recurrence approximately; as grows their poles converge to the function’s poles and the approximants converge to the function throughout the disc, the poles themselves excepted. That is de Montessus de Ballore’s theorem, from 1902.
The denominator is the recurrence the coefficients eventually obey, and finding it is the same act as locating the poles. It is the rational counterpart of reading a sequence’s growth off its generating function’s nearest singularity: the same information, read in the other direction — from the coefficients to the singularity — and turned into something that can be evaluated at a point.
For a simple pole the whole thing can be done by hand. The coefficients of are the powers , each twice the one before. The recurrence is the denominator , and the approximant of order is the function exactly. The Taylor sum of degree at , outside the radius of one half, is , and the function there is : every additional term doubles an error that was already the wrong sign.
The example is special in one way — the function is rational — and ordinary in the way that matters. For a function that is not a quotient, the approximant cannot land its poles on a singularity it cannot reproduce exactly. What it does instead is the most interesting thing in the subject.
The poles line up along the cut
The logarithm’s singularity at is not a pole. It is a branch point, and continuing around it changes its value by , so the function cannot be made single-valued on the plane without cutting it along a ray — conventionally everything to the left of .
A quotient of polynomials has poles and nothing else, so it cannot have a cut. What the diagonal approximants do is build one out of poles.
No pole leaves the real axis and none lands to the right of . As the order grows, the poles crowd towards and spread out along the ray, and in the limit they fill it. The denominator is imitating the cut by standing poles along it, and everywhere off the cut — including all of the real axis to the right of , including and — the approximants converge.
The positions are not arbitrary, and the reason they are what they are is the connection the subject is best known for. The logarithm can be written as an integral,
and approximating that integral by a weighted sum of values of the integrand gives — a quotient of polynomials with poles at . If the sample points and weights are those of Gaussian quadrature on , the rule integrates every polynomial of degree up to exactly, and expanding in powers of shows that this is exactly the statement that the sum matches the series through coefficients. So the Padé approximant of order is times the -point Gaussian quadrature of that integral, and its poles are at minus the reciprocals of the Gauss nodes. The figure finds the poles as roots of the exact denominator and, separately, finds the nodes by Newton’s method on Legendre’s recurrence, and the two lists agree.
The nodes are the roots of the Legendre polynomials, which are what Gram–Schmidt manufactures from the powers of x under the product . A question about extending a logarithm past its radius and a question about straightening a basis of polynomials have the same answer, and it is a set of points that has been known since 1814.
The same coefficients, spent twice
The comparison worth making is not between an approximant and a Taylor sum of the same degree. It is between the two ways of spending the same coefficients: the approximant uses of them, exactly as many as the Taylor sum of degree .
The factor of nine is not a coincidence of the drawing. For a function of this kind the error of the diagonal approximants shrinks geometrically, at a rate fixed by how far the point is from the cut in a sense that the cut itself defines. At a point on the positive axis the rate per order is . At the square root is , the ratio is , and its square is — which is what the ten measured errors settle on.
The formula also says what happens far away. At the square root is a little over , the rate per order is about , and the approximants still converge — slowly, but at a point a hundred times beyond the radius, where the Taylor series’ hundredth term is larger than . And it says the rate only degrades as approaches the cut, never as it merely grows.
Inside the radius too
The case for the quotient is not only that it goes where the series cannot. Where both converge, it spends the coefficients better.
The Taylor sum at gains two terms per order, each term about a quarter of the size of the one two before, so a factor of four per order is what it can manage. The approximant’s rate from the same formula is about , and the measured sequence is approaching it. Sixteen coefficients buy seven digits one way and sixteen the other.
That is the practical version of the whole story, and it is why rational approximation is used in numerical libraries for functions like the exponential of a matrix, where the coefficients are cheap and each evaluation of the approximant is a linear solve. The series is the source of the information; how that information is turned into a number is a separate decision, and the polynomial is not the only candidate.
When there is nothing to imitate
The poles of an approximant are where it spends its extra freedom, so a function with no singularities at all is a good test of whether the freedom is being wasted.
For the exponential, whose Taylor series converges everywhere, the diagonal approximants still exist and still beat the Taylor sums of the same size — their poles simply move away from the origin as the order grows, off into the complex plane, where they do no harm on any bounded region. The quotient does not invent singularities the function lacks. It keeps them at a distance that grows with the order, which is the rational version of the Taylor series’ infinite radius.
Those approximants to have a history out of proportion to their size. Hermite used rational approximations to exponentials in 1873 to prove that is transcendental — the first number shown to be so that was not built for the purpose — and the proof works by showing the approximations are too good for to satisfy any polynomial equation with whole-number coefficients. The same objects that are computationally convenient turned out to encode an arithmetic fact.
A method of approximation is also a source of inequalities, and inequalities are what arithmetic proofs are made of. Approached too fast to be algebraic is the same mechanism with fractions in place of rational functions.
Where the method needs care
The approximant need not exist in the form asked for. For there is no approximant of order with : its coefficient of is nought, the linear system is singular, and the figure above uses for that reason. The general theory arranges approximants in a table and handles the missing entries, which occur in square blocks.
Noise in the coefficients produces false poles. If the coefficients are measured or rounded, the approximant can acquire a pole and a zero almost on top of each other, cancelling except in a tiny neighbourhood — a Froissart doublet. It is harmless away from itself and catastrophic near it, and it is why the figures here solve for the denominators in exact rational arithmetic. For the logarithm the linear system is a close relative of the Hilbert matrix, and in double precision it has lost most of its digits by order eight.
Convergence is not uniform in general. For functions whose singularities are poles, the diagonal approximants converge in a weaker sense — outside sets that can be made as small as wanted, but whose position can wander — and spurious poles can appear anywhere for a few orders before moving off. The clean picture of poles lined up on a cut belongs to functions built as integrals like the one above, and for those it is a theorem. For functions with several branch points the cuts the approximants choose are the ones of least capacity, which need not be the ones anybody would draw.
There is no remainder formula. A Taylor sum comes with an error term that has one unknown in it, and replacing the unknown by its worst case gives a guaranteed bound from the derivatives alone. Nothing of the kind exists for a general quotient: the error of an approximant depends on how well its poles have imitated singularities it cannot see, and the coefficients used to build it do not say. What practice does instead is compare consecutive orders and treat their agreement as evidence — a reasonable heuristic that a Froissart doublet can fool, since two orders can agree everywhere except near a false pole that only one of them has. For functions defined by integrals like the logarithm’s there are guaranteed bounds, because at a positive two neighbouring sequences of approximants close in on the true value from opposite sides and bracket it — but that is a property of the function, not of the method.
And the approximant does not know where the function is defined. On the cut itself the approximants do not converge to anything sensible, and nothing about a single approximant warns that a point is near a cut rather than merely near a pole.
Padé’s table, and the integrals behind it
The approximants are older than the name. Jacobi and Cauchy wrote down rational interpolants; Frobenius studied the systematic arrangement of approximants by numerator and denominator degree in 1881; Hermite used them for the exponential in 1873. Henri Padé’s thesis of 1892 organised them into the table that bears his name and studied its structure, including the blocks of repeated entries.
The connection to integrals and continued fractions is Stieltjes’, in the work of 1894 that also founded the moment problem. He showed that functions given by integrals of the form above have continued-fraction expansions whose truncations are exactly the diagonal approximants, and that their poles interlace and lie on the support of the integral — which is the lined-up picture, proved thirty years before anybody could have drawn it.
That order of events is common in analysis and worth noticing. The objects were computed for a century before anybody could say which functions they converge for, and the answer, when it came, was a statement about where the function’s singularities are rather than about the approximants.
What an overlapping curve hides
In the opening figure the two quotients lie on top of the logarithm, and that is the point of the figure and its limitation. An error of is a fraction of a pixel, so the drawing can distinguish a method that works from one that fails and cannot distinguish two that work. The rate figures exist because the curves could not show it.
The pole figure shows a window of the ray, and the approximants’ outermost poles lie beyond it: the drawing shows that the poles are on the cut and cannot show them filling it, which is the claim about the limit.
And nothing here shows the function’s other values. Past the cut the logarithm continues onto further sheets, differing by multiples of , and the approximants — being single-valued — see only the sheet they were built on. The walk around the singularity that reaches the other sheets is a different construction, and no quotient of polynomials performs it.
Still open: a series with no radius at all
Everything here assumed the series converges somewhere. Some do not converge anywhere except at nought — their coefficients grow like factorials — and they arise constantly: as expansions of integrals, of solutions to differential equations, of quantities in physics. Such a series has radius zero, so every method that reads a function off its disc of convergence says there is nothing to read.
A series that converges nowhere takes up that case, and it turns out to be a case where the Taylor sum is excellent if it is stopped at the right place, and where the Padé approximants — which need only coefficients, not convergence — converge anyway. Laurent series, which allow negative powers and describe a function around a singularity rather than avoiding it, are the other direction the same question points, and the equation a sequence satisfies is where the location and kind of a singularity turn into the growth of the coefficients.
Changing the form, not the centre
The move here is worth separating from the mathematics, because it applies well beyond series.
The Taylor series failed at and there were two responses. One changed where the information came from — a new centre, new derivatives, a walk of discs. The other kept the information and changed what was built from it. The second is cheaper, since the coefficients were already known, and in this case it was also better, since it reached every point off the cut at once rather than one disc at a time.
When an approximation fails, ask whether it is the information or the form that is inadequate. A polynomial through evenly spaced points failed because of where the points were, which was the information; a polynomial built from coefficients at nought failed because it was a polynomial, which was the form. The two failures look identical on a graph, and the repairs are completely different.
What links here
Computed from the collection, not written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The size of a number with no formula — both name analytic continuation, approximation, convergence
- A square wave built entirely out of round ones — both name convergence, orthogonality
- The nearest point of a flat thing — both name approximation, orthogonality
- The staircase that is not the diagonal — both name approximation, convergence
- What a map does to a circle — both name approximation, orthogonality
- Where Newton's method goes instead — both name complex numbers, convergence
Named objects
A dashed tag is an object no other essay names yet.
Analytic continuationApproximationComplex numbersConvergenceOrthogonalityPade approximantPolynomial approximationRadius of convergenceSingularityTaylor series