An integral that cannot be a whole number
Worth reading first: A tail too small to be a whole number · Adding up rectangles until they stop being rectangles.
The proof that is not a fraction fits in a paragraph because the series for hands over both halves of the squeeze at once: the factorial makes the head whole, and the same factorials make the tail small.
has no such series. Every classical expression for it converges too slowly to be squeezed, and for two thousand years nobody could show it was not a fraction — Aristotle assumed it was, Archimedes bracketed it, and the question stayed open until Lambert settled it in 1761 with a continued fraction for the tangent. The proof that fits on a page is Ivan Niven’s, from 1947, and it is the same template with a manufactured multiplier.
The template, and what has to be supplied
The argument that settled has a shape: assume the number is , find a multiplier that makes some computable quantity a whole number, and show that quantity is strictly between zero and one.
For the multiplier was and the quantity was a series tail. Here the multiplier is a polynomial and the quantity is an integral, and both have to be constructed. Suppose with and whole numbers, and set
Two claims are needed, and they are the two panels of the figure.
is a whole number. This is the hard half, and the polynomial was built for it.
for large . This is the easy half, and it is a bound on a positive function times a bounded one.
Together they are the contradiction, because there is no whole number strictly between zero and one — the same observation that finished the last rung.
Why the integral is a whole number
Integrating by parts repeatedly turns the integral into a sum of values of and its derivatives at the two endpoints, with alternating signs. The polynomial has degree , so after steps the derivatives vanish and the process stops; what is left is
So is whole as soon as every derivative of at and at is whole, and that is what the polynomial’s peculiar shape buys.
At : expand and every term has degree at least , so the derivatives of order below all vanish. The derivative of order at zero is the coefficient of times , divided by the in the denominator — and is a whole number. The coefficients themselves are whole because and are.
At : the polynomial satisfies , which is a two-line check on the definition, so the derivatives at are the derivatives at up to sign. Nothing more is needed.
The right-hand panel of the figure computes those derivatives at in exact whole-number arithmetic for a stated and reports that every one of them is whole. That is the step usually asserted in the textbook proof, and it is the step a reader has no reason to believe on being told.
Why the integral is small
The other half is a bound, and it needs nothing clever.
On the interval, , so — the product of two positive factors each at most its own maximum. And is between and there. So
The numerator grows exponentially in and the denominator grows factorially, so the bound goes to zero — and past some computable it is below one. That can be found by search rather than estimated, and the figure finds it.
Note what has just happened. The number was free all along; the polynomial exists for every degree. The proof chooses after seeing and , large enough for the bound to bite. That order matters: for any particular there is a fraction for which the bound is useless, and for any particular there is an for which it is not. The quantifiers are the proof.
Where the difficulty actually sits
It is worth being precise about which part of this was hard, because the proof is short and its shortness is misleading.
The bound is routine. Anyone who has seen an exponential lose to a factorial can produce it, and the tail argument for is the same comparison doing the same job.
The integrality is not routine at all, and the polynomial is not something a reader would think of. It has to vanish to order at both endpoints (so that low derivatives are zero), have whole-number coefficients after multiplying by (so the high ones are whole), be symmetric about the midpoint (so one endpoint’s computation serves for both), and be small on the interval (so the bound works). Four requirements, one construction, and the construction was found in 1947 rather than in 1761.
Lambert’s original proof is quite different. It develops as a continued fraction, shows the expansion is infinite for every non-zero rational , and concludes that forces to be irrational. It is a better proof in the sense that it proves more — every non-zero rational has irrational tangent, so the circular functions take no non-zero rational to a rational — and a worse one in the sense that convergence of a continued fraction with function entries takes real analysis to justify.
Niven’s is the one that fits on a page, and the reason it fits is that all the ingenuity is in a formula that can be written in one line and checked in three.
The polynomial, and why every requirement is forced
The construction looks arbitrary until each of its features is matched to the line of the proof that needs it, and then there is nothing arbitrary left.
Why and rather than any other pair of factors. The two factors vanish at and at respectively, each to order . That is what kills the low derivatives at both endpoints and leaves only the ones the factorial makes whole.
Why the same exponent on both. It makes , so the arithmetic at one endpoint serves for the other. Different exponents would double the work and buy nothing.
Why inside rather than outside. Writing rather than keeps every coefficient a whole number. This is where the assumption that is a fraction is actually used, and it is used exactly once.
Why divide by . Without it the derivatives are still whole and the integral is not small — the factorial is what turns into something that goes to zero. It is also what makes the derivatives just whole rather than divisible by a large factor, which is not needed but is where the elegance is.
Why and not something else. Two reasons. It vanishes at both and , which is what makes the integration by parts terminate cleanly, and its own derivatives cycle with period four, which is what produces an alternating sum of even derivatives rather than a mess. Any function with those two properties would do, and is the one that has them because is defined by it.
Five features, five lines of the proof. A construction with no spare parts is usually a construction someone worked backwards from the contradiction, and this one visibly was.
What it does not settle
The same three limitations as the last rung, and one of them is sharper here.
It does not give transcendence. could still, for all this argument says, be a root of some whole-number polynomial of high degree. That it is not is Lindemann’s theorem of 1882, and it is a much heavier result — it needs the transcendence of for algebraic non-zero , and Euler’s identity to connect the two.
That distinction is not academic. Squaring the circle needs transcendence and not irrationality: the constructible numbers include plenty of irrational ones, so the impossibility of squaring the circle waited on Lindemann rather than on Lambert, a hundred and twenty years later. Doubling the cube is the contrasting case, where irrationality is easy and the impossibility still needed a statement about degrees rather than about fractions.
It does not measure how irrational is. The proof shows a quantity is small; it extracts no approximation statement. What is known — that cannot be approached faster than about the eighth power of the denominator — comes from entirely different machinery, of the kind the rung above introduces, and the exponent has been reduced slowly and painfully over decades.
It says nothing about , , or . Whether any of those is irrational is open. At least one of and is transcendental, because if both were algebraic then and would be roots of a quadratic with algebraic coefficients — but which one, nobody knows.
A worked instance of the arithmetic
It helps to see the integrality claim on actual numbers rather than in general.
Take , which is not and is close enough for the picture, and . The polynomial is . Expanding by the binomial theorem gives eight terms with whole-number coefficients; multiplying by shifts them to degrees through ; and the derivative of order at zero is that term’s coefficient times , over .
The ratio is a product of consecutive whole numbers and is therefore whole. Every derivative comes out whole, and the figure prints the first several of them — numbers with a dozen digits, exactly, in big-integer arithmetic, because the claim is whole and a floating-point computation of a twelve-digit quantity cannot tell a whole number from its neighbour.
The size of those numbers is worth noticing. They are enormous, and is under one. The integral is a very small difference of very large whole numbers, and that is not a defect of the presentation — it is what the proof is doing. A quantity assembled out of numbers near and known to lie strictly between zero and one has been forced to be a whole number, and there is none available.
How long it took, and what the delay was about
The gap between Archimedes and Lambert is nineteen centuries, and it is worth asking what the missing ingredient was, because the answer is not ingenuity.
Archimedes bracketed between and by inscribing and circumscribing ninety-six-sided polygons, and the method is exact and mechanical: it can be pushed to any accuracy by doubling the sides. What it produces is a sequence of fractions, and no sequence of fractions decides whether its limit is one.
The missing ingredient was an infinite process that says something about a number’s kind rather than its size. Series, continued fractions and integrals all do that; polygon-doubling does not, because every stage of it is a rational approximation and rational approximations are exactly what the question is about.
That is why the seventeenth century’s arrival of infinite series changed the situation and the eighteenth century’s mastery of them settled it. Euler had ’s continued fraction in 1737; Lambert had the tangent’s in 1761; and Niven’s polynomial is a nineteenth-century object used in 1947 to make the answer fit on a page.
There is a general lesson in the shape of it and this site keeps meeting it. A question about what a number is cannot be answered by better approximations to it. Counting what has no formula is the same observation about the primes: an enormous table of them was available for a century before anything was proved, and the proof when it came was about a completely different object.
The same squeeze, on a number nobody could reach
It is worth saying what the template has and has not managed since, because was not the end of the line and the pattern of the successes is informative.
Every for rational non-zero , and every , fell to Lambert’s machinery in the eighteenth century. The continued fractions were there and the numbers were reachable.
— the sum of the reciprocal cubes — held out until 1978. Euler had evaluated the sum of the reciprocal squares as in 1735, which settles that one, and the odd exponents resisted everything for two hundred and forty years. Apéry’s proof is the same template: a quantity that would have to be a whole number, and a demonstration that it is too small. Producing the quantity took two auxiliary sequences satisfying a recurrence nobody had seen, and the talk announcing it was widely disbelieved.
is open. So is the irrationality of the Euler–Mascheroni constant, of , and of — no, that last one is settled, by Gelfond and Schneider, and it is worth naming because it shows how uneven the ground is: a number that looks far more exotic than is known and a sum of reciprocal fifth powers is not.
The unevenness has a reason. The template needs a quantity that is forced to be whole, and that comes from a structure rather than from a number. Where a number sits inside a series with factorial denominators, or inside an integral against a polynomial, or inside a linear form with algebraic coefficients, there is something to force. Where it does not, there is nothing to try, and has been in that position for a very long time.
What the picture cannot show
The left panel draws and not , and the integral is of the second. The difference does not matter for the bound — is between zero and one — but the curve on the page is not the integrand, and a reader looking for the area under the drawn curve is looking at the wrong quantity by a factor of up to .
The right panel prints derivatives at zero and says every one is whole. It cannot show why, which is the binomial expansion and the ratio of factorials, and the reason is that the why is arithmetic rather than shape.
And neither panel can show the choice of , which is the logical heart. The proof works because is chosen after and ; a picture at fixed degree shows a single case, and the quantifier is invisible in every one of them.
Where the ladder goes next
Above: transcendence. Being no root of any whole-number polynomial is a stronger property than being no fraction, it needs a different kind of argument, and the accessible version of that argument is about how fast a number can be approached rather than about whether anything hits it exactly.
Two debts. Lambert’s continued fraction for the tangent is quoted and not drawn, and it is the historically decisive proof. And Lindemann’s theorem is named three times here as the thing that settles squaring the circle, and no essay on this site proves it — which is honest, since it is several pages of exponential polynomials, and is still a promise.
What Niven built
When a number arrives without a series that is both whole and small, build a polynomial that is.
Every requirement on Niven’s answers one line of the template. It vanishes to high order so the low derivatives are whole; it carries so the high ones are; it is symmetric so one endpoint suffices; and it is tiny so the integral is. The proof is one paragraph because the construction absorbed all of it, which is the usual arrangement when a two-thousand-year-old question turns out to have a one-page answer.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The square that cannot shrink — both name irrationality, proof by contradiction
Named objects
A dashed tag is an object no other essay names yet.
DerivativeFactorialIntegralityIntegration by partsIrrationalityPiPolynomialProof by contradictionSqueezeTranscendence