A rectangle cut by a curve
Worth reading first: Area is the undoing of slope · An endless region with a finite area.
The formula for integration by parts is usually introduced as a consequence of the product rule. If and are functions, then , and integrating both sides gives
evaluated between the endpoints. It is then used as a manoeuvre: an integral that cannot be done is traded for one that can, by moving a derivative from one factor to the other.
That derivation is correct and it hides a picture that is simpler than the product rule. When is an increasing function of , the two integrals in the formula are two regions of the plane, and the formula says that together they fill a rectangle.
The two regions a curve makes
Look at what each integral measures. , taken from to , is the area between the curve and the horizontal axis — the ordinary area under a graph, adding up vertical strips. , taken from to , is the same construction with the roles of the axes swapped: the area between the curve and the vertical axis, adding up horizontal strips.
For an increasing curve those two regions do not overlap, and together with the small rectangle in the corner they tile the big rectangle whose far corner is the curve’s endpoint. So
That is the whole formula, and the picture proves it without any derivative. In the figure the region below the curve has area , the region beside it , and the difference of the two rectangles is , which is their sum. Both regions were measured by strips, independently, and the rectangle was not consulted.
What the formula does is visible too. If one of the two regions is hard to measure and the other is easy, the rectangle converts the easy one into the hard one. There is nothing more to integration by parts than that: the curve cuts a rectangle into two pieces, and knowing one piece and the rectangle is knowing the other.
The usual form, with and as functions of some third variable , is the same picture with the curve described parametrically. As runs from to the point traces the curve, , , and the formula becomes the familiar
The product rule is the same rectangle, growing
The picture and the product rule are not two proofs of one formula that happen to agree. They are the same observation at two scales.
Take the rectangle with sides and , and let both sides grow a little, by and by . The new rectangle is the old one plus three pieces: a strip of width and height down one side, a strip of height and width along the top, and a small corner of size where the two strips meet. So
Divide by the change in whatever and depend on and let the changes shrink. The corner is a product of two small quantities, so it vanishes faster than either strip, and what survives is the product rule, . It is the same corner that completing the square had to pay for — there it was the whole point, and here it is the part that disappears.
Now add up the growth all along the curve, from the small rectangle to the big one. The strips of the form are vertical slices under the curve, and they add up to ; the strips of the form are horizontal slices beside it, and they add up to ; the corners add up to nothing. The total growth is the big rectangle minus the small one. Integration by parts is the product rule summed, and the picture at the top of this essay is what the sum looks like when every strip has been laid in place.
The same accounting works with no limit at all. For two sequences, the change in a product from one step to the next is exactly — the corner is absorbed by taking one step later — and adding over gives summation by parts, the discrete formula Abel used to control sums whose terms oscillate. In the picture the curve becomes a staircase and the two regions become two interlocking stacks of rectangles, and the corners, instead of vanishing, are exactly accounted for by the half-step shift. The same move turned the fundamental theorem into a telescoping sum, and it is no accident that it works here: both formulas are statements that differences add up to the change between endpoints.
The logarithm, measured from its side
The classic first example is the integral of , which has no obvious antiderivative. The textbook trick writes as and integrates by parts. The picture does something more transparent.
The region to the left of the logarithm is the region under the exponential, seen sideways, because describes the same curve with the axes swapped. The area under from to is , which needs nothing but the fact that the exponential is its own derivative. So the area under the logarithm is the rectangle less : .
In general this gives the integral of any inverse function for free. If is increasing and its integral is known, then the integral of is a rectangle minus the integral of :
The arcsine, the arctangent, the logarithm and every -th root are all inverse functions of things with easy integrals, and every one of their integrals comes out this way with no cleverness. The formula was published in this form by Charles-Ange Laisant in 1905, which is surprisingly late for something that is, in the picture, one rectangle and two pieces of it.
A rectangle the curve does not reach
The picture has a second use when the rectangle and the curve do not quite agree — and it produces an inequality that sits under a great deal of analysis.
Start both regions at the origin, with an increasing curve through . Take the region under the curve up to and the region beside it up to , where is not necessarily the height the curve has at .
When the two regions meet at the corner and fill the rectangle exactly, which is the parts formula again. When is anything else, one region sticks out past the rectangle’s edge and the two together more than cover it.
That is Young’s inequality: for any increasing with ,
with equality exactly when . The proof is the second figure: the rectangle is covered by the two regions, sometimes with overlap to spare, never with a gap.
Choosing turns it into the version that gets used. The two areas are and , where , so . For that is the elementary fact that , the gap between the arithmetic and geometric means; for other exponents it is the inequality from which Hölder’s inequality follows, and Hölder’s inequality is the tool that bounds the size of a product by the sizes of its factors throughout analysis — including the Cauchy–Schwarz inequality as the case , and the pairing of exponents and is the same duality that turns circles into diamonds and squares. A good deal of analysis rests on one rectangle that a curve either fills or overfills.
The factorials are areas
Integration by parts earns its reputation when it is applied repeatedly, and the first example is the one Euler used to extend the factorial beyond the whole numbers.
Consider the area under over the whole half-line. For it is the area under , which is . For larger , integrate by parts with and : the boundary term vanishes at both ends, and what remains is
Each area is times the one before. Starting from , the areas are — the factorials.
The picture shows why the recurrence is plausible before any formula. Each curve peaks further to the right and higher than the last, and its area is dominated by the region around the peak at , where the height is . That peak height times a width of about is roughly the area — which is Stirling’s approximation read off the curve, and the width comes from treating the peak as a bell, exactly as every bell’s area comes from one number.
What the area formulation adds is that it makes sense when is not a whole number. The area under is perfectly definite, and so the factorial of one half exists; it turns out to be , a consequence of the bell’s area. The function defined this way — shifted by one, for historical reasons — is Euler’s gamma function, , and the recurrence is the one integration by parts produced above, now valid for every positive . The requirement that the region have finite area is the threshold for improper integrals applied at the end near zero, where blows up exactly when .
π from powers of a sine
The second classic repetition produces as an infinite product, and it is the same move with a different pair of functions.
Let be the area under over a quarter turn, from to . Integrating by parts once — splitting as and using — gives
With and , every follows: the even ones are times a fraction and the odd ones are fractions alone.
Now use the ordering the picture displays. Every curve lies below the one before, so the areas decrease: . Divide through by ; the right-hand ratio tends to , so the ratio of the even area to the odd one is squeezed to as well. Writing that ratio out with the recurrence and solving for gives
John Wallis’s product of 1656. After ten factors it is ; after a thousand, , still short of — the partial products approach from below and the error after factors is about , which is slow. Wallis found it by a long process of interpolation that he could not justify; the argument by areas is Euler’s, and it needs only the recurrence and the fact that the curves are nested.
There is a connection hiding here that is worth pointing out. Wallis’s product and Stirling’s formula are two views of one fact. The constant in Stirling’s formula can be derived from Wallis’s product, and historically that is how de Moivre and Stirling pinned it down: de Moivre had the formula for with an unknown constant, and Stirling identified the constant as from Wallis. The powers of a sine over a quarter turn and the areas under are both computed by integrating by parts, and they meet at the same .
What the rectangle requires
The picture has a hypothesis the formula does not, and it is worth seeing where it breaks.
The rectangle argument needs the curve to be increasing, so that the region under it and the region beside it do not overlap. If the curve rises and then falls, the region beside it is traced twice — once on the way up and once on the way down, with opposite signs — and “the area between the curve and the vertical axis” has to be read as a signed quantity. The formula survives, because it is proved by the product rule and the product rule does not care about monotonicity; the picture needs the signs put in by hand.
Neither of the repeated applications, the factorials and the sine powers, is a single rectangle at all. Each step of the recurrence is one application of the formula to a different pair of functions, and a picture of the step would be a different rectangle for each. What the figures show instead is the outcome — the sequence of areas and the pattern among them — and the recurrence is the reason for the pattern, stated in the text rather than drawn.
What the pictures cannot show
Boundary terms that vanish at infinity. The factorial calculation throws away because it is zero at both ends, and the zero at infinity is a statement about a limit — that eventually beats every power — which no figure of finite width can display. The curves in the factorial figure visibly approach the axis by ; that they stay there, and faster than grows, is taken from the exponential’s properties, not from the drawing.
Which choice of parts works. Integration by parts is a choice — which factor to differentiate and which to integrate — and a bad choice makes the problem worse rather than better. The rectangle shows why the formula is true for any split; it gives no hint of which split will help. Choosing and in the factorial integral raises the power of at every step instead of lowering it, and the recurrence runs away rather than terminating. The skill is in the choice, and no picture of the formula contains it.
The measured areas are for whole numbers only. The factorial figure computes areas at . The claim that the same construction gives a smooth function of , and that its value at a half is , is proved from the integral, not seen in the table.
Still open: how much repetition can tell
Applied once, integration by parts moves a derivative across a product. Applied times to the remainder of a Taylor polynomial, it produces the exact integral form of the error of a Taylor approximation, and it is how the error in approximating a smooth function by polynomials is controlled. Applied to a sum rather than an integral, in the discrete form called summation by parts, it is how a power series is read from inside at the edge of its interval. Applied to a Fourier coefficient, each application gains a factor of the frequency in the denominator, which is exactly why a smooth function’s coefficients fall off quickly and a function with a corner’s fall off only like .
The systematic version is the Euler–Maclaurin formula, which relates a sum to an integral by applying parts over and over, with the Bernoulli numbers appearing as the coefficients. Its correction terms form an asymptotic series that usually diverges: taking more corrections improves the approximation up to a point and then makes it worse, and where the best stopping point lies depends delicately on the function. For the factorial the series is Stirling’s, and the optimal truncation and its error are known precisely. For a general sum there is no universal answer, and when an Euler–Maclaurin series can be summed by some regularisation to give an exact result is a question that in several directions is still being worked out.
Two pieces of one rectangle
The formula for integration by parts has two faces and they should be kept together. As a formula it is the product rule integrated, a tool for trading one integral for another, and its power comes from repetition: the factorials, the gamma function, Wallis’s product and Stirling’s constant all fall out of applying it again and again to a well-chosen pair.
As a picture it is a curve cutting a rectangle. The area under the curve and the area beside it are two pieces of one shape, and knowing either piece and the rectangle is knowing the other. That picture gives the integral of every inverse function at no cost, and when the rectangle is moved off the curve it gives Young’s inequality — an overfilled rectangle — and through it the inequalities that bound products throughout analysis.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The area that names the number — both name area, fundamental theorem, integral, logarithm
- A circle unrolled into a triangle — both name area, pi
- A rectangle grown on two sides — both name area, integral
- An integral that cannot be a whole number — both name factorial, pi
- Counting what has no formula — both name integral, logarithm
- Round is not the only way to be the same width — both name area, pi
Named objects
A dashed tag is an object no other essay names yet.
AntiderivativeAreaFactorialFundamental theoremInequalityIntegralInverse functionLogarithmPiRecurrence