Algebra

One number under every bell

The area under e^(−x²) has no formula in terms of the usual functions, and yet the area under e raised to any downward quadratic is known exactly. Completing the square in the exponent moves and squeezes every such curve into the same one, so a single number — √π — pays for all of them.

Worth reading first: Completing the square, by completing a square · Where two roots run into each other.

The curve y=ex2y = e^{-x^2} is the most studied bell in mathematics, and the area under it has no formula. Not a difficult formula — none. There is no combination of powers, logarithms, exponentials and trigonometric functions whose derivative is ex2e^{-x^2}, and that is a theorem, due to Liouville, rather than a report that nobody has found one. The fundamental theorem converts an area into a search for an antiderivative, and here the search provably fails.

And yet the total area under the curve, from one end of the line to the other, is known exactly: it is π\sqrt{\pi}. So is the area under e3x2+5xe^{-3x^2 + 5x}, and under e(x7)2/50e^{-(x - 7)^2/50}, and under every curve of the form ee raised to a downward-opening quadratic. None of those has an antiderivative either. All of them are known, and they are known because one move reduces every one of them to the first.

The move is completing the square, performed in the exponent.

A quadratic in the exponent, completed. Two panels sharing an x-axis. Above, the parabola −x² + 2x with its top at x = 1 marked. Below, e raised to that parabola: a bell centred at the same x = 1, with peak height e^1, beside the faint unmoved bell e^(−x²).
Fig. 1 Above, the exponent x2+2x-x^2 + 2x, a downward parabola whose top is at x=1x = 1 with height 11. Below, ee raised to it: a bell centred at the same place, with its peak lifted to e1e^1. The faint curve is the plain bell ex2e^{-x^2}, and the new one is that bell moved one unit right and stretched upward by a factor of ee.

The exponent is a parabola with a top

Read the exponent first, because everything happens there.

x2+2x-x^2 + 2x is a quadratic, and completing its square rewrites it as (x1)2+1-(x - 1)^2 + 1. The completed form says two things at once: the parabola’s highest point is at x=1x = 1, and its height there is 11. Everywhere else it is lower by exactly (x1)2(x - 1)^2.

Now put that in an exponent. Since eA+B=eAeBe^{A + B} = e^A e^B, a constant added to the exponent becomes a constant multiplying the result:

ex2+2x=e1e(x1)2.e^{-x^2 + 2x} = e^{1} \cdot e^{-(x - 1)^2}.

That is the whole content of the top figure, and the two panels share an axis so the correspondence can be read vertically. The top of the parabola becomes the peak of the bell, at the same xx. The height of the top becomes a multiplier on the whole bell. And the part that curves away, (x1)2-(x - 1)^2, is the plain bell’s exponent, shifted.

The general statement has the same shape. For any a>0a > 0 and any bb,

ax2+bx=a(xb2a)2+b24a,-ax^2 + bx = -a\left(x - \frac{b}{2a}\right)^2 + \frac{b^2}{4a},

so

eax2+bx=eb2/4aea(xb/2a)2.e^{-ax^2 + bx} = e^{b^2/4a} \cdot e^{-a(x - b/2a)^2}.

Every such curve is the basic bell, moved to b/2ab/2a, squeezed sideways by aa, and lifted by eb2/4ae^{b^2/4a}. The quantity b2/4ab^2/4a is the same “half the coefficient, squared” that paid for the missing corner of the square, now paying for the height of a bell.

A quadratic in the exponent, completed. Two panels sharing an x-axis. Above, the parabola −0.5x² − 2x with its top at x = −2 marked. Below, e raised to that parabola: a bell centred at the same x = −2, with peak height e^2, beside the faint unmoved bell e^(−0.5x²).
Fig. 2 The same construction with a wider bell pulled the other way: the exponent 12x22x-\tfrac12 x^2 - 2x completes to 12(x+2)2+2-\tfrac12(x + 2)^2 + 2, so the bell is centred at 2-2, wider than the plain one by a factor of 2\sqrt 2, and lifted by e2e^2, about 7.47.4.

The second figure is worth comparing with the first rather than reading on its own. Nothing about the bell’s shape was ever in the linear term bxbx: that term only decides where the top of the exponent sits and how high. The shape — how quickly the curve falls away from its peak — lives entirely in aa, the coefficient of x2x^2. That division of labour is visible in the exponent before any exponential is taken, and completing the square is simply the act of reading it off.

Moving and squeezing, and what each does to area

Each of the three operations changes the area in a way that can be stated without computing anything.

Moving the bell sideways does nothing to its area. A region slid along the axis is the same region. So the area under e(xc)2e^{-(x - c)^2} is the area under ex2e^{-x^2} for every cc.

Squeezing it sideways divides the area. The curve eax2e^{-ax^2} is eu2e^{-u^2} with u=axu = \sqrt a\, x, which is the plain bell with its horizontal axis compressed by the factor a\sqrt a. Every vertical strip becomes a\sqrt a times narrower and keeps its height, so the area is divided by a\sqrt a.

Lifting it multiplies the area. Multiplying every height by eb2/4ae^{b^2/4a} multiplies the area by the same factor.

Put together, with II standing for the area under the plain bell:

eax2+bxdx=eb2/4a1a  I.\int_{-\infty}^{\infty} e^{-ax^2 + bx}\,dx = e^{b^2/4a}\,\sqrt{\frac{1}{a}}\; I.

Every bell's area from one number. 5 curves of the form e^(−ax² + bx), moved and squeezed, with a table of their areas beside them. Each measured area matches √(π/a)·e^(b²/4a): the areas range from 1.772 to 4.818.
Fig. 3 Five curves of the form eax2+bxe^{-ax^2 + bx}: the plain bell, two lifted and moved copies, a squeezed and lifted one, and a widened one. The table gives each measured area and that area divided by π\sqrt\pi. Every ratio is exactly 1/aeb2/4a\sqrt{1/a}\,e^{b^2/4a} — the three operations’ factors multiplied — so one number underlies all five areas.

The table makes the claim concrete. The areas range from about 1.771.77 to 4.824.82, and when each is divided by π\sqrt\pi what remains is 11, e1/4e^{1/4}, ee, 12e\tfrac12 e and 22 — the lift and the squeeze, and nothing else. The areas were measured by adding up thin strips under each curve, with no formula consulted, and the agreement holds to nine decimal places in every row.

That is what “reduces to one integral” means in practice. The family has two free parameters and infinitely many members, and the area of every member is π\sqrt{\pi} times a factor the completed square hands over for free. The one number nobody can get from an antiderivative has to be found once, by some other means, and then it never has to be found again.

A curve with no antiderivative, and a table anyway

It is worth being exact about what the completion does and does not deliver, because it is easy to overstate.

It delivers the area under the whole curve. It does not deliver the area under part of it. The area under ex2e^{-x^2} from 00 to 11, say, is a perfectly definite number, about 0.74680.7468, and completing the square is no help in finding it, because the missing antiderivative is still missing. The function that answers the partial question is given a name and defined as the area: the error function, erf(t)=2π0tex2dx\operatorname{erf}(t) = \frac{2}{\sqrt\pi}\int_0^t e^{-x^2}\,dx, scaled so that it runs from 00 to 11. Its values were tabulated by hand in the nineteenth century, by adding up strips exactly as the table above was produced, and every statistics textbook still prints some version of that table.

So the situation is peculiar and worth stating plainly. Every partial area of every bell reduces, by the completion, to a value of one tabulated function. Every whole area reduces to one number. Neither reduction produces an antiderivative, and neither needs one: a reduction to a single standard quantity is as good as a formula for anybody who has the quantity, and the reason the bell family is so tractable is that all of it reduces to so little.

The same thing is true of the exponential itself, in a way that is easy to forget. exe^x is not a formula either; it is a named function, defined by a property and tabulated, and it feels like a formula only because it is so familiar. erf\operatorname{erf} is less familiar and no less definite.

Why the one number is the square root of π

The other means is one of the best tricks in analysis, due to Poisson, and it works by making the problem bigger.

The area I=ex2dxI = \int e^{-x^2}\,dx is hard. Its square is

I2=ex2dxey2dy=e(x2+y2)dxdy,I^2 = \int e^{-x^2}\,dx \int e^{-y^2}\,dy = \iint e^{-(x^2 + y^2)}\,dx\,dy,

which is the volume under a surface standing over the whole plane — a round hill of height 11 at the origin, falling away the same way in every direction. The height at a point depends only on its distance rr from the origin, er2e^{-r^2}, and a round hill is best measured in rings rather than in strips.

The bell's area squared, cut into rings. On the left a disc shaded darker towards its centre, standing for the surface e^(−x² − y²), cut into 12 concentric rings. On the right the curve 2πr·e^(−r²), the volume each ring contributes per unit width, with the area under it shaded; that area is π, and the bell's own area is its square root.
Fig. 4 Left: the hill ex2y2e^{-x^2 - y^2} seen from above, shaded darker where it is higher and cut into twelve rings of equal width. Right: the volume each ring holds per unit of its width, 2πrer22\pi r\,e^{-r^2}, with the area under that curve shaded. The area is π\pi; the bell’s own area, measured separately, squares to the same number to nine places.

A thin ring at radius rr and width drdr has circumference 2πr2\pi r, so its area is 2πrdr2\pi r\,dr and the volume above it is about 2πrer2dr2\pi r\,e^{-r^2}\,dr. Adding up the rings:

I2=02πrer2dr.I^2 = \int_0^\infty 2\pi r\, e^{-r^2}\,dr.

And now the problem has changed character completely. The ring integrand 2πrer22\pi r\,e^{-r^2} has an antiderivative, πer2-\pi e^{-r^2}, because the extra factor of rr is exactly what the chain rule produces when er2e^{-r^2} is differentiated. So the integral is π(10)=π\pi(1 - 0) = \pi, and I=πI = \sqrt\pi.

The right-hand panel shows why the rings succeed where the strips failed. Near the centre the rings are tiny, so the hill’s greatest height contributes almost nothing; the ring curve starts at zero, rises to a peak just past r=0.7r = 0.7, and falls away. It is a completely different curve from the bell, with the same total — and unlike the bell it is the derivative of something elementary. The circle does the work: the factor 2πr2\pi r that turns strips into rings is the circumference, and it is where the π\pi in the answer comes from.

That is why a π\pi appears in the normal distribution, which has no circle anywhere in its statement. The density 12πσe(xμ)2/2σ2\frac{1}{\sqrt{2\pi}\,\sigma} e^{-(x - \mu)^2/2\sigma^2} is the bell moved to μ\mu and squeezed by a=1/(2σ2)a = 1/(2\sigma^2), and the constant in front is exactly the reciprocal of π/a=2πσ\sqrt{\pi/a} = \sqrt{2\pi}\,\sigma — the area the completed square predicts. The bell that assembles itself out of coin flips carries a π\pi because two independent bells, placed at right angles, make a round hill.

Two bells multiplied are a bell

The completion does more than compute areas. It explains why the family of bells is closed under an operation that, on its face, should make a mess.

Multiply two bells. Their exponents add, and the sum of two quadratics is a quadratic. Completing that sum gives a single bell back, with its own centre and its own width.

Two bells multiplied are a bell. Two bell curves, centred at -1 and 2 with widths 1 and 0.6, and their product rescaled to peak height one: a single narrower bell centred at 1.21, between the two and nearer the narrower one.
Fig. 5 A bell centred at 1-1 with width 11, a narrower one centred at 22 with width 0.60.6, and their product, rescaled to peak height one. The product is a single bell, centred at about 1.211.21 — between the two, but much nearer the narrower one — and narrower than either. Divided by the predicted bell, the product is the same constant at every point sampled.

The completed square names the answer precisely. Writing each bell as e(xmi)2/2si2e^{-(x - m_i)^2 / 2s_i^2} and calling wi=1/si2w_i = 1/s_i^2 the precision of each, the product’s exponent is

w1(xm1)2+w2(xm2)22,-\frac{w_1 (x - m_1)^2 + w_2 (x - m_2)^2}{2},

and completing the square in xx gives a bell with precision w1+w2w_1 + w_2, centred at

m=w1m1+w2m2w1+w2.m = \frac{w_1 m_1 + w_2 m_2}{w_1 + w_2}.

So precisions add, and the new centre is the average of the old ones weighted by precision. In the figure the narrow bell has almost three times the precision of the wide one, so the product sits nearly three-quarters of the way towards it. And because precisions add, the product is always narrower than both factors: combining two bells can only sharpen.

This one computation is behind a large share of practical statistics. When two independent measurements of the same quantity each carry a bell-shaped uncertainty, the combined uncertainty is their product, and the rule for combining them — weight each by one over its variance — is the completed square above. It is also why a sum of two independent normal variables is normal again: the density of a sum is a convolution, the convolution of two bells involves an integral of a product of bells, and completing the square in the integration variable collapses it to a single bell whose variance is the sum of the two.

Several variables, and a determinant where the width was

The completion is not a one-dimensional trick, and the version in several variables is where it becomes most useful.

In two variables a downward quadratic exponent is (ax2+2bxy+cy2)+(linear terms)-(ax^2 + 2bxy + cy^2) + (\text{linear terms}), and its level curves are ellipses — tilted, stretched, centred somewhere. Completing the square in two variables means doing what the one-variable case did, twice: first gather every term containing xx into a square, which leaves a quadratic in yy alone, and then complete that. The result is a centre, which absorbs the linear terms exactly as b/2ab/2a did, and a sum of two squares in new tilted coordinates. That repeated completing is Sylvester’s procedure for reducing a quadratic form, and the requirement that the exponent open downward in every direction is the requirement that both squares come out with minus signs.

Then the volume under the resulting hill is the product of two one-dimensional bells, and each contributes a π\sqrt\pi divided by the square root of its own squeeze. The product of the squeezes is the determinant of the matrix (abbc)\begin{pmatrix} a & b \\ b & c \end{pmatrix}, so

e(ax2+2bxy+cy2)dxdy=πacb2.\iint e^{-(ax^2 + 2bxy + cy^2)}\,dx\,dy = \frac{\pi}{\sqrt{ac - b^2}}.

In nn variables the answer is πn/2\pi^{n/2} divided by the square root of the determinant. The width that the single aa controlled in one dimension is controlled by a whole matrix in many, and the number that measures it is the determinant — the factor by which a linear map scales volumes, which is exactly what it is doing here, because the tilted, stretched hill is the round one pushed through a linear map.

This is the formula behind the multivariate normal distribution and behind a large part of theoretical physics, where the integrals of ee to a quadratic in many variables are called Gaussian integrals and are almost the only ones that can be done exactly. A theory whose central integral is Gaussian is a theory that can be solved; the others are solved by expanding around a Gaussian and hoping the corrections are small.

The same shift, taken into the imaginary

The completion has one more application, and it is the one that makes the bell special among all curves.

Multiply the bell by a wave, cos(kx)\cos(kx), and ask for the area. For k=0k = 0 it is π\sqrt\pi. As kk grows the wave oscillates faster under the bell, positive and negative lobes cancel more and more thoroughly, and the area shrinks.

A bell times a wave. The curves e^(−x²)cos(kx) for k = 0, 2, 4. The first is the bell itself; the others oscillate inside it and their areas shrink as 1.772, 0.652, 0.032, following √π·e^(−k²/4).
Fig. 6 The bell multiplied by cos(kx)\cos(kx) for k=0k = 0, 22 and 44. The faster the wave, the more its lobes cancel: the areas, measured by thin strips, are 1.772451.77245, 0.652050.65205 and 0.032460.03246, and πek2/4\sqrt\pi\, e^{-k^2/4} gives the same three numbers. The area shrinks as a function of kk that is itself a bell.

The measured areas follow πek2/4\sqrt\pi\, e^{-k^2/4} exactly, and the formula comes from the same completion with an imaginary linear term. Since cos(kx)\cos(kx) is the real part of eikxe^{ikx}, the area is the real part of

ex2+ikxdx,\int e^{-x^2 + ikx}\,dx,

and completing the square with b=ikb = ik gives eb2/4π=ek2/4πe^{b^2/4} \sqrt\pi = e^{-k^2/4}\sqrt\pi, because (ik)2=k2(ik)^2 = -k^2. The “lift” factor eb2/4e^{b^2/4}, which made the areas in the earlier table larger, becomes a factor that makes this one smaller, for no reason except that the square of an imaginary number is negative.

This is the statement that the Fourier transform of a bell is a bell. Breaking a curve into waves of every frequency — the analysis that builds a square wave out of round ones — assigns to each frequency kk the amount of that wave present, and for the bell the answer is another bell, in kk rather than in xx. A narrow bell in xx has a wide one in kk and a wide one in xx has a narrow one in kk, with the product of the two widths fixed. That trade-off is the uncertainty principle of signal processing, and in quantum mechanics it is Heisenberg’s; for the bell, and only the bell, it is achieved with equality.

What the pictures cannot show

The imaginary shift is not drawn, because it cannot be. The completion moves the bell’s centre to ik/2ik/2, a point off the real line, and the justification that the area is unchanged by such a move is Cauchy’s theorem about integrals in the complex plane — the region under a curve has become a contour integral, and no strip picture represents it. The fourth figure measures the areas and finds they agree with the formula; the argument that they must is invisible in it.

The rings are a picture of a volume, and the volume is only suggested. The disc on the left is shaded, not raised, and the reader is asked to imagine the hill. What the figure actually establishes is the right-hand panel’s area, a one-dimensional integral; the step from the double integral to the ring integral — that a round region can be cut into rings with no volume lost — is the change to polar coordinates, and it is stated rather than drawn.

Every integral here runs over the whole line, and every figure stops at the edge of the canvas. The areas were measured over ranges wide enough that the tails beyond them are smaller than 102010^{-20}, which is why the agreement holds to nine places. But the curves themselves are drawn over a few units, and a reader could not tell from the drawing that the tails are negligible rather than merely off the page. For the bell they are negligible, falling faster than any power; for a curve with heavier tails the same picture would be quietly wrong.

Where the completion stops working: integrals that are only nearly bells

Completing the square computes integrals of ee to a quadratic exactly. Most integrals worth doing have a different exponent, and the natural question is what survives when the exponent is only approximately quadratic.

The answer is Laplace’s method, and it is the completed square used as an approximation. If an integrand eNf(x)e^{N f(x)} has a single sharp peak where ff is largest, then near that peak ff is close to its own Taylor quadratic, and as NN grows the peak sharpens until everything away from it is negligible. The integral then approaches the bell integral for that quadratic — completed square, π\sqrt\pi and all. Applied to n!=0xnexdxn! = \int_0^\infty x^n e^{-x}\,dx, whose integrand peaks at x=nx = n, it produces Stirling’s formula n!2πn(n/e)nn! \approx \sqrt{2\pi n}\,(n/e)^n, and the 2π\sqrt{2\pi} in that approximation is this essay’s π\sqrt\pi, arrived at by the same route.

Two questions remain genuinely delicate there. How good the approximation is depends on the next terms of the Taylor expansion, and the corrections form an asymptotic series that usually diverges — useful when truncated early, meaningless when summed. And when the peak is not sharp, or there are several, or the maximum sits on the edge of the range, the method needs modification case by case. In several variables the same completion runs on a quadratic form, the π\sqrt\pi becomes πn/2\pi^{n/2} divided by the square root of a determinant, and how fast a bell arrives in the central limit theorem is a question about exactly those correction terms.

One corner, paid for once

The square that was completed with a missing corner and the bell whose area is π\sqrt\pi are the same operation in two settings. In the square, the corner’s size (b/2)2(b/2)^2 was forced by symmetry, and paying for it turned an awkward shape into one whose side could be read off. In the exponent, the same b2/4ab^2/4a is forced by the same symmetry, and paying for it turns an awkward curve into a moved, squeezed, lifted copy of one standard curve.

What makes the exponent version remarkable is how much it leverages. The one integral that has no antiderivative is computed once, by a trick that borrows a circle; every other integral of the family then follows from arithmetic. The products of bells, the sums of normal quantities, and the frequencies inside a bell all come out as bells again, because in every case the operation adds exponents, and a sum of quadratics can always be completed back into a single square.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

Completing the squareComplex numberse, the numberFourier analysisIntegralNormal distributionPiQuadratic polynomialsScalingVolume