Analysis

The corners go first

Fourier was not decomposing waves for the pleasure of it. He was solving the flow of heat, and the whole apparatus exists because each harmonic fades at a rate set by the square of its frequency — which is why a sharp profile smooths instantly and why the flow cannot be run backwards.

Worth reading first: A square wave built entirely out of round ones · Where the coefficients come from.

Heat a metal ring so that half of it is hot and half is cold, then leave it alone. What happens next is completely determined, and it is determined one harmonic at a time.

A square profile of heat, spreadingThe same profile at four times, each drawn from the same harmonics with each one damped by the exponential of minus its frequency squared times the time. The corners go first.−ππ-11position around the ringtemperatureat the startt = 0.005t = 0.05t = 0.4a square profile on a ring, left to spread: harmonic m fades by exp(−m²t), so the 41th term is gone 1,681 times fasterthan the firstthe corners disappear immediately and the shape that remains is a single sine — which is why running the flowbackwards is hopeless: the information in the corners has been divided by a number this large
Fig. 1 A square profile of temperature on a ring, at four moments. The corners are gone almost immediately; what survives after a while is a single sine wave, fading.

The sharp edges vanish first, and they vanish fast. What is left after a short interval is a rounded shape, and what is left after a long one is a single hump — the first harmonic and nothing else, decaying steadily toward a uniform ring.

Why a harmonic is the right thing to track

Temperature spreads according to a rule relating the change over time to the curvature in space:

ut=2ux2.\frac{\partial u}{\partial t} = \frac{\partial^2 u}{\partial x^2}.

Where the profile is bent upwards, it warms; where it bends down, it cools; where it is straight, nothing happens. That is the whole physical content, and as a statement about a general profile it is unhelpful, because the change at each point depends on the shape near that point and the shape is changing everywhere at once.

Now try it on a single harmonic. Differentiating sin(mx)\sin(mx) twice gives m2sin(mx)-m^2\sin(mx): the same shape, scaled. So a profile that starts as a pure sine of frequency mm has a rate of change proportional to itself, which is the defining property of the exponential, and its amplitude must decay as

em2t.e^{-m^{2}t}.

A sine is an eigenfunction of the operation governing the flow. It does not change shape, only size. That single fact is what makes the whole problem tractable: for a sine the partial differential equation is an ordinary one, and its solution is a decaying exponential with a rate that is the square of the frequency.

The target, multiplied by one harmonic at a timeFour panels, each showing the square wave multiplied by a single sine. The areas cancel exactly except against the harmonics the wave actually contains.−ππ-11× sin 1xarea = 4.000coefficient 1.273−ππ-11× sin 2xarea = 0no such harmonic in it−ππ-11× sin 3xarea = 1.333coefficient 0.424−ππ-11× sin 4xarea = 0no such harmonic in iteach panel is the target multiplied by one harmonic, with the area above the axis in one colour and the area below in the otheragainst its own harmonic the two do not cancel; against any other they cancel exactly, which is what makes one coefficient extractable without disturbing the rest
Fig. 2 The decomposition that makes the trick usable. Any starting profile is split into harmonics by projecting onto each one in turn, and the split is possible because the harmonics do not interfere.

The method is therefore three steps: split the initial profile into harmonics, let each one decay at its own rate, add them back up. Nothing about the flow couples the harmonics, so the middle step is a list of independent multiplications, and the difficulty of the problem has been moved entirely into the projection.

The square of the frequency is everything

The rate is m2m^2 and not mm, and that exponent decides the character of the whole subject.

The ninth harmonic decays 8181 times faster than the first. The forty-first decays 1,6811{,}681 times faster. At the time when the fundamental has faded by a factor of ee, the forty-first harmonic has faded by e1681e^{1681}, which is a number with seven hundred digits.

The spectrum of the square waveOne bar per harmonic, its height the size of that harmonic's coefficient. The 11 non-zero coefficients fall away like 1 over m to the power 1.00.0.320.640.951.27123456789131721harmoniccoefficientthe square wave, harmonic by harmonic: 11 of the first 21 are non-zero, and their sizes fall like 1 / m^1.00the exponent is fitted to the bars rather than taken from the formula, and it is what says how smooth the target is:a jump gives 1, a corner gives 2, and a smooth function falls faster than any power
Fig. 3 The starting profile’s spectrum: odd harmonics falling away like 1/m1/m. The flow multiplies each of these bars by em2te^{-m^2 t}, which at any positive time drives the right-hand end of this picture to nothing while barely touching the left.

So the flow is not a gentle smoothing that acts on everything a little. It is a filter that annihilates the high frequencies and leaves the low ones almost untouched, and its cutoff moves down through the spectrum as time passes. Fine detail is not blurred; it is deleted.

A square profile of heat, spreadingThe same profile at four times, each drawn from the same harmonics with each one damped by the exponential of minus its frequency squared times the time. The corners go first.−ππ-11position around the ringtemperatureat the startt = 0.001t = 0.01t = 0.3a square profile on a ring, left to spread: harmonic m fades by exp(−m²t), so the 41th term is gone 1,681 times fasterthan the firstthe corners disappear immediately and the shape that remains is a single sine — which is why running the flowbackwards is hopeless: the information in the corners has been divided by a number this large
Fig. 4 The same profile at a different set of times, with one very early frame. Even at t=0.001t = 0.001 the corners have rounded — the harmonics above the thirtieth are already down by a factor of a thousand — while the overall shape is essentially untouched.

This is why the profile becomes smooth immediately rather than gradually. At any time t>0t > 0 the coefficients carry a factor of em2te^{-m^2 t}, which falls faster than any power of mm, so the profile at any positive time is infinitely differentiable — no matter how ragged it was at the start. The solution is smoother than its own initial condition, instantly, and a square profile with jumps in it becomes an analytic function in no time at all.

Where the heat goes

The total amount of heat on the ring does not change — nothing is being added or taken away — and the decomposition says where it sits.

The zeroth coefficient, the average temperature, is multiplied by e02t=1e^{-0^2 t} = 1 at every time: it never decays, because a uniform profile has no curvature anywhere and therefore no reason to change. Every other coefficient decays. So the flow drives any starting profile toward its own average, and the average is fixed by the start.

A square profile of heat, spreadingThe same profile at four times, each drawn from the same harmonics with each one damped by the exponential of minus its frequency squared times the time. The corners go first.−ππ-11position around the ringtemperatureat the startt = 0.4t = 1.2a square profile on a ring, left to spread: harmonic m fades by exp(−m²t), so the 41th term is gone 1,681 times fasterthan the firstthe corners disappear immediately and the shape that remains is a single sine — which is why running the flowbackwards is hopeless: the information in the corners has been divided by a number this large
Fig. 5 The late stages, on a longer clock. By t=0.4t = 0.4 only the fundamental is left in any visible amount; from there the shape does not change at all and only its height does, falling by a factor of ee every unit of time.

That late behaviour is worth naming because it recurs across the whole of applied mathematics. After a transient in which the fast modes die, a linear system is described by its slowest surviving mode and nothing else — the shape stops changing and only the amplitude moves. The rate at which a room cools, a circuit settles, a population equilibrates or a numerical scheme converges is, after a while, a single exponential whose exponent is the smallest eigenvalue that has not already vanished.

The profile drawn here starts as a square wave and ends as a sine, and nothing about the square wave survives except its overall amplitude. Everything that distinguished it has been thrown away, and what is left is the shape the domain itself prefers.

What cannot be undone

The same exponent makes running the flow backwards impossible in a precise and unusually vivid way.

To recover the past, each coefficient would have to be multiplied by e+m2te^{+m^2 t}. The first harmonic would be scaled up by a modest factor; the forty-first by e1681e^{1681}. Any uncertainty in the measured present — and there is always some — is amplified by that same enormous factor, so an error of one part in 101210^{12} in a high harmonic becomes an error of astronomical size in the reconstructed past.

Backwards diffusion is the standard example of an ill-posed problem: a solution exists, it is unique, and it does not depend continuously on the data. The third condition is the one that matters in practice, since data always has noise in it, and a problem that magnifies noise without bound cannot be solved by any method whatever.

That is a satisfying answer to a question that looks physical. Why is heat flow irreversible? Not because of any statistical argument about molecules, at this level of description, but because the operation that would reverse it is unbounded — the information in the fine structure has been divided by numbers with hundreds of digits, and dividing back is not something that can be done to a measurement.

The practical version of the same fact is that deblurring a photograph is hard and unblurring a very blurred one is impossible. A blur is a convolution, convolution is multiplication in the frequency domain, and undoing it is division by a number that is nearly zero wherever the blur was effective.

The problem that produced the method

Fourier was not a pure mathematician and did not present this as mathematics.

He had been Napoleon’s scientific adviser in Egypt and then prefect of Isère, and he worked on heat because it was a problem of the moment — the propagation of heat in solids was of interest to metallurgy, to geology and to the question of the age of the Earth, which Fourier’s own analysis later informed. His 1807 memoir on the subject contained the decomposition, the eigenfunction argument above, and the solution; it was refused publication.

The objection was not to the physics. Lagrange, on the committee, did not accept that an arbitrary function could be written as a sum of sines, and Fourier’s argument for it was not sound. The memoir was awarded a prize in 1811 with an explicit note about its lack of rigour and was not printed until 1822, as part of Théorie analytique de la chaleur.

Two things came out of that dispute, and neither is the heat equation. The first is the century of analysis described at the end of the first rung of this ladder — convergence, continuity, the definition of a function and of an integral, all of it provoked by a question Fourier’s method made unavoidable. The second is the recognition that Fourier had found something more general than a solution to one equation: a method for any linear problem whose domain has symmetries, which is most of them.

It is a fair example of a result whose importance was not visible from either side of the argument. Fourier thought he had solved heat conduction; his critics thought he had made an error about series. He had done both, and the residue was a technique.

What it costs

The method needs the whole problem to be linear and the domain to be one the harmonics fit.

On a ring, the sines and cosines are exactly the functions that come back to themselves after a full turn, so the decomposition is available. On an interval with fixed ends, the right family is the sines that vanish at both ends; on a disc it is Bessel functions; on a sphere, spherical harmonics. In each case the family is the set of eigenfunctions of the same operation on that domain, and finding it is the whole difficulty. On a domain with an awkward shape there is no closed-form family at all, and the modes must be computed numerically — which is what a great deal of scientific computing is.

Linearity is the harder condition. The whole approach rests on the flow acting on each harmonic independently, and that fails the moment the equation has a term in u2u^2 or a conductivity that depends on the temperature. Then harmonics interact — energy moves between frequencies — and no amount of decomposition separates the problem. That interaction is exactly what makes turbulence hard, and it is the reason the Fourier method transformed the linear parts of physics and left the nonlinear parts as difficult as it found them.

The same equation, run as a blur

There is a second reading of the solution that makes the smoothing tangible and is worth having beside the harmonic one.

The spectrum of the triangle waveOne bar per harmonic, its height the size of that harmonic's coefficient. The 11 non-zero coefficients fall away like 1 over m to the power 2.00.0.200.410.610.81123456789131721harmoniccoefficientthe triangle wave, harmonic by harmonic: 11 of the first 21 are non-zero, and their sizes fall like 1 / m^2.00the exponent is fitted to the bars rather than taken from the formula, and it is what says how smooth the target is:a jump gives 1, a corner gives 2, and a smooth function falls faster than any power
Fig. 6 A profile with corners but no jumps, whose coefficients fall like 1/m21/m^2. Under the flow these are multiplied by the same em2te^{-m^2t} — the starting smoothness decides how much is there to lose, and the flow decides how fast it goes.

Multiplying every coefficient by em2te^{-m^2 t} is, in the original variable, a convolution: the profile at time tt is the starting profile smeared with a bell-shaped kernel whose width grows like t\sqrt{t}. Diffusion is a Gaussian blur, with a radius that spreads as the square root of the elapsed time.

That is exactly the growth rate a random walk’s spread has, and the coincidence is not one: the density of a random walker obeys this equation, so the blur is the distribution of where a walker has got to. The two subjects share their equation and therefore share every consequence — the t\sqrt{t}, the smoothing, the irreversibility, and the fact that a sharp initial condition becomes smooth at once.

The two readings answer different questions and it is worth having both. The frequency reading says which features survive; the blur reading says how far the influence of a point has spread. Neither is more fundamental, and moving between them is the transform of the previous rung doing what it exists for.

Where the picture needs a condition

The figures show a ring, which is chosen so the profile is periodic and the harmonics apply directly.

A bar with its ends held at a fixed temperature has different modes — sines that vanish at both ends — and a bar with insulated ends has cosines. The qualitative story is identical in all three cases, with the same em2te^{-m^2t}, and the equilibrium differs: a ring settles at its average temperature, a bar with cooled ends settles at zero, and an insulated bar settles at its average. Nothing in the pictures distinguishes these, and the choice of boundary condition is as much a part of the problem as the equation.

There is also a scale hidden in the figures. Writing the equation without a diffusion constant sets the length and time units together, so the numbers labelled tt are dimensionless. In a real material the decay rate of the mm-th mode is αm2\alpha m^2 with α\alpha the thermal diffusivity, which for copper is about 10410^{-4} square metres per second and for wood about a hundredth of that — a factor that changes every time in the caption and nothing about the shape of the story.

What the picture cannot show

The claim being made is about infinitely many harmonics, and the drawing sums forty-one of them. At t=0t = 0 that truncation is visible: the flat stretches ripple, and the jump overshoots by the nine percent that never goes away. So the first curve in each figure is not the initial condition; it is a partial sum standing in for it.

At every later time the truncation is invisible, and for a good reason — the terms left out are multiplied by em2te^{-m^2 t} with mm above forty-one, and at any drawn time those are below the width of the line. The figures are therefore least accurate exactly where they are most confident-looking, which is at t=0t = 0, and the assertion behind them checks the later frames rather than the first.

What no figure here shows is the amplification going the other way. The essay’s strongest claim is that the backwards flow multiplies errors by e+m2te^{+m^2t}, and a picture of that would be a picture of a curve leaving the canvas immediately in every direction. It is a claim about numbers of a size no drawing accommodates.

The ladder from here

Below: the series, the projection that supplies its coefficients and the transform that removes the need for periodicity. Above and outward: the wave equation, where the same decomposition gives modes that oscillate rather than decay — because there the second time derivative gives e±imte^{\pm imt} instead of em2te^{-m^2t}, and nothing is lost, which is why sound carries information and heat does not.

Sideways: the m2m^2 is the reason a random walk’s diffusion and this flow are the same object, since the walk’s density obeys precisely this equation; and the smoothing seen here is the mechanism behind every Gaussian blur, which is what the flow’s solution operator is.

The right coordinates make the problem disappear

The lasting point is the shape of the method rather than the answer.

The equation is hard because it couples every point of the profile to its neighbours. The decomposition finds a set of directions along which the coupling vanishes — where the operation acts by simple multiplication — and in those coordinates the problem is not solved so much as dissolved. There is nothing left to solve: each coefficient decays on its own, and the answer is written down.

That manoeuvre is one of the most valuable in the subject and it wears different names in different places. Here it is Fourier analysis; in linear algebra it is finding the directions a map leaves alone; in the study of a repeated process it is the eigenvalues that decide what survives. Each is the same instruction: do not attack the operator, find the things it does not mix up.

Fourier’s contemporaries objected to his method’s rigour and were right to; his sense of what the objects were for was better than theirs. The apparatus is not a clever way to write down a square wave. It is a change of coordinates in which an intractable physical process becomes a list of independent exponentials, and everything else this ladder contains is a consequence of it being available.

What links here

Computed from the collection, not written here: the essays that point at this one.

Named objects

A dashed tag is an object no other essay names yet.

DiffusionEigenfunctionExponential decayFourier analysisHeat equationIrreversibilitySpectrum