The ripples that make a series run away
Worth reading first: A square wave built entirely out of round ones · A plucked string keeps its corners.
A square wave built out of round ones ends with a list of facts about convergence that the pictures could not show. The Fourier series of a continuous function can diverge at a point — du Bois-Reymond, 1873. Averaging the partial sums instead of taking them as they stand repairs it — Fejér, 1900. Both are stated there, as facts. This essay is about why they are true, and the reason turns out to be one picture and one number.
The picture is a curve called the Dirichlet kernel, which is what the -th partial sum does to any function whatever. The number is the area under its absolute value. That area grows without limit, but only as fast as the logarithm of — so slowly that the divergent series below takes terms to reach — and it is the whole of the obstruction. Continuity was never going to be enough, because no condition on the size of a function can beat a multiplier that grows. What averaging does is make the multiplier stop growing, and it does that by making the kernel positive.
A partial sum is an average against a fixed curve
The -th partial sum keeps the harmonics from to . Written out with the coefficients as projections and the order of sum and integral exchanged, it is
The closed form for is a geometric series summed round the circle. What the formula says is that does the same thing to every function: it slides the fixed curve along and averages against it. Everything about how partial sums behave, for every function at once, is a property of that one curve.
The spike in the middle does what a partial sum is supposed to do. It is tall and narrow, its signed area is exactly one, and averaging a function against it picks out the function’s value near . If the kernel were only the spike, partial sums would converge for every continuous function.
The ripples are the trouble. Away from the centre, oscillates between roughly , and that envelope does not depend on . As grows the ripples crowd together but do not get lower. They alternate in sign, so their signed contributions nearly cancel, and for a smooth function that cancellation is what saves the day. But a function is free to have its own ups and downs, and a function whose ups meet the kernel’s ups and whose downs meet its downs gets no cancellation at all.
The most a partial sum can amplify
Take a function that never exceeds one in size and ask how large can be. The average against is largest when equals wherever is positive and wherever it is negative, and then the average is the area under :
These are the Lebesgue constants, and is exactly the factor by which the -th partial sum can magnify a function — the worst case, over every function of size at most one, of how large the partial sum can come out. The sign function is not continuous, but it can be made continuous by replacing each jump with a steep ramp, and the loss is as small as the ramps are narrow.
The figure’s function looks like a square wave that has been told where to switch, and that is exactly what it is. Its ten-term partial sum at the origin is — the function it came from never exceeded one. Nothing about the function is extreme. It is continuous, bounded, and changes sign twenty times. What is extreme is its alignment with the kernel, and any fixed has such a function.
The same function is harmless to other partial sums. Its sign pattern was cut to match the ripples of , which sit at spacing ; the ripples of sit ten times closer, and against them the function’s slow switching averages out, so its partial sum at the origin falls from at ten terms to at thirty and at a hundred, settling on the function’s own value of one. A bad function for one is an ordinary function for the others. That is why the Lebesgue constants alone do not yet give a divergent series: a single function has to be bad for infinitely many at once, which means it has to carry sign patterns aligned with kernels at infinitely many scales, each small enough not to spoil continuity and each large enough to push its own partial sum up. Stacking them without letting them interfere is the whole difficulty of the construction below, and it is solved by putting the scales so far apart that each one has finished before the next begins.
The number is computable in closed form, which is how the figures check it: , a formula of Fejér’s, compared at small with a direct quadrature of that shares none of its arithmetic. is , and the constants rise from there.
Growing like a logarithm
How fast they rise is the heart of the matter, and it can be read off the picture. Away from the spike, is divided by . The numerator oscillates so fast that it contributes its average value, . So the ripples contribute roughly
and the spike, whose area is bounded, contributes a constant. The Lebesgue constants grow like — without limit, but only as fast as the harmonic series, and for exactly the same reason: an integral of .
On a logarithmic axis the constants are a straight line of slope per unit of , which is per factor of ten. At ten thousand terms the amplification is ; at a million it would be ; at about . Nobody will ever see the amplification become large on a computer, and it is unbounded all the same.
The same slow growth appears in polynomial interpolation at well-chosen nodes, where the best possible interpolation still has a Lebesgue constant growing like a logarithm. That is not a coincidence. A theorem of Lozinski and Kharshiladze says that every linear way of producing a trigonometric polynomial of degree from a continuous function, if it leaves such polynomials alone, amplifies some function by at least . The Fourier partial sum is the best of all of them, and even the best one cannot keep its amplification bounded.
From unbounded amplification to a divergent series
A different bad function for each is not yet a single function whose series diverges. The step from one to the other is a principle about linear operations that is worth stating in full, because it is one of the few places in analysis where an existence proof arrives from nothing but a size count.
If each of a sequence of linear operations on continuous functions is bounded, but their bounds are not, then some single continuous function is made unbounded by the sequence. That is the uniform boundedness principle, proved in general by Banach and Steinhaus in 1927. Applied to the partial sums evaluated at , whose bounds are the , it produces a continuous with : a continuous function whose Fourier series diverges at .
The proof is not a construction. It uses a completeness argument due to Baire to show that the functions whose partial sums at stay bounded are, in a precise sense, a negligible set — so negligible that most continuous functions, in that sense, have Fourier series that diverge at . The functions a reader draws, and the functions whose series anybody computes, are the exceptions.
That is a stronger conclusion than du Bois-Reymond’s, and it came fifty years later. His 1873 counterexample was a construction, and so was a much cleaner one that Fejér gave in 1910. Fejér’s is worth drawing, because every piece of it can be computed.
Fejér’s function, one block at a time
The building block is a trigonometric polynomial
chosen to be a small function with a large partial sum. It is small because the sum never exceeds in size, for any and any . That constant is , the Wilbraham–Gibbs integral, and it is the same number that fixes the square wave’s 9% overshoot: the overshoot at a jump is of its height. So every lies between and .
Its partial sums are large because multiplying out turns into cosines: at frequency and at frequency . Its value at is zero, since the two halves cancel. But a partial sum that stops at frequency has collected every and none of the , so at it equals . The block holds a hidden harmonic sum, and the truncation exposes it.
Now add blocks at wildly separated frequencies, shrunk by so that the total stays small: with . The shrinking makes the sum converge uniformly, so is continuous, and . The separation means that when the partial sums reach the -th block, the earlier blocks are complete and contribute nothing at , and the later ones have not started. So the partial sum at climbs to about as it passes , and falls back to nothing after .
The figure uses four blocks and so draws a function whose series converges — every finite sum of blocks is a trigonometric polynomial. With the blocks continued for ever, the spikes grow like , the partial sums at are unbounded, and the series diverges there, while the function it came from is continuous and zero at that very point. The fourth spike sits at terms — about eighteen billion billion — and reaches only . The divergence is real and it is invisible to computation, which is why it was found by argument rather than by observation.
Averaging makes the kernel positive
Fejér’s more famous result came first and is the repair. Instead of the partial sums, take their running average,
which is the averaging Cesàro used to give Grandi’s series a value. Averaging partial sums means averaging their kernels, and the average of through has a closed form that is a square:
A square is never negative. So has no negative ripples at all, its absolute area equals its signed area, and that is one, for every . The amplification that grew like a logarithm is now stuck at exactly one. More than that, an average against a positive kernel with area one is a genuine weighted average of the function’s values, so always lies between the smallest and largest values of — which is why the averaged sums of the square wave rise to the jump from underneath without ever overshooting. And as grows the positive bump concentrates at the centre, so the weights concentrate near , and continuity of does the rest: the averages converge to , uniformly, for every continuous .
Fejér proved it in 1900, while still a student in Budapest. A corollary is Weierstrass’s theorem that every continuous periodic function is a uniform limit of trigonometric polynomials, since each is one. The coefficients were never the problem. The partial sums use them in a way whose kernel has negative ripples, and the averaged sums use the same coefficients in a way whose kernel does not.
What the curves cannot show
They cannot show the divergent function. The fourth block of Fejér’s function oscillates at frequency , and no drawing of the function could resolve it; what is drawn is its partial sums at a single point, computed from harmonic numbers. The function itself is known only through its formula and the bound that makes it continuous.
They cannot show where divergence can happen. The picture shows divergence at one point. How large the set of such points can be for a continuous function is a hard question with a two-sided answer: Carleson’s theorem of 1966 says the series of any continuous function — indeed of any function with finite energy — converges everywhere except on a set of length zero, and Kahane and Katznelson showed the same year that every set of length zero is the divergence set of some continuous function. The logarithmic growth of drawn above is far too weak to reach Carleson’s theorem, whose proof needs a much finer analysis of how the kernel’s ripples interact with the function’s own oscillations.
And they cannot show the uniform boundedness argument. It is a proof by category: it shows that the functions with bounded partial sums at are too few to be everything, without exhibiting one that is left out. No figure represents “most continuous functions”, and the figure that does exist, Fejér’s, represents a single exception built by hand.
A race between two numbers
The Lebesgue constants do more than produce counterexamples. They also say exactly which functions are safe, through an inequality that fits on one line and turns the whole question into a race.
Let be the smallest possible uniform error between and any trigonometric polynomial of degree — the best that any method whatsoever could do with harmonics. Pick the polynomial that achieves it. The partial sum leaves alone, since is already its own series cut off at , so
and the right-hand side is a function of size at most plus that same function amplified by at most . Hence
This is Lebesgue’s inequality, and it reframes everything above. The partial sums converge whenever the best possible approximation error shrinks faster than the amplification grows — whenever . The obstruction is not that the partial sums approximate badly; it is that they can be worse than the best approximation by a factor of , and only a function that is barely approximable at all can lose that race.
How approximable a function is depends on how smooth it is, and that dependence was worked out by Dunham Jackson in 1911: if changes by at most over any distance , then is at most a constant times . Put the two together and the partial sums converge uniformly as soon as . That is the Dini–Lipschitz test, and it covers every function whose graph is no steeper than any power of the distance — every Hölder continuous function, and so every function anybody writes down with a formula. The counterexamples have to be continuous with a modulus that decays more slowly than , which is an extraordinarily weak form of continuity, and Fejér’s blocks are built with exactly that slowness: their frequencies run away so fast that each one only just fits under the bound.
So the logarithm cuts both ways. It is small enough that a trace of smoothness defeats it, and large enough that continuity alone cannot.
Where this leads: how fast, not whether
For functions with any smoothness at all, then, convergence is settled, and the question that matters in practice is not whether a series converges but how fast, and that is decided by how quickly the coefficients decay, which is decided by how smooth the function is: a jump costs , a corner , as the plucked string showed, and every further derivative buys another power.
The larger lesson travels. Whenever a limit is taken by a sequence of linear operations — interpolations, quadratures, discretisations — the right first question is whether their sizes stay bounded. If they do not, the uniform boundedness principle guarantees a continuous input that defeats the method, whether or not anyone has found it, and the cure, when there is one, is to replace the operations by averages with positive weights.
One number decides it
The -th Fourier partial sum is an average against the Dirichlet kernel, and the area of that kernel’s absolute value, the Lebesgue constant , grows like without bound. Any fixed therefore has a continuous function of size one that its partial sum amplifies by , and the uniform boundedness principle turns that sequence of bad functions into one continuous function whose series diverges at a point — as Fejér’s explicit function, with blocks at , does in plain view.
Averaging the partial sums averages the kernels, and the average is a square: positive, of area one, with no ripples below zero. The amplification stops growing, the averaged sums stay within the function’s range, and they converge for every continuous function. The coefficients were fine all along; the way the partial sums weigh them was not.
When a method fails for some input nobody can draw, measure the method’s size first — an unbounded size is a proof that the input exists.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A curve with a corner at every point — both name continuity, convergence, counterexample
- The staircase that is not the diagonal — both name continuity, convergence, counterexample
- Uniform, except on a small set — both name continuity, convergence, counterexample
- Which curves have a length at all — both name continuity, convergence, counterexample
- A curve that has area — both name continuity, counterexample
- A limit can jump at every fraction — both name continuity, counterexample
Named objects
A dashed tag is an object no other essay names yet.
Cesaro summationContinuityConvergenceCounterexampleDirichlet kernelFourier analysisGrowth rateLebesgue constantUniform boundedness