Analysis

A curve with a corner at every point

Continuity means a curve can be drawn without lifting the pen. Differentiability means it has a tangent. The first was assumed to nearly imply the second until 1872, when Weierstrass exhibited a curve that is continuous everywhere and has a tangent nowhere — and it is a sum of cosines.

Worth reading first: The slope of a single point · A square wave built entirely out of round ones.

A tangent is what a curve looks like from close up. Zoom in far enough on any ordinary curve and it becomes indistinguishable from a straight line; the slope of that line is the derivative, and the definition is a limit of chord slopes as the two points come together.

The question this essay is about is whether every curve that can be drawn without lifting the pen has such a line. For most of the nineteenth century the answer was assumed to be yes except at a scatter of exceptional points — corners, cusps, places where something obviously went wrong.

6 cosines, and a curve with no tangent anywhere. Partial sums of a sum of cosines whose amplitudes shrink geometrically and whose frequencies grow faster. Each term adds finer detail; the curve converges and its slopes do not.
Fig. 1 Partial sums of a sum of cosines whose heights are halved at each step and whose frequencies are tripled. Each new term adds detail at a finer scale and smooths nothing; the curve converges and the slopes do not.

The curve drawn there is continuous at every point of the line and has a tangent at none of them. It was published by Weierstrass in 1872, it caused a considerable amount of distress, and it is built out of the most familiar functions there are.

Two races, run at different speeds

The construction adds cosine waves. The kth term has height aᵏ and frequency bᵏ, with a below 1 so the heights shrink and b a whole number above 1 so the waves get faster.

Two things then happen at once, and the whole design is in the ratio between them.

The heights shrink geometrically, so the sum converges. The terms are bounded by aᵏ, which adds to a finite total, so the partial sums close in on a limit — and they do it uniformly, meaning the worst error anywhere is bounded by the tail of the geometric series. A uniform limit of continuous functions is continuous, so the result is a curve that can be drawn.

The slopes grow geometrically, so nothing converges there. Differentiating the kth term multiplies its height by its frequency, giving a slope of size about (ab)ᵏ. If ab is above 1 those grow without bound, and the sequence of derivatives of the partial sums diverges everywhere.

That divergence is not by itself a proof — a sequence of derivatives can misbehave while the limit function is perfectly differentiable. What Weierstrass showed, and what takes real work, is that the chord slopes of the limit itself do not settle at any point.

His own condition was stronger than ab > 1: he required b to be an odd whole number and ab to exceed 1 + 3π/2, which is about 5.71. Hardy showed in 1916 that ab ≥ 1 with ab > 1 is enough, and dropped the requirement on b being an integer. The first curve drawn here uses a = 0.5 and b = 3, so ab = 1.5 — comfortably inside Hardy’s condition and outside Weierstrass’s own. The picture is of a curve its inventor could not have proved anything about, and the two below it move each dial in turn to show that the product is what the condition is made of.

6 cosines, and a curve with no tangent anywhere. Partial sums of a sum of cosines whose amplitudes shrink geometrically and whose frequencies grow faster. Each term adds finer detail; the curve converges and its slopes do not.
Fig. 2 The same construction with the heights barely shrinking at all. Each term is 0.90.9 times as tall as the one before and three times as fast, so ab=2.7ab = 2.7 instead of 1.51.5 — and the curve still converges, because 0.90.9 is under one and a geometric series with ratio under one has a finite total whatever the frequencies are. Six terms are drawn, the last sampled at 1,2151{,}215 points for its 122122 oscillations. Convergence is settled by the heights alone; roughness is settled by the product.

Zooming in, and finding the same thing

The clearest evidence that no tangent exists is what happens under magnification.

The same corner, magnified 1,000 times. Four windows on the same curve, each a tenth the width of the last. The curve looks the same at every magnification, which is what having no tangent anywhere amounts to.
Fig. 3 The same curve in four windows, each one tenth the width of the last, and each drawn with as many terms as its own sampling resolves. A curve with a tangent would look straighter at every step. Here the width falls by a thousand and the height of the window falls by seventy-eight, so the picture is essentially unchanged.

An ordinary curve flattens under magnification: halve the window’s width and the curve’s departure from a straight line falls by a factor of four, so a window kept in proportion to the curve’s own variation looks flatter and flatter. That flattening is differentiability, stated in a way that can be looked at.

Here the flattening does not happen. Magnifying by ten reduces the vertical range of the window by a factor of about four, not ten, so each panel has the same amount of wiggling in it as the last. The figure computes those ranges and asserts that they fall by less than a factor of six per tenfold zoom, which is the picture’s version of the theorem.

The exponent hiding in that factor is worth naming, since it is the sharpest thing that can be said about the curve. The height of a window of width w is proportional to w raised to the power log(1/a)/log b, which for these values is log 2/log 3, about 0.63. A differentiable curve has exponent 1; a curve with exponent below 1 has chord slopes of size w⁰·⁶³/w = w⁻⁰·³⁷, which grows without limit as w shrinks. The whole theorem is that one exponent being less than one.

What is happening is that the terms with frequency below the window’s width are nearly straight lines across it and contribute nothing to the shape, while the terms with higher frequency are a smaller copy of the whole picture. Zooming in trades one set of terms for another and leaves the shape statistically the same — which is self-similarity, and it is why this curve is one of the first objects the word fractal was invented for.

Measuring the chords

The definition of a derivative is a limit of chord slopes, so the way to test it is to take chords.

Chords that settle, and chords that do not. The slope of the chord from a fixed point, at spacings that halve, for an ordinary curve and for one built from cosines. The first sequence converges; the second grows without bound.
Fig. 4 The same measurement on two curves: the slope of the chord from a fixed point to one a distance h away, for h halving nine times. On the left the slopes close in on a number. On the right they swing further and further and are still growing when the measurement stops.

For an ordinary curve the chord slopes settle: the figure’s left panel closes on the tangent’s slope, and the differences shrink by half with each halving of h, which is what a well-behaved limit looks like.

For the rough curve, the chord slopes do not merely fail to converge — they grow. Halving the spacing brings a finer term into play whose contribution to the chord slope is ab times the previous one, so the sequence of measurements diverges geometrically. That is a stronger failure than a corner, where the two one-sided limits exist and disagree; here neither one-sided limit exists at all.

The difference is worth stating carefully because it is where the nineteenth-century intuition failed. A corner is a failure at one point that a picture makes obvious. This is a failure at every point that no picture makes obvious, because the curve looks perfectly reasonable at any magnification a page can carry.

The same measurement, on the other dial

The chord measurement above was made on one member of the family. It is worth seeing the family move, since the whole claim is about a product rather than about one curve.

4 cosines, and a curve with no tangent anywhere. Partial sums of a sum of cosines whose amplitudes shrink geometrically and whose frequencies grow faster. Each term adds finer detail; the curve converges and its slopes do not.
Fig. 5 The other dial moved instead: the heights halve as before and the frequencies now multiply by seven, so ab=3.5ab = 3.5. Four terms are drawn, the last sampled at 1,6001{,}600 points for its 172172 oscillations, and four are already enough to make the curve unreadable as a shape where six were needed at b=3b = 3. The roughness arrives faster and it is the same roughness — neither knob decides on its own, their product does.

The left-hand panel of the chord figure is the rung this essay stands on: chords through a fixed point, each strictly closer than the last, with slopes closing in on a number — and the generator requires each chord to be strictly closer than its predecessor, because a sequence of secants that did not close would be showing nothing.

The point of setting the two panels side by side is that they are the same measurement. Nothing in the procedure changes: pick a point, pick a second one nearby, divide the rise by the run, and bring the second point in. What differs is entirely in the function being measured, and the measurement’s failure to settle is the only evidence available that something is wrong. There is no separate test for differentiability; there is the limit, and either it exists or it does not.

A point with two slopes. Secants to |x| at zero, taken from each side. Every one from the right has slope 1 and every one from the left has slope −1, at every distance, so the quotients never settle.
Fig. 6 The failure everybody meets first: a corner, where the chords from the left settle on one number and the chords from the right settle on another. Two limits exist and disagree, which is a far gentler failure than having no limit at all.

What it broke

The example is famous less for what it says than for what it stopped people from saying.

Before it, arguments in analysis routinely assumed that a continuous function was differentiable except at isolated points, and proofs were built on that assumption without comment. Ampère had published an attempted proof that continuity implies differentiability except on a small set; textbooks had asserted it. The counterexample made those arguments unusable overnight and forced the subject to define its terms.

The reaction at the time was not admiration. Hermite wrote of turning away in fear and horror from the lamentable plague of functions with no derivatives; Poincaré called such examples a monstrosity invented to embarrass the mathematics of the previous generation. That reaction is worth recording because it turned out to be exactly backwards: the exceptional case is the differentiable curve, not this one. In any reasonable sense of most, most continuous functions have a tangent nowhere.

Making that precise took another fifty years — Banach and Mazurkiewicz, in 1931, showed that the functions with a derivative somewhere form a vanishingly small subset of all continuous functions. The curve in the figures is not a monster. It is typical, and the smooth curves everybody had been thinking about are the specimens.

What has to be given up to be this rough

The example is usually presented as showing that continuity buys nothing, and that overstates it. Continuity alone buys nothing; continuity plus one more modest condition buys a derivative almost everywhere, and knowing which conditions those are says exactly what the curve had to sacrifice.

A monotone function is differentiable almost everywhere. That is Lebesgue’s theorem, and it holds with no smoothness assumed: a function that never decreases has a tangent at every point outside a set of measure zero, however many jumps and corners it has. So a nowhere-differentiable curve must fail to be monotone on every interval, however short — there is no stretch, anywhere, on which it only rises.

The condition can be weakened further and the conclusion survives. A function of bounded variation — one expressible as a difference of two monotone functions, equivalently one whose total up-and-down movement over an interval is finite — is also differentiable almost everywhere. So the curve here must have unbounded variation on every interval.

That has a consequence worth stating in the essay’s own terms: the graph has infinite length between any two points. Not merely a large length, and not merely at fine scales — take any two points on the curve, however close, and the arc between them is not rectifiable. The wiggling adds up, and it adds up to infinity in every window.

Which is what the magnification figure was showing without saying so. Each tenfold zoom brings in terms whose total up-and-down movement is a constant fraction of the previous panel’s, so the movement accumulates without bound rather than converging. A curve of finite length would have that quantity settle, and settling is exactly what does not happen.

So the price of having no tangent anywhere is not exotic. It is oscillation, at every scale, in every window, forever — and the two positive theorems above say that oscillation is the only way to pay it. Anything that fails to oscillate on some interval has a derivative on most of that interval, whatever else it does.

This also explains the two neighbouring failures. A corner oscillates nowhere and is differentiable everywhere but one point. A staircase function is monotone and therefore differentiable almost everywhere despite having infinitely many jumps. Neither can be made nowhere-differentiable by piling on more of the same defect, because neither defect is oscillation.

And it sharpens the historical remark above. When Poincaré called the example a monstrosity, the reasonable reply was already available: it is the only shape a continuous function can take if it is to have no tangent anywhere, and its features are forced rather than contrived.

Built from the same parts as a square wave

The construction uses cosines with multiplied frequencies, which is exactly the vocabulary of a Fourier series, and the comparison with a familiar one is instructive.

A square wave from 7 sine waves. The sum of the first 7 harmonics, compared with the square wave it approaches.
Fig. 7 A square wave assembled from seven sine waves. The amplitudes fall like one over the frequency, which is slow enough to produce a jump in the limit and fast enough to keep the sum bounded — a different point on the same dial.

A square wave built from round ones has amplitudes falling like 1/k with frequencies k: the product of amplitude and frequency is constant, so the slopes neither grow nor shrink, and the limit has jumps but is smooth between them.

Weierstrass’s curve pushes the same dial one notch further. The amplitudes fall geometrically, which is much faster, so the sum is not merely bounded but continuous everywhere; the frequencies grow geometrically and faster still, so the product grows. Fast decay in one column and faster growth in the other, and the result is continuous with no smooth stretch anywhere.

Reading a Fourier series’ coefficients as a statement about smoothness is a standard technique, and this is its extreme case. Coefficients decaying like 1/k² give a differentiable limit; like 1/k, a limit with jumps; geometrically with a fast enough frequency growth, a limit that is continuous and nowhere differentiable. The smoothness of the answer is legible in how quickly the coefficients die, and nothing else about them matters.

Where it turns up outside mathematics

A path that is continuous, nowhere differentiable and self-similar is exactly what a physical process looks like when it has structure at every scale.

The trace of a particle in a fluid, watched under a microscope, has that character: at any magnification the path has as much detail as at any other, and asking for its velocity at an instant is asking for a limit that does not exist. Brownian motion is the mathematical model, and its paths are continuous and nowhere differentiable with probability one — the same pair of properties, arrived at by an entirely different route.

The relationship to a random walk is direct. A walk’s position after n steps has typical size √n, so over a time interval of length h the change is about √h, and the chord slope is √h/h = 1/√h, which grows without bound as h shrinks. That is the same divergence the figure measures, produced by scaling rather than by a series.

The practical consequence is worth naming, because it changes what can be asked of a model. A quantity whose path has this character has no instantaneous rate of change, so any equation written with one is meaningless as it stands — which is why stochastic calculus had to be built from scratch rather than borrowed, and why the rules in it are not the rules of ordinary calculus. The chain rule acquires an extra term, and the extra term is exactly the roughness that the h scaling measures.

A second consequence runs the other way. A measurement of a rough path taken at finite resolution always produces a finite slope, and that slope depends entirely on the resolution — halve the sampling interval and the measured slope changes by a predictable factor. Anybody measuring the roughness of a coastline, a price series or a surface is measuring an exponent rather than a value, and the exponent is the only thing that survives a change of instrument.

What it costs

Everything drawn here is a partial sum, since an infinite series cannot be evaluated — and how many terms to draw is not a free choice, which is the interesting part of the cost.

The kth term oscillates bᵏ times per two units, so a picture of a partial sum has a finest wiggle whose width is fixed by the term count and the span. Sampling below about eight points per wiggle does not draw the function; it draws an interference pattern between the function and the sampling grid, which looks convincingly rough and is not the curve. The first version of the hero here did exactly that — six terms at b = 7 is sixteen thousand oscillations, sampled four thousand times — and it produced a figure that was both wrong and 248 KB of coordinates.

Each curve is now sampled at ten points per oscillation of its own finest term, which is 1,215 samples for the six-term curve’s 122 oscillations, and the generator refuses to draw any curve at fewer than eight. The magnification figure applies the same rule per panel: each window is drawn with as many terms as its sampling resolves — five, seven, nine and twelve of them — and the terms left out are asserted to contribute less than a twentieth of what the window shows.

So the pictures are honest at the scales drawn and cannot be extended indefinitely. The theorem is about all scales at once, and no drawing reaches past the resolution of its own grid.

What the picture cannot show

The claim is that no point of the line has a tangent, and every figure examines finitely many points. The chord measurement is taken at one place; the zoom is centred on one place. Nothing about the drawings rules out a single well-behaved point somewhere else, and the proof that there is none is a piece of analysis with no picture.

Nor can the figures distinguish a curve with no tangent from a curve whose tangent is very hard to see. A curve made from a finite number of terms does have a tangent everywhere; it is a finite sum of cosines and perfectly smooth. Every picture on this page is of such a curve — the six-term partial sum has a derivative at every point, and that derivative is merely large. What is drawn is a smooth approximation to a non-smooth limit, and the non-smoothness is entirely in the passage to the limit that no drawing performs.

And the self-similarity is only statistical. The curve does not repeat exactly under magnification, so a reader looking for an exact copy in the deepest panel of the zoom will not find one.

The ladder from here

Rungs above: Hardy’s version of the theorem, with the condition on b removed. The Takagi function, which is nowhere differentiable, built from triangle waves rather than cosines, and easier to analyse. The Hausdorff dimension of the graph, which is 2 + log a/log b — a number between 1 and 2, so the curve is genuinely rougher than a line and genuinely thinner than a region. Brownian motion, where the same pair of properties holds with probability one. The Banach–Mazurkiewicz theorem, which says that nowhere-differentiable is the typical case. Functions differentiable everywhere with an unbounded derivative, and functions differentiable exactly once. And the space-filling curves, which are continuous and worse.

The shape of the idea

Two limits are involved and the essay is about which of them commutes with which.

Summing the series is one limit. Taking the derivative is another. For a finite sum they interchange freely — the derivative of a sum is the sum of the derivatives, and every partial sum here is differentiable. What Weierstrass’s example shows is that the interchange fails in the limit, and it fails so badly that one side does not exist.

That is the standing shape of a counterexample in analysis: not a strange object built for its own sake, but a demonstration that two operations everybody performs without thinking cannot be performed in either order. The rearranged series makes the same point about addition and reordering; the staircase that is not the diagonal makes it about limits and length.

Each of them is remembered as a monster and each is really a piece of bookkeeping about which limits may be swapped. Since analysis is almost entirely the study of interchanging limits, the monsters are the subject’s actual content and the well-behaved cases are the corner it happens to have started in.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

ContinuityConvergenceCounterexampleDifference quotientDifferentiabilityFourier seriesFractal dimensionSelf-similarityUniform convergenceWeierstrass function