Analysis

A curve with a corner at every point

Continuity means a curve can be drawn without lifting the pen. Differentiability means it has a tangent. The first was assumed to nearly imply the second until 1872, when Weierstrass exhibited a curve that is continuous everywhere and has a tangent nowhere — and it is a sum of cosines.

Worth reading first: The slope of a single point · A square wave built entirely out of round ones.

A tangent is what a curve looks like from close up. Zoom in far enough on any ordinary curve and it becomes indistinguishable from a straight line; the slope of that line is the derivative, and the definition is a limit of chord slopes as the two points come together.

The question this essay is about is whether every curve that can be drawn without lifting the pen has such a line. For most of the nineteenth century the answer was assumed to be yes except at a scatter of exceptional points — corners, cusps, places where something obviously went wrong.

6 cosines, and a curve with no tangent anywherePartial sums of a sum of cosines whose amplitudes shrink geometrically and whose frequencies grow faster. Each term adds finer detail; the curve converges and its slopes do not.-0.4-0.20.20.4-0.50.511.52xthe sum1 term2 terms3 terms6 termseach term has 0.5 times the height of the one before and 3 times the frequency, so ab = 1.5 — above one, and thatis the condition6 terms are drawn, the last at 1,215 samples for its 122 oscillations; the sum converges because the heights area geometric series, and the slopes do not because they are multiplied by 1.5 at each step
Fig. 1 Partial sums of a sum of cosines whose heights are halved at each step and whose frequencies are tripled. Each new term adds detail at a finer scale and smooths nothing; the curve converges and the slopes do not.

The curve drawn there is continuous at every point of the line and has a tangent at none of them. It was published by Weierstrass in 1872, it caused a considerable amount of distress, and it is built out of the most familiar functions there are.

Two races, run at different speeds

The construction adds cosine waves. The kth term has height aᵏ and frequency bᵏ, with a below 1 so the heights shrink and b a whole number above 1 so the waves get faster.

Two things then happen at once, and the whole design is in the ratio between them.

The heights shrink geometrically, so the sum converges. The terms are bounded by aᵏ, which adds to a finite total, so the partial sums close in on a limit — and they do it uniformly, meaning the worst error anywhere is bounded by the tail of the geometric series. A uniform limit of continuous functions is continuous, so the result is a curve that can be drawn.

The slopes grow geometrically, so nothing converges there. Differentiating the kth term multiplies its height by its frequency, giving a slope of size about (ab)ᵏ. If ab is above 1 those grow without bound, and the sequence of derivatives of the partial sums diverges everywhere.

That divergence is not by itself a proof — a sequence of derivatives can misbehave while the limit function is perfectly differentiable. What Weierstrass showed, and what takes real work, is that the chord slopes of the limit itself do not settle at any point.

His own condition was stronger than ab > 1: he required b to be an odd whole number and ab to exceed 1 + 3π/2, which is about 5.71. Hardy showed in 1916 that ab ≥ 1 with ab > 1 is enough, and dropped the requirement on b being an integer. The curves drawn here use a = 0.5 and b = 3, so ab = 1.5 — comfortably inside Hardy’s condition and outside Weierstrass’s own. The pictures are of a curve their inventor could not have proved anything about.

Powers of 0.5, added upA bar for each term of a geometric series with the running total drawn over it, approaching but never reaching the horizontal line at 2.1234567891000.511.52termsvalue1 / (1 − 0.5) = 2the terms are 1, 0.5, 0.25, 0.125, … and the running total closes on 2what is missing after 10 terms is 1.95e-3, which is the next term over one minus the ratio
Fig. 2 Why the heights are harmless: a geometric series with ratio below one has a finite total, whatever the frequencies of the things being added. Convergence of the curve is settled by this picture and has nothing to do with the shape of the terms.

Zooming in, and finding the same thing

The clearest evidence that no tangent exists is what happens under magnification.

The same corner, magnified 1,000 timesFour windows on the same curve, each a tenth the width of the last. The curve looks the same at every magnification, which is what having no tangent anywhere amounts to.1 wideand 2.895 tall0.1 wideand 0.722 tall0.01 wideand 0.198 tall0.001 wideand 0.037 tallthe same curve around x = 0.34, in windows 1, 0.1, 0.01, 0.001 wide, each drawn with as many terms as its ownsampling resolves — 5, 7, 9, 12 of thema curve with a tangent would look straighter at every step — here the height of the window falls by 77.9 while the widthfalls by 1,000
Fig. 3 The same curve in four windows, each one tenth the width of the last, and each drawn with as many terms as its own sampling resolves. A curve with a tangent would look straighter at every step. Here the width falls by a thousand and the height of the window falls by seventy-eight, so the picture is essentially unchanged.

An ordinary curve flattens under magnification: halve the window’s width and the curve’s departure from a straight line falls by a factor of four, so a window kept in proportion to the curve’s own variation looks flatter and flatter. That flattening is differentiability, stated in a way that can be looked at.

Here the flattening does not happen. Magnifying by ten reduces the vertical range of the window by a factor of about four, not ten, so each panel has the same amount of wiggling in it as the last. The figure computes those ranges and asserts that they fall by less than a factor of six per tenfold zoom, which is the picture’s version of the theorem.

The exponent hiding in that factor is worth naming, since it is the sharpest thing that can be said about the curve. The height of a window of width w is proportional to w raised to the power log(1/a)/log b, which for these values is log 2/log 3, about 0.63. A differentiable curve has exponent 1; a curve with exponent below 1 has chord slopes of size w⁰·⁶³/w = w⁻⁰·³⁷, which grows without limit as w shrinks. The whole theorem is that one exponent being less than one.

What is happening is that the terms with frequency below the window’s width are nearly straight lines across it and contribute nothing to the shape, while the terms with higher frequency are a smaller copy of the whole picture. Zooming in trades one set of terms for another and leaves the shape statistically the same — which is self-similarity, and it is why this curve is one of the first objects the word fractal was invented for.

Measuring the chords

The definition of a derivative is a limit of chord slopes, so the way to test it is to take chords.

Chords that settle, and chords that do notThe slope of the chord from a fixed point, at spacings that halve, for an ordinary curve and for one built from cosines. The first sequence converges; the second grows without bound.h/1h/4h/16h/64h/256-2-112a curve with a tangenth/1h/4h/16h/64h/256-20-101020a curve with nonethe same measurement on both: the chord from x = 0.34 to x = 0.34 + h, for h halving 9 times down to 9.8e-4on the left the slopes close on 1.842; on the right they reach 18.9 and are still growing, which is what itmeans for a tangent not to exist
Fig. 4 The same measurement on two curves: the slope of the chord from a fixed point to one a distance h away, for h halving nine times. On the left the slopes close in on a number. On the right they swing further and further and are still growing when the measurement stops.

For an ordinary curve the chord slopes settle: the figure’s left panel closes on the tangent’s slope, and the differences shrink by half with each halving of h, which is what a well-behaved limit looks like.

For the rough curve, the chord slopes do not merely fail to converge — they grow. Halving the spacing brings a finer term into play whose contribution to the chord slope is ab times the previous one, so the sequence of measurements diverges geometrically. That is a stronger failure than a corner, where the two one-sided limits exist and disagree; here neither one-sided limit exists at all.

The difference is worth stating carefully because it is where the nineteenth-century intuition failed. A corner is a failure at one point that a picture makes obvious. This is a failure at every point that no picture makes obvious, because the curve looks perfectly reasonable at any magnification a page can carry.

The ordinary case, for comparison

It is worth putting the well-behaved picture next to the rough one, since the whole claim is a comparison.

Secants closing on the tangent to x²Secant lines through x = 1 and a second point 1.5, 1, 0.6, 0.3, 0.1 away, with the slope of each. They approach 2, the derivative there.0.511.522.5123456xh = 1.5 slope 3.5000h = 1 slope 3.0000h = 0.6 slope 2.6000h = 0.3 slope 2.3000h = 0.1 slope 2.1000
Fig. 5 Secants closing on a tangent in the ordinary way: five chords through a fixed point, each closer than the last, with their slopes approaching the derivative. Every step of that sequence is a measurement, and the sequence settles.

The figure above is the rung this essay stands on. Five chords, each strictly closer than the last, with slopes closing in on a number — and the generator requires each chord to be strictly closer than its predecessor, because a sequence of secants that did not close would be showing nothing.

The point of setting the two side by side is that they are the same measurement. Nothing in the procedure changes: pick a point, pick a second one nearby, divide the rise by the run, and bring the second point in. What differs is entirely in the function being measured, and the measurement’s failure to settle is the only evidence available that something is wrong. There is no separate test for differentiability; there is the limit, and either it exists or it does not.

A point with two slopesSecants to |x| at zero, taken from each side. Every one from the right has slope 1 and every one from the left has slope −1, at every distance, so the quotients never settle.-2-1.5-1-0.50.511.52-0.50.511.52xslope 1 from the rightslope −1 from the left
Fig. 6 The failure everybody meets first: a corner, where the chords from the left settle on one number and the chords from the right settle on another. Two limits exist and disagree, which is a far gentler failure than having no limit at all.

What it broke

The example is famous less for what it says than for what it stopped people from saying.

Before it, arguments in analysis routinely assumed that a continuous function was differentiable except at isolated points, and proofs were built on that assumption without comment. Ampère had published an attempted proof that continuity implies differentiability except on a small set; textbooks had asserted it. The counterexample made those arguments unusable overnight and forced the subject to define its terms.

The reaction at the time was not admiration. Hermite wrote of turning away in fear and horror from the lamentable plague of functions with no derivatives; Poincaré called such examples a monstrosity invented to embarrass the mathematics of the previous generation. That reaction is worth recording because it turned out to be exactly backwards: the exceptional case is the differentiable curve, not this one. In any reasonable sense of most, most continuous functions have a tangent nowhere.

Making that precise took another fifty years — Banach and Mazurkiewicz, in 1931, showed that the functions with a derivative somewhere form a vanishingly small subset of all continuous functions. The curve in the figures is not a monster. It is typical, and the smooth curves everybody had been thinking about are the specimens.

Built from the same parts as a square wave

The construction uses cosines with multiplied frequencies, which is exactly the vocabulary of a Fourier series, and the comparison with a familiar one is instructive.

A square wave from 7 sine wavesThe sum of the first 7 harmonics, compared with the square wave it approaches.−π−π/2π/2π-117 terms
Fig. 7 A square wave assembled from seven sine waves. The amplitudes fall like one over the frequency, which is slow enough to produce a jump in the limit and fast enough to keep the sum bounded — a different point on the same dial.

A square wave built from round ones has amplitudes falling like 1/k with frequencies k: the product of amplitude and frequency is constant, so the slopes neither grow nor shrink, and the limit has jumps but is smooth between them.

Weierstrass’s curve pushes the same dial one notch further. The amplitudes fall geometrically, which is much faster, so the sum is not merely bounded but continuous everywhere; the frequencies grow geometrically and faster still, so the product grows. Fast decay in one column and faster growth in the other, and the result is continuous with no smooth stretch anywhere.

Reading a Fourier series’ coefficients as a statement about smoothness is a standard technique, and this is its extreme case. Coefficients decaying like 1/k² give a differentiable limit; like 1/k, a limit with jumps; geometrically with a fast enough frequency growth, a limit that is continuous and nowhere differentiable. The smoothness of the answer is legible in how quickly the coefficients die, and nothing else about them matters.

Where it turns up outside mathematics

A path that is continuous, nowhere differentiable and self-similar is exactly what a physical process looks like when it has structure at every scale.

The trace of a particle in a fluid, watched under a microscope, has that character: at any magnification the path has as much detail as at any other, and asking for its velocity at an instant is asking for a limit that does not exist. Brownian motion is the mathematical model, and its paths are continuous and nowhere differentiable with probability one — the same pair of properties, arrived at by an entirely different route.

The relationship to a random walk is direct. A walk’s position after n steps has typical size √n, so over a time interval of length h the change is about √h, and the chord slope is √h/h = 1/√h, which grows without bound as h shrinks. That is the same divergence the figure measures, produced by scaling rather than by a series.

The practical consequence is worth naming, because it changes what can be asked of a model. A quantity whose path has this character has no instantaneous rate of change, so any equation written with one is meaningless as it stands — which is why stochastic calculus had to be built from scratch rather than borrowed, and why the rules in it are not the rules of ordinary calculus. The chain rule acquires an extra term, and the extra term is exactly the roughness that the h scaling measures.

A second consequence runs the other way. A measurement of a rough path taken at finite resolution always produces a finite slope, and that slope depends entirely on the resolution — halve the sampling interval and the measured slope changes by a predictable factor. Anybody measuring the roughness of a coastline, a price series or a surface is measuring an exponent rather than a value, and the exponent is the only thing that survives a change of instrument.

What it costs

Everything drawn here is a partial sum, since an infinite series cannot be evaluated — and how many terms to draw is not a free choice, which is the interesting part of the cost.

The kth term oscillates bᵏ times per two units, so a picture of a partial sum has a finest wiggle whose width is fixed by the term count and the span. Sampling below about eight points per wiggle does not draw the function; it draws an interference pattern between the function and the sampling grid, which looks convincingly rough and is not the curve. The first version of the hero here did exactly that — six terms at b = 7 is sixteen thousand oscillations, sampled four thousand times — and it produced a figure that was both wrong and 248 KB of coordinates.

Each curve is now sampled at ten points per oscillation of its own finest term, which is 1,215 samples for the six-term curve’s 122 oscillations, and the generator refuses to draw any curve at fewer than eight. The magnification figure applies the same rule per panel: each window is drawn with as many terms as its sampling resolves — five, seven, nine and twelve of them — and the terms left out are asserted to contribute less than a twentieth of what the window shows.

So the pictures are honest at the scales drawn and cannot be extended indefinitely. The theorem is about all scales at once, and no drawing reaches past the resolution of its own grid.

What the picture cannot show

The claim is that no point of the line has a tangent, and every figure examines finitely many points. The chord measurement is taken at one place; the zoom is centred on one place. Nothing about the drawings rules out a single well-behaved point somewhere else, and the proof that there is none is a piece of analysis with no picture.

Nor can the figures distinguish a curve with no tangent from a curve whose tangent is very hard to see. A curve made from a finite number of terms does have a tangent everywhere; it is a finite sum of cosines and perfectly smooth. Every picture on this page is of such a curve — the six-term partial sum has a derivative at every point, and that derivative is merely large. What is drawn is a smooth approximation to a non-smooth limit, and the non-smoothness is entirely in the passage to the limit that no drawing performs.

And the self-similarity is only statistical. The curve does not repeat exactly under magnification, so a reader looking for an exact copy in the deepest panel of the zoom will not find one.

The ladder from here

Rungs above: Hardy’s version of the theorem, with the condition on b removed. The Takagi function, which is nowhere differentiable, built from triangle waves rather than cosines, and easier to analyse. The Hausdorff dimension of the graph, which is 2 + log a/log b — a number between 1 and 2, so the curve is genuinely rougher than a line and genuinely thinner than a region. Brownian motion, where the same pair of properties holds with probability one. The Banach–Mazurkiewicz theorem, which says that nowhere-differentiable is the typical case. Functions differentiable everywhere with an unbounded derivative, and functions differentiable exactly once. And the space-filling curves, which are continuous and worse.

The shape of the idea

Two limits are involved and the essay is about which of them commutes with which.

Summing the series is one limit. Taking the derivative is another. For a finite sum they interchange freely — the derivative of a sum is the sum of the derivatives, and every partial sum here is differentiable. What Weierstrass’s example shows is that the interchange fails in the limit, and it fails so badly that one side does not exist.

That is the standing shape of a counterexample in analysis: not a strange object built for its own sake, but a demonstration that two operations everybody performs without thinking cannot be performed in either order. The rearranged series makes the same point about addition and reordering; the staircase that is not the diagonal makes it about limits and length.

Each of them is remembered as a monster and each is really a piece of bookkeeping about which limits may be swapped. Since analysis is almost entirely the study of interchanging limits, the monsters are the subject’s actual content and the well-behaved cases are the corner it happens to have started in.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

ContinuityConvergenceCounterexampleDifference quotientDifferentiabilityFourier seriesFractal dimensionSelf similarityUniform convergenceWeierstrass function