Analysis

Uniform, except on a small set

The functions xⁿ settle on their limit at every point and never uniformly — the trouble is all in a strip next to 1. Throw the strip away and the convergence is uniform on what is left, however thin the strip. Egorov proved that this always happens on an interval, Lusin proved the matching fact about a single function, and a bump sliding off along the whole line shows why both need a set of finite length to start from.

Worth reading first: A limit can jump at every fraction · Covering a set from outside.

At every point left of 1, the powers xnx^n fall to 00; at 1 they stay at 1. They converge at every point and not uniformly: however large nn is, some point just left of 1 has xnx^n close to 1 while the limit there is 0. The largest gap between a member and the limit never shrinks.

But the failure is confined. Everything that goes wrong happens in a strip next to 1, and the strip can be as thin as anyone likes. Remove it and what is left converges uniformly.

xⁿ is uniform off a strip, with N = 14, 29, 59 for 3 strips. The functions x to the n on the unit interval with a band of half-width 0.05 around zero. For each of 3 strips next to 1, the member at which every later one stays in the band away from the strip.
Fig. 1 xnx^n against the band of half-width 0.050.05 round the limit 00. Once the strip next to 11 is removed, a single member count NN puts every later member inside the band: the strip 0.20.2 needs N=14N = 14, 0.10.1 needs N=29N = 29, 0.050.05 needs N=59N = 59.

Off the strip [1δ,1][1 - \delta, 1] the largest value of xnx^n is (1δ)n(1 - \delta)^n, and that falls to 00 as nn grows. So for a tolerance of 0.050.05, the members from NN on are all within the band as soon as (1δ)N(1-\delta)^N is below 0.050.05. For a strip of width 0.20.2 that is N=14N = 14; for 0.10.1, N=29N = 29; for 0.050.05, N=59N = 59. Each strip has a finite NN, and the NN grows without bound as the strip shrinks, which is why the strip cannot be thrown away entirely.

Egorov’s theorem

Dmitri Egorov proved in 1911 that this is not a feature of xnx^n. If functions on an interval converge at every point, then for any length δ\delta, however small, there is a set of total length less than δ\delta off which the convergence is uniform. Carlo Severini had found the same result a year earlier, in a paper on series of orthogonal functions that went unnoticed; the theorem usually carries Egorov’s name.

The hypothesis is weak and the conclusion strong. Nothing is assumed about continuity — the theorem holds for any functions for which lengths of the relevant sets make sense — and nothing about monotonicity or bounds. Convergence at each point separately is enough to buy uniform convergence almost everywhere, with “almost” meaning off a set that can be made as short as desired.

Littlewood later put it as the third of his three principles of analysis: every set of finite length is nearly a finite union of intervals, every function is nearly continuous, and every convergent sequence of functions is nearly uniformly convergent. Each “nearly” means “except on a set of small length”, and the second and third principles are the two theorems on this page.

The proof shows exactly where the length of the interval enters. For a tolerance 1/k1/k and a stage NN, collect the points at which every member from NN onwards is within 1/k1/k of the limit. As NN grows these collections increase, and together they contain every point, because at every point the sequence eventually stays within 1/k1/k. On an interval of finite length, an increasing family of sets that fills it must eventually leave out less than any given length. So for each kk there is an NkN_k whose collection misses less than δ/2k\delta/2^k.

Throw away all the missed points, for all kk at once. The total thrown away is less than δ/2+δ/4+δ/8+=δ\delta/2 + \delta/4 + \delta/8 + \cdots = \delta — the square that fits the whole geometric series — and on what is left, for every tolerance 1/k1/k, every member from NkN_k on is within 1/k1/k of the limit. That is uniform convergence.

Why the small set cannot have no length

Egorov’s set is small in length, and it is natural to hope for more — that the convergence might be uniform off a set of length zero, as small as a set can be and still be something. For xnx^n the hope fails, and the reason is instructive.

A set of length zero contains no interval, so its complement meets every interval, and in particular meets every interval (1η,1)(1 - \eta, 1) next to 1. At a point xx of the complement close enough to 1, xnx^n is as close to 1 as desired for any fixed nn, while the limit there is 0. So on the complement the largest gap is still 1 for every member, and the convergence is not uniform there. Whatever is removed has to contain a whole strip next to 1, of some positive length.

That is why the theorem is stated for every positive δ\delta rather than for δ=0\delta = 0. The small set is small in the strongest sense available — shorter than any prescribed length — but its length is never zero unless the convergence was uniform to begin with, up to a set that cannot matter. The distinction between “almost uniform” and “uniform almost everywhere” is exactly this, and only the first is true.

It also shows that the strip in the first figure is not an artefact of the choice of xnx^n. Any sequence of continuous functions converging to a limit with a jump must fail uniformly near the jump, because a uniform limit of continuous functions is continuous; so any set whose removal makes the convergence uniform must contain a neighbourhood of the jump, cut back from both sides. Egorov’s theorem says only that such neighbourhoods can be taken short.

The spike, trimmed

A second sequence shows the same repair in a different place.

a triangle of area one at 4 values of n, and the limit. Several members of the sequence a triangle of area one drawn on one pair of axes with the function they settle on, and the largest gap between each member and that limit reported.
Fig. 2 A triangle of area one at n=2n = 2, 44, 88 and 1616, standing on the interval near 00 and growing taller as it narrows. Every column settles to zero while the largest gap to the limit goes 22, 44, 88, 1616.

The travelling spike rises to height nn over a base of width 2/n2/n next to 00. At each point x>0x > 0 it eventually passes, and the values are 00 from then on; at 00 itself they are always 00. So the sequence converges to 00 everywhere, and not uniformly — the largest gap is nn, which grows rather than shrinking.

Egorov’s theorem promises a thin strip whose removal makes it uniform, and the strip is [0,δ)[0, \delta). Once 2/n2/n is less than δ\delta, the whole spike lives inside the strip and every member is exactly 00 off it. The convergence off the strip is not merely uniform but eventually constant.

What the removal does not repair is the integral. Each spike has area 1 while the limit has area 0, and the area lives inside the strip that was thrown away. Removing a small set restores uniform convergence and with it every conclusion uniform convergence supports on what is left; it cannot restore conclusions about the whole interval that depend on what happens inside the small set. That limitation is the reason integration theory needed a separate theorem — dominated convergence — which asks for a single integrable bound on all the members, a bound the spike fails.

Why the length has to be finite

The step in the proof that used the interval’s finite length was the claim that an increasing family of sets that fills the interval eventually leaves out little. On the whole line it is false, and so is the theorem.

A bump that escapes along the line. Four members of a sequence of triangular bumps of height one moving right along the line; each point is eventually zero, while every member has height one somewhere.
Fig. 3 A bump of height 11 sliding right along the whole line, at n=1n = 1, 33, 55 and 77. Every point is passed once and then left at 00, so the limit is 00 everywhere; but every member reaches 11 somewhere.

Let the nn-th member be a bump of height 1 centred at nn. Every point is eventually behind the bump, so the sequence converges to 0 everywhere. For the convergence to be uniform off a set EE, the members would have to be below any tolerance off EE from some stage on, and a member is below a half only outside the middle unit of its bump. So EE would have to contain the middle of every bump from that stage on — infinitely many separated intervals of length one. No set of finite length does that, let alone a small one.

The collections from the proof show the failure directly. For tolerance a half and stage NN, the points where every member from NN on is small are the points left of N1/2N - 1/2. These increase and fill the line, as the proof requires, but each misses an infinite length. An increasing family filling a finite interval runs out of room to miss anything; on the line there is always more room.

The sliding bump is the travelling spike moved from the edge of a finite interval to infinity. Both concentrate their bad behaviour somewhere; in one case the somewhere is a strip of small length, and in the other it is a region of infinite length out towards infinity.

A single function, nearly continuous

Egorov’s theorem is about a sequence. Nikolai Lusin proved in 1912 the corresponding statement about one function: any function for which lengths of level sets make sense agrees with a continuous function except on a set of arbitrarily small length — more precisely, it becomes continuous when restricted to a closed set whose complement is as short as desired. The two theorems are related: Lusin’s is proved by approximating the function by step functions and applying Egorov’s to the approximations.

The most discontinuous function met so far is the indicator of the rationals, which jumps at every point. Lusin’s theorem applies to it too, and the small set is easy to exhibit.

Every fraction covered, in total length under 0.4. The fractions with denominator up to 7 on the unit interval, each inside a covering interval whose length halves from one fraction to the next; the uncovered remainder carries no fraction.
Fig. 4 The 1919 fractions with denominator up to 77, each covered by an interval whose length halves from one fraction to the next, starting from 0.40.4. All the covers together, and those of the fractions after them, fit in total length 0.40.4; here the covered part is 0.2500.250 of the interval.

List the fractions of the interval in any order that reaches them all — by denominator, say, which is how every fraction can be counted. Cover the first by an interval of length ε/2\varepsilon/2, the second by one of length ε/4\varepsilon/4, the jj-th by one of length ε/2j\varepsilon/2^j. The covers have total length at most ε\varepsilon, which is the reason the rationals have no length at all. Remove them. What is left is closed, has length at least 1ε1 - \varepsilon, and contains no fraction.

On what is left the indicator of the rationals is the constant 0, and a constant is continuous. So the function that is discontinuous at every point becomes continuous after removing a set shorter than any chosen ε\varepsilon.

The phrase “becomes continuous” needs reading carefully. The restriction to the closed set is continuous as a function on that set. The original function is still discontinuous at every one of those points, because each of them has fractions arbitrarily close — fractions that were removed. Lusin’s theorem does not say the function has continuity points; Baire’s does, for limits of continuous functions, and the indicator of the rationals has none. The two results measure different things, and they coexist here without conflict.

Convergence that settles nowhere

Egorov’s theorem sits in the middle of a list of ways a sequence of functions can converge, and the list has a member weaker than all the others.

The typewriter sequence: 15 members, a block halving each level. Rows showing successive members of the typewriter sequence, each the indicator of a dyadic interval, sweeping across the unit interval level by level with the first member of each level marked.
Fig. 5 The first 1515 members of the typewriter sequence, one per row: a block of height 11 stepping across the interval, halving in width at each new level. The first block of each level is marked; the widths fall from 11 to 1/81/8 over the rows shown.

The typewriter sequence is the indicator of a block that sweeps across the interval, halves in width, and sweeps again. At level kk there are 2k2^k blocks of width 1/2k1/2^k. The block widths fall to 0, so the set where a member differs from 0 has length falling to 0 — the sequence converges to 0 in measure. But every point is covered by exactly one block in every level, so at every point the values are 1 infinitely often and 0 infinitely often. The sequence converges at no point at all.

The typewriter is also a warning about what length alone can see. At every stage the member differs from 0 on a set of length 1/2k1/2^k, a quantity that falls to 0 steadily; a reader told only those lengths would conclude that the sequence settles. It settles in the sense of length and in no other, because the set where it differs from 0 keeps moving, and every point is visited again at every level.

So convergence in measure does not imply convergence at points, and Egorov’s theorem does not apply — there is no pointwise limit to be uniform towards. What survives is a theorem of Frigyes Riesz from 1909: a sequence converging in measure has a subsequence converging at almost every point. Here the first block of each level — the ones marked — is such a subsequence, the indicators of [0,1/2k][0, 1/2^k], which converge to 0 at every point except 0.

Five kinds, and the arrows between them

With the typewriter the list is complete, and the implications between its members are exactly the theorems on this page and the one before it.

Five kinds of convergence and the arrows between them. A diagram of five boxes — uniform, pointwise, almost uniform, in measure, and a subsequence converging almost everywhere — joined by implication arrows, with the conditions and the counterexamples to the converses.
Fig. 6 Five ways a sequence of functions can converge, each arrow an implication labelled with the condition it needs. The converses fail: xnx^n converges at every point and not uniformly, and the typewriter converges in measure and at no point.

Uniform convergence implies convergence at every point. On an interval of finite length, convergence at every point — or at almost every point — implies uniform convergence off a small set, which is Egorov. That implies convergence in measure, since the set where a member is far from the limit is eventually inside the small set. And convergence in measure implies convergence of a subsequence at almost every point, which is Riesz.

None of the converses holds, and each is broken by one of the sequences above. xnx^n converges at every point and not uniformly. The sliding bump converges at every point of the line and is uniform off no small set, which is Egorov’s hypothesis of finite length failing. The typewriter converges in measure and at no point. A diagram of implications between definitions can look like bookkeeping; with a counterexample on each missing arrow it is a map of which hypothesis buys what.

Two further arrows are worth adding in words. Uniform convergence off a small set implies convergence at almost every point, since a point outside the small sets of every size is a point where the members converge — so the box for uniform convergence off a small set leads back up to convergence at points, give or take a set of length zero. And the arrow from convergence in measure to a subsequence is the one that makes measure-theoretic arguments work in practice: a proof that only controls lengths can always pass to a subsequence and recover pointwise information, which is how most theorems about limits of integrals are eventually brought back to statements about points.

What the figures cannot show

Every figure shows finitely many members of a sequence and finitely many of the fractions, and the theorems are about all of them. That xnx^n needs N=59N = 59 for the strip 0.050.05 is a computation; that some NN exists for every strip is the fact that (1δ)n(1-\delta)^n falls to 0, which is proved and not drawn.

The Lusin figure shows the covers of nineteen fractions. The claim is about the covers of all of them, and the remaining covers are shorter than the pixels available to draw them: the twentieth has length 0.4/2200.4/2^{20}, less than a millionth. That the uncovered set contains no fraction at all is a consequence of the construction reaching every fraction, which a list can only begin.

And the escaping bump is drawn on a window of length nine. The failure of Egorov’s theorem on the line is a statement about what happens outside every window, and a picture of a window cannot contain it. What the figure can show is the mechanism — the bad behaviour moving away rather than concentrating — and the argument supplies the rest.

The question it leaves: when a limit and an integral can be swapped

Egorov’s theorem restores uniform convergence off a small set, and the spike showed what it does not restore: the integral of the limit need not be the limit of the integrals, because the mass of the members can hide in the small set. The question of exactly when the two can be exchanged is the central question of integration theory, and the answers are the convergence theorems of Lebesgue’s integral.

Monotone convergence says the exchange is safe for an increasing sequence, which is how the indicator of the rationals gets integral zero as the limit of spike functions. Dominated convergence says it is safe when a single integrable function bounds every member, and the travelling spike has no such bound. Between those two theorems lies the whole practical content of Egorov’s, because the standard proof of dominated convergence runs through it: remove the small set, use uniform convergence on the rest, and use the bound to control what was removed. Which is the step the spike cannot pass, since its members have no common bound for the small set to be controlled by.

Nearly everything, nearly uniform

The picture this page draws is that pointwise convergence is not as weak as the first example suggested. A sequence that converges at every point of an interval converges uniformly off a set as short as anyone wants, and a single function, however discontinuous, is continuous on a closed set of nearly full length. Both statements are about length rather than about gaps, the counterpart of Baire’s theorem, which described the same limits by where they may jump.

The two conditions that cannot be dropped are finite length, which the escaping bump shows, and convergence at points, which the typewriter shows. With them, the difference between pointwise and uniform convergence is a strip that can be made as thin as desired. Without them, it can be the whole of the space.

The same pattern recurs throughout analysis. A property that fails at a few points, or on a set of zero length, or off a set that can be made as short as desired, is treated as though it held everywhere, and the theorems that justify doing so are exactly the ones drawn here. The question they answer is not whether a limit is well behaved but where its misbehaviour is allowed to live — near a jump, inside a strip, on a set of fractions — and in each case the answer can be measured, which is what makes the misbehaviour harmless rather than merely rare.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

ContinuityConvergenceCounterexampleGeometric seriesMeasurePointwise convergenceRational numberUniform convergence