Which functions can be added up
Worth reading first: Covering a set from outside · Adding up rectangles until they stop being rectangles.
Riemann’s integral is defined by squeezing: approximate the area from above and from below, refine the partition, and if the two approximations meet, that common value is the integral. The question this rung answers is exactly which functions the squeeze works on.
The answer is Lebesgue’s, from 1904, and it could not have been stated before the previous rung: a bounded function on a bounded interval is Riemann integrable exactly when the set of points where it is discontinuous has measure zero.
The squeeze, stated precisely
Partition the interval into cells. On each cell take the largest value the function attains and multiply by the cell’s width; add those and the result is the upper sum. Take the smallest value instead and the result is the lower sum.
Refining a partition can only lower the upper sum and raise the lower one, so the upper sums have an infimum and the lower ones a supremum, and the first is never below the second. The function is Riemann integrable when they are equal.
So the whole question is whether the gap between them can be driven to nothing, and the gap on a cell is the function’s oscillation there — the difference between its largest and smallest values. The integral exists exactly when the total oscillation, weighted by cell widths, can be made arbitrarily small.
That reformulation is what makes the criterion findable: oscillation is a local quantity, and a function is continuous at a point exactly when its oscillation on small cells around that point goes to zero.
Two indicators, one partition
The clearest test case is an indicator function — one on a set, zero off it — because its oscillation on a cell is one where the cell meets both the set and its complement, and zero otherwise.
Take the middle-thirds set. Its indicator is discontinuous at every point of the set, since every neighbourhood of such a point contains removed intervals. So the discontinuities are exactly the set, which has measure zero, and the criterion says the function is integrable with integral zero.
Take the fat Cantor set. Its indicator is discontinuous at every point of that set too, and the set has measure about . So the criterion says it is not integrable — the upper and lower sums never meet.
Two functions of identical shape, one integrable and one not, distinguished by the measure of a set the picture cannot show a difference in. That is the whole content of the rung, and it is why the fat Cantor set was constructed in the first place.
The lower sum is nothing in both cases
Half the computation is immediate and it is worth doing first, because it isolates where the difference lives.
Neither set contains an interval. So every cell of every partition, however fine, contains points outside the set, where the indicator is zero. The smallest value on every cell is therefore zero, and the lower sum is zero for every partition.
The figure asserts exactly that: no cell of the drawn partition lies wholly inside either set. That is the property being nowhere dense gives, and both sets have it.
So integrability is decided entirely by the upper sum, and integrability would mean the upper sum comes down to zero. For the middle-thirds set it does; for the fat one it cannot, because the upper sum is at least the set’s own measure.
Why the upper sum bottoms out at the measure
The lower bound is the part worth arguing carefully, since it is what makes the failure permanent rather than a matter of trying harder.
The cells that meet the set cover the set, so their total width is at least the set’s outer measure — that is the definition of outer measure, which takes the smallest total over all coverings. And the upper sum is exactly the total width of the cells that meet the set, since the indicator’s maximum there is one.
So the upper sum is bounded below by the measure, at every partition, however fine. No refinement escapes a bound that comes from the definition of the measure being refined against.
For a set of measure zero the bound is zero and says nothing; the upper sum still has to be shown to come down, which it does. For a set of positive measure the bound is positive and the sum is trapped above it forever.
How slowly the good case converges
The middle-thirds indicator is integrable, and the rate is worth looking at because it is much slower than one might expect.
At equal cells, the number that meet the middle-thirds set is roughly , since is the set’s dimension. So the upper sum is about , which goes to zero and does so at a crawl: at twenty-four thousand cells it is still .
The figure’s sweep shows exactly that. Six partition sizes spanning a factor of a thousand bring the upper sum from over a half down to — genuine convergence, visibly not finished.
A convergence that slow is why the case cannot be settled by looking at one partition, and it is why the figure draws a sequence rather than a single picture. At ninety-six cells the two upper sums are both a substantial fraction of the interval, and neither number tells a reader which of them is going anywhere.
Where the fat set’s upper sum settles, and why it never quite arrives
The other curve is worth reading as carefully, because it does move and its limit is not zero.
At twenty-four cells the fat set’s approximation meets nearly every cell, so the upper sum is close to one. As the partition refines, the cells that meet only the removed middles drop out, and the sum comes down towards the measure of the set — — which it approaches from above and never passes.
The figure prints the last value at and draws a dashed line at the set’s own length, and the gap between them is the whole story of the convergence: the upper sum is still counting cells that overlap the removed middles, and refining removes them one scale at a time.
Both curves are converging and only one is converging to zero, which is the honest description and is more useful than saying one converges and the other does not. Integrability is a statement about the limit, and the limits here are and .
There is a small asymmetry worth naming. The fat set’s upper sum reaches within a per cent of its limit by six thousand cells; the thin set’s is still a factor of twenty above its limit at twenty-four thousand. The one that converges to something does so quickly, and the one that converges to nothing takes forever — which is the opposite of what one would guess.
What the criterion says about the rationals
The other standard example runs the other way, and it is worth putting beside these.
The function that is one at the rationals and zero elsewhere is discontinuous everywhere, so its discontinuity set is the whole interval, of measure one. Not integrable — and here the failure is obvious, since the upper sum is one and the lower is zero at every partition.
The function that is at the rational in lowest terms and zero at the irrationals is continuous at every irrational and discontinuous at every rational. Its discontinuity set is the rationals, of measure zero, so it is integrable, with integral zero.
A function discontinuous at a dense set can be integrable, which is the same surprise the previous rung’s covering delivered, in the form it takes for functions.
A restatement in terms of continuity points
There is a second reading of the criterion that is easier to use, and it comes from taking the complement.
A bounded function is Riemann integrable exactly when it is continuous almost everywhere — at every point outside a set of measure zero. Stated that way, the criterion says the integral notices only how much discontinuity there is, never where it is or how dense.
Two consequences follow immediately. A monotone function has at most countably many discontinuities, so every monotone bounded function is integrable — a result that used to need its own proof. And a function with only jump discontinuities on a countable set is integrable, however the jumps are arranged.
What is not implied is anything about the function’s values. The criterion is about continuity alone, so a wildly oscillating function that is nevertheless continuous almost everywhere integrates fine, and a function that takes only two values can fail.
A limit of continuous functions need not be continuous, and the criterion is what makes that fact dangerous: a sequence of integrable functions can converge pointwise to a non-integrable one, so the Riemann integral does not survive limits. Repairing that is the other half of what Lebesgue’s integral is for.
What Lebesgue’s integral does instead
The criterion is a diagnosis, and the treatment is a different integral.
Riemann partitions the domain into cells and asks for the function’s range on each. Lebesgue partitions the range into bands and asks for the measure of the set where the function lands in each. For the indicator of the fat Cantor set that is trivial: the function is one on a set of measure and zero elsewhere, so the integral is and there is nothing to squeeze.
The reason the second approach works where the first does not is that it never needs the set to be an interval. Riemann’s cells are intervals by construction, and a set that is nowhere dense forces every cell to contain both values; Lebesgue’s bands are arbitrary measurable sets, and the fat Cantor set is one.
So the answer to which functions can be added up depends on which integral is meant, and the criterion above is exactly the description of the gap between the two. Every Riemann integrable function is Lebesgue integrable with the same value; the converse fails, and the fat Cantor set’s indicator is the smallest witness.
Volterra’s function, and why this mattered at the time
The criterion looks like tidying, and it was prompted by a genuine crisis.
Volterra built in 1881 a function differentiable everywhere on , with bounded derivative, whose derivative is discontinuous on a fat Cantor set. By the criterion, that derivative is not Riemann integrable.
So there is a function with everywhere, bounded, and no Riemann integral of . The fundamental theorem of calculus, in the form integrating a derivative recovers the function, has no content for this pair — not because the derivative misbehaves at infinity or blows up, but because its discontinuities are arranged in a set of positive measure containing no interval.
That is a failure at the centre of the subject rather than at its edges, and it is the reason a new integral was wanted. Lebesgue’s handles Volterra’s function, and the fundamental theorem holds for it in the form the derivative deserves.
Oscillation, and the sets where it is large
The proof of the criterion is worth sketching because the object it introduces is useful elsewhere.
Define the oscillation of at a point as the limit, as neighbourhoods shrink, of the difference between the supremum and infimum there. The function is continuous at a point exactly when its oscillation there is zero, so the discontinuity set is the union over of the sets where the oscillation is at least .
Each of those sets is closed, and a countable union of them is the whole discontinuity set. If each has measure zero the union does; conversely if the union has measure zero each does. So the criterion reduces to controlling one closed set at a time, and on a closed set of measure zero a compactness argument produces a finite subcover of small total length — which is where the upper sum’s descent comes from.
Compactness is doing the work again, exactly as it did in showing an interval’s measure is its length. The two arguments are the same argument, and both are the reason closed bounded sets are the ones the theory is comfortable with.
The order the ideas arrived in
The criterion is dated 1904 and its ingredients arrived over half a century, in an order worth recording.
Riemann defined the integral in 1854, in his habilitation thesis, and gave a condition for integrability that was correct and unusable: he required that for every , the total length of the cells on which the oscillation exceeds can be made arbitrarily small. That is Lebesgue’s criterion with no word for the quantity it is about.
Smith in 1874 and du Bois-Reymond in 1875 built the sets showing the condition has teeth. Volterra in 1881 built the derivative that fails it. Jordan in 1892 gave the first systematic notion of content, which was not countably additive and so could not state the criterion. Borel in 1898 introduced countable additivity for a restricted family of sets, Lebesgue in 1901 defined measure for all sets, and in 1904 the criterion could finally be said.
So the statement was in Riemann’s own thesis and nobody could see it for fifty years, because the phrase has measure zero did not exist. That is a common pattern in this subject: the theorem waits for a noun.
The lesson for a reader is that Riemann’s condition and Lebesgue’s are the same condition, and the difference is entirely one of language. The word turned an unusable criterion into a usable one without changing what it says.
What the pictures cannot show
The sweep stops at twenty-four thousand cells and the theorem is about all of them. The middle-thirds curve is at and heading down; that it reaches zero is the criterion, and no finite sweep establishes it.
The two shaded rows look the same. At ninety-six cells the fat set’s row is denser than the thin one’s, which is a fact about that partition rather than about the sets. The distinguishing quantity is the limit of the shaded fraction, and a limit is not a picture.
And the approximations are the drawn objects, not the sets. Every row shades the cells meeting a stage-eleven approximation, which contains the set. The bound the argument needs — that the upper sum never falls below the true measure — is asserted against the limiting measure, computed as an infinite product, rather than against the approximation’s.
Where the ladder goes next
Everything so far has assumed that outer measure adds up on the sets in play, and the last rung shows the assumption cannot be made in general. A set with no size at all exists, its rational translates tile an interval, and the arithmetic of countably many copies of one length lands nowhere the answer could be.
Sideways: the fat Cantor set is the object whose indicator broke the integral, and the rectangles the integral is built from are what the criterion is about.
What is worth carrying away
A criterion is worth more than a list of examples, and it usually arrives late.
Riemann defined his integral in 1854 and the exact condition for it to exist was found fifty years later, once the language to state it existed. In between there were examples of functions that could not be integrated and no way to say what they had in common.
The pattern is that a subject invents its notion of negligible only when forced. Measure zero was invented to answer this question and then turned out to be the right notion of negligible for the whole of analysis — which is what usually happens when a definition is built to make one theorem come out.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A curve that has area — both name cantor set, continuity, limit, measure
- Almost none of it left, and still uncountably many — both name cantor set, measure, measure zero
- A ball whose outside is not one — both name cantor set, limit
- Area is the undoing of slope — both name continuity, limit
- The curve that is its own slope — both name continuity, limit
- The slope of a single point — both name continuity, limit
Named objects
A dashed tag is an object no other essay names yet.
Cantor setContinuityDense setLimitMeasureMeasure zeroThe integralUpper sum