Adding up rectangles until they stop being rectangles
There is an area under a curve. Rectangles have areas; this shape does not, or not yet. The curve is not straight, so none of the formulas for areas of straight-sided things apply. That is the entire problem, and the solution is to cover the region with shapes that are straight-sided, accept that the cover is wrong, and then make it less wrong without limit.
Wrong in a controlled way
The rectangles above use the curve’s height at the left edge of each strip. Since this curve is rising, each rectangle is too short, and the total underestimates. Using the right edge instead overestimates by a comparable amount. Neither is the answer; both bracket it.
The useful observation is not that the estimate is close. It is that the error is controlled by the width of the strips. Slide the rectangles’ tops up and down within each strip and the total can only move by as much as the curve rises across that strip. Halve the strip width and, for a well-behaved curve, halve that slack. The error is not merely small — it is small in a way that can be made as small as required by turning one dial.
For the curve drawn here the exact area is , and the left sums climb toward it from below: , then , then . They are still short at thirty rectangles, and they will still be short at thirty thousand. The point is not the specific numbers. The point is that the sequence of numbers has a limit, that the limit is , and that it does not depend on how the region was chopped up.
That last clause is doing more work than it appears to.
What “the integral exists” claims
Here is the definition, more or less as Riemann gave it in 1854. Chop the interval into strips, not necessarily equal ones — the freedom matters, in the way that the freedom to choose any and mattered to the dissection proof. In each strip, pick any point and use the curve’s height there for the rectangle. Add up. Now consider all possible choppings and all possible choices of sample point, and ask what happens as the widest strip shrinks to zero.
If every such sequence of sums approaches the same number, the function is integrable and that number is the integral.
The insistence on every choice is the substance. It would be much easier to define the integral using, say, equal strips and left endpoints — that always produces some number for any bounded function whatever. But such a definition would be worthless, because it would assign an “area” to functions that do not have one in any meaningful sense, and the answer would change if the strips were chosen differently.
The function that demonstrates this is Dirichlet’s: when is rational and when it is irrational. Chop the interval however finely. In every strip, no matter how narrow, there are both rationals and irrationals. Sample the rationals and every rectangle has height 1, giving a total of 1. Sample the irrationals and every rectangle has height 0, giving 0. The sums do not converge on anything; they depend entirely on the sampling, forever, at every level of refinement.
So Dirichlet’s function has no Riemann integral. Not because the calculation is hard, but because there is no number for the calculation to produce. This is worth knowing about, because it is the point at which “area under the curve” stops being a description of something obviously there and becomes a claim requiring proof.
It is worth being clear that this is not a pathology invented to be awkward. Dirichlet’s function is the simplest thing that distinguishes “has an area” from “the area has not been worked out yet”, and until it existed there was no way to say what integrability meant. It is the object that turned a word into a definition. Counterexamples of that kind do real work: the Möbius band plays exactly the same role for the word side.
The rectangles are a scaffold
A common misreading of the picture is that the integral is approximately the sum of a great many very thin rectangles — that with enough of them, the staircase is near enough the curve.
That is not the claim. There is no number of rectangles at which the staircase becomes the curve; at every finite stage the staircase is still a staircase, with corners, and its total is still wrong. What the limit asserts is that the sequence of wrong answers converges, and the integral is defined as the thing it converges to.
The distinction matters because it is the difference between a numerical method and a definition — the same distinction as between checking that no route across Königsberg works and proving that none can. Slicing into a million rectangles is a way of computing an approximate answer. The integral is the exact answer, and it is exact because the limit is exact, not because a million is a large number.
Infinitely many things are being added, each of them infinitesimally small, and the product of those two infinities is an ordinary finite number. That combination made people nervous for about two hundred years, with good reason, and the limit definition exists precisely to make it respectable. Newton and Leibniz got the right answers in the 1660s and 1670s — much as Galileo had read the odd-number rule off an inclined plane without the machinery to justify it; the rigorous account of why they were right took until Cauchy and Riemann in the nineteenth century.
The paradox that made the limit necessary
The limit definition looks like fussiness until the alternative is tried, and the alternative was tried for a century.
Bonaventura Cavalieri, working in the 1630s, treated a region as being made of its vertical lines — indivisibles, with no width at all — and compared two regions by pairing off their lines. The method works beautifully and gives right answers: two solids with equal cross-sections at every height have equal volume, which is Cavalieri’s principle and is still taught under that name.
Then it gives a wrong one. Take a rectangle and cut it along a diagonal into two triangles of visibly different size — say the diagonal from one corner to a point partway along the opposite side. Every vertical line meets the left triangle in some segment and the right triangle in some segment, and the two families of segments can be put in one-to-one correspondence. Pair each line of one triangle with a line of the other and every pair has been matched, with nothing left over on either side. The triangles must therefore have equal area.
They do not. One is visibly larger.
The pairing is real and the conclusion is false, and the reason is that a region is not the set of its lines in any sense that respects area. Each line has zero width; a matching of lines carries no information about how densely they are packed, and the two triangles pack theirs differently. Adding up zero-width objects is not an operation with an answer, and no amount of care in the pairing repairs it.
Riemann’s rectangles fix exactly this. A rectangle has a width, so the sum is a sum of genuine areas and is a genuine number at every stage; the vanishing happens in a limit taken afterwards rather than being built into the objects being added. That is the whole methodological difference, and it is why the definition is a limit of finite sums rather than a sum of infinitesimal ones — a distinction that took two centuries to become non-negotiable, and that Cavalieri’s paradox is the cheapest demonstration of.
The same limit, one dimension up, converges to whatever it likes
There is a natural expectation that the procedure generalises: chop finer, add up, take the limit, and the answer is the area of whatever curved thing was being approximated. For a curve over an interval it works. For a surface it fails, and it fails in a way that is worth seeing because it retroactively explains how much the one-dimensional case was getting for free.
Take a cylinder and inscribe a triangulated surface in it — vertices arranged in rings around the cylinder, triangles zig-zagging between adjacent rings. Every vertex lies exactly on the cylinder. As both the number of rings and the number of vertices per ring grow, every triangle shrinks to nothing and the whole surface converges to the cylinder, in the ordinary sense that no point of it is far from the cylinder.
The total area of the triangles does not converge to the cylinder’s area. It converges to whatever the construction is told to converge to. Let the ring count grow like the square of the vertex count and the triangle areas total something strictly larger than the cylinder’s; grow it faster and the total goes to infinity. The same sequence of surfaces, in the same shrinking sense, with a limiting area that depends entirely on the ratio between two counts that both go to infinity.
This is the Schwarz lantern, and Hermann Schwarz produced it in the 1880s to demolish exactly the definition a reader would have written down. What goes wrong is that the triangles, while small, are increasingly tilted relative to the surface they approximate — nearly edge-on, so each one contributes far more area than the patch of cylinder beneath it. Closeness of position was never the same as closeness of orientation, and area is a quantity that cares about the second.
The rectangles under a curve escape this only because they are constrained to be vertical: there is no orientation left to go wrong. That constraint reads as a convenience of the drawing and is doing load-bearing work, and the one-dimensional case’s good behaviour is not evidence of anything about the general method. This is the sharpest available answer to why not simply define area as a limit of approximations: for surfaces, that definition assigns no number at all.
The shortcut nobody expects
Having gone to this trouble, the practical situation is comic: almost nobody ever computes an integral by adding up rectangles.
The reason is the fundamental theorem of calculus, which observes that integration and differentiation undo each other. If a function has derivative , then the area under from to is just . Two evaluations and a subtraction, instead of an infinite limiting process.
That this works is genuinely startling. Rates of change and accumulated totals are not obviously related; one is about what a function is doing at a single instant, the other about everything it has done across an interval. The theorem says they are inverse operations. The intuition, once seen, is hard to unsee: the accumulated area, considered as a function of where the region stops, grows at exactly the rate given by the curve’s current height — because pushing the right edge out by a hair adds a sliver of area whose height is the curve and whose width is the hair.
So the rectangles are scaffolding. They establish that the thing exists and what it means, and then the fundamental theorem provides the ladder that makes the scaffolding unnecessary for calculation.
The same instinct, elsewhere
Cut into pieces, count the pieces, let the pieces shrink. It is the same instinct as stacking odd numbers into a square — where the pieces are shells and never shrink at all — and it is precisely the instinct behind estimating by dropping needles on a floor, where the pieces are random trials and the limit is a probability rather than an area.
What the picture cannot show
Every figure here draws a finite number of rectangles, and the integral is defined by what happens as that number grows without bound. The gap between those two things is the entire content of the definition, and no drawing crosses it.
Worse, the pictures are actively misleading in one respect: they make convergence look obvious. Thirty rectangles hug the curve so closely that the eye reports the job as done, when in fact the sum is still short by 0.14 and will still be short at thirty million. The rate matters as much as the fact — and the rate is invisible. Measured on this curve at thirty strips: the left sum is out by , the trapezoid rule by , the midpoint rule by , and Simpson’s rule by , which is zero to the precision the arithmetic carries. Doubling to sixty strips divides the left error by and the midpoint and trapezoid errors by — first order against second, exactly as advertised. Four pictures at would look nearly identical to each other and span twelve orders of magnitude in accuracy.
The Simpson entry should not be glossed over, because it is the example flattering the method rather than the method being that good. Simpson’s rule integrates any cubic exactly, and the curve drawn throughout this essay is , a quadratic. Its error here is therefore not small but zero, and it is zero at four strips as much as at thirty. On a curve the rule cannot integrate exactly the error falls like — still far better than anything else on the list, but a rate rather than an identity.
That is worth watching for generally. A numerical claim demonstrated on one example may be reporting the method’s behaviour or the example’s, and the two look identical in a table of numbers. Here the give-away is that the error did not improve when rose, which is not what an error does.
The ladder from here
Rungs above this one: the midpoint and trapezoid rules, and why one is unexpectedly good. Simpson’s rule, and the polynomial that explains it. The fundamental theorem given its own picture, with the sliver of area whose height is the curve. Improper integrals, where the region is unbounded and the answer is finite anyway. Arc length, surface area, and volumes of revolution — the same limit applied to different slices. Lebesgue’s construction, chopping the range instead of the domain. Measure zero, and the exact criterion for Riemann integrability. And numerical quadrature, where the question stops being whether the limit exists and becomes how few evaluations can get within a tolerance — the point at which this ladder meets Monte Carlo.
Riemann’s construction was superseded, incidentally. Lebesgue’s integral, from 1902, chops up the range rather than the domain — a change of viewpoint of the same kind as reading a matrix by its columns rather than its entries — grouping together all the places where the function has roughly the same value, rather than all the places that are near each other. It handles Dirichlet’s function without difficulty, assigning it the integral 0, and it is the definition analysts actually use. But the rectangles came first, and they are still the picture people carry.
What links here
Computed from the collection, not written here: the essays that point at this one.
- Area is the undoing of slope
- A circle unrolled into a triangle
- A sum whose terms vanish and whose total does not
- The same terms, in a different order, adding to whatever is asked
- A square wave built entirely out of round ones
- Bayes' theorem is a picture of a square
- Completing the square, by completing a square
- Getting pi by dropping needles on the floor
- and 12 more
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A bell curve assembled out of coin flips — both name convergence, limit
- A walk that always comes home, until it does not — both name convergence, limit
Named objects
A dashed tag is an object no other essay names yet.
AreaConvergenceDirichletDirichlet's functionFundamental theoremIntegrabilityLebesgue integralLimitRiemann sum