Algebra

The roots of the slope stay inside

Mark the roots of a polynomial in the complex plane and stretch a band around them. However the roots are arranged, the roots of the derivative land inside the band — never outside, never on a new frontier. The reason is a balance of pushes, the same reason makes the derivative of a cubic mark the foci of an ellipse nobody asked for, and a question about how far inside the roots must sit has been open since 1958.

Worth reading first: Every power sum, from the coefficients alone · A loop that cannot miss the middle.

A polynomial’s roots are points in the complex plane, and so are the roots of its derivative. The derivative has one root fewer, and it is natural to ask where they go relative to the originals. They could, for all anyone has said so far, be anywhere: the derivative is a different polynomial with different coefficients, and the roots of different polynomials are not obliged to have anything to do with each other.

They are obliged, and the obligation is simple to state. Every root of the derivative lies inside the convex hull of the roots of the polynomial — the smallest convex region containing them, what a band stretched round the roots would enclose. That is the Gauss–Lucas theorem. Carl Friedrich Gauss wrote down the physical reason in the 1830s, and Félix Lucas published the theorem with a proof in 1874.

The roots of the derivative inside the hull of the roots. Three polynomials of degree five, each drawn as its roots with their convex hull shaded, and the four roots of its derivative marked. Every root of the derivative lies inside the hull.
Fig. 1 Three polynomials of degree five. The orange points are the roots, the shaded polygon their convex hull, and the blue points the four roots of the derivative, found from the derivative’s own coefficients. Every blue point is inside the hull, the two sets share an average, and when the roots all lie on a line the hull is a segment and the blue points sit between them.

The blue points in the figure were not placed; they were computed by taking the derivative of each polynomial, coefficient by coefficient, and finding its roots numerically. Where they landed was then checked against the hull. In all three panels they are well inside, and in the middle one — three roots bunched in one corner — two of the four sit in the bunch, as though the derivative could see the cluster and responded to it.

The pushes that have to cancel

Every power sum, from the coefficients alone used a single identity to extract power sums from coefficients: the derivative divided by the polynomial is a sum of simple fractions, one per root,

p(z)p(z)=k=1n1zzk.\frac{p'(z)}{p(z)} = \sum_{k=1}^{n}\frac{1}{z - z_k}.

There it was expanded at infinity. Read at an ordinary point instead, it proves the theorem.

Take a root ww of the derivative that is not also a root of the polynomial. At ww the left side is zero, so the fractions on the right cancel. Each fraction 1/(wzk)1/(w - z_k) is the complex conjugate of (wzk)/wzk2(w - z_k)/|w - z_k|^2, so conjugating the whole equation gives

k=1nwzkwzk2=0.\sum_{k=1}^{n}\frac{w - z_k}{|w - z_k|^2} = 0.

Each term is a vector pointing from the root zkz_k towards ww, with length one over the distance. A root of the derivative is a point where these pushes, one away from each root, balance exactly.

The push of five roots, and the four places it balances. A field of arrows showing, at each point of the plane, the direction of the sum of pushes away from each root of a degree-five polynomial. The four points where the pushes cancel are the roots of the derivative, and they lie in the shaded convex hull of the roots.
Fig. 2 At each point of the plane, the direction of k(zzk)/zzk2\sum_k (z - z_k)/|z - z_k|^2 for the five roots of the first panel: a push away from every root of pp, weakening with distance. The four points where it vanishes are the roots of pp', in blue. Outside the shaded hull every push has a part pointing away from it, so nothing out there can balance.

Now the hull is forced. Solve the balance equation for ww: collecting the terms gives wkck=kckzkw\sum_k c_k = \sum_k c_k z_k with weights ck=1/wzk2c_k = 1/|w - z_k|^2, all positive. So

w=kckzkkck,w = \frac{\sum_k c_k\, z_k}{\sum_k c_k},

a weighted average of the roots with positive weights. And any weighted average of points with positive weights lies in their convex hull — that is one of the definitions of the hull, and it is the arithmetic behind a centre is three weights, where every point of a triangle is some balance of its three corners. A root of the derivative that is also a root of the polynomial is in the hull trivially. So all of them are.

The field of arrows shows the same argument as a picture. From any point outside the hull there is a line separating that point from all the roots — the fact at the heart of a wall between two bodies — and every push has a positive component across that line, away from the roots. Pushes that all lean the same way cannot add to zero. Inside the hull the arrows point every which way, and in four places they cancel.

Gauss’s version was physical. Put an equal electric charge at each root of a polynomial, in a flat world where the force between charges falls off as one over the distance rather than one over its square. The points where a test charge would feel no force are exactly the roots of the derivative. The theorem is then the observation that a test charge placed outside a cloud of repelling charges is always pushed further away.

The flat case is Rolle’s theorem

When the roots all lie on the real line, the hull is the segment from the smallest to the largest, and the theorem says the derivative’s roots are real and lie in that segment. That is something every calculus course proves by a different route.

5 real roots, and a turning point in every gap between them. The graph of a polynomial whose roots are all real, with the roots of its derivative marked at the peaks and troughs between them: one in each gap between neighbouring roots.
Fig. 3 A polynomial with five real roots, drawn as a graph, and the four roots of its derivative at the peaks and troughs between them. Between each pair of neighbouring roots the graph turns once: Rolle’s theorem supplies one turning point per gap, and the derivative’s degree leaves no room for a second.

Rolle’s theorem says that between two roots of a smooth function its slope must vanish somewhere, since the graph rises and has to come back down, or falls and has to come back up. With five distinct real roots there are four gaps, so at least four turning points, and a derivative of degree four has at most four roots. So there is exactly one in each gap — the roots of the derivative interlace the roots of the polynomial.

Gauss–Lucas is the statement that survives when the roots leave the line. There are no gaps to count in the plane, and nothing like “between” in the ordinary sense; what survives is the convex hull, which on a line is exactly the stretch from the first root to the last. The interlacing is lost, and the containment is kept. That is a common fate for statements about real numbers carried into the complex plane: the order goes, and a convexity statement remains.

A cubic draws its own ellipse

For three roots the hull is a triangle and the derivative has two roots, and here the theorem sharpens into something that looks like a coincidence and is not.

Marden's theorem: the derivative's roots are the foci of the midpoint ellipse. Three triangles, each the three roots of a cubic. Inside each is drawn the ellipse touching every side at its midpoint, and the two roots of the cubic's derivative are marked: they are the ellipse's foci.
Fig. 4 Three cubics, each with its roots at the corners of a triangle. Inside each is the ellipse that touches every side at its midpoint, and the two roots of the cubic’s derivative, in blue, are its foci. For the equilateral triangle the ellipse is the inscribed circle and the two foci coincide at the centre.

Marden’s theorem: if the roots of a cubic are the corners of a triangle, the roots of its derivative are the two foci of the ellipse inscribed in the triangle that touches each side at its midpoint. That ellipse is called the Steiner inellipse, it is the largest-area ellipse fitting inside the triangle, and its centre is the triangle’s centroid. Jörg Siebeck proved the theorem in 1864; it is named after Morris Marden, whose writing on the geometry of polynomials in the 1940s made it widely known.

The figure checks the claim in the form an ellipse is defined by. An ellipse with foci f1f_1 and f2f_2 is the set of points whose distances to the two foci add to a fixed total, like the string and two pins of the gardener’s construction. So the ellipse drawn is: take the two roots of the derivative as foci, measure the total distance to one midpoint, and draw every point with the same total. The figure then measures the total at the other two midpoints — equal, to seven decimal places — and at two hundred points along each side, where it is never smaller. An ellipse whose foci were computed from a polynomial’s derivative passes through all three midpoints and touches the sides there.

Two things follow at once. The centre of the ellipse is the midpoint of the two foci, which is the average of the derivative’s roots, which is — by what the coefficients already know — the same as the average of the cubic’s roots: the centroid. And for the equilateral triangle, whose cubic can be taken as z31z^3 - 1, the derivative is 3z23z^2, with a double root at the centre, so the two foci merge and the ellipse is a circle.

Taking the derivative again, and again

Gauss–Lucas applies to the derivative as much as to the polynomial, so it can be applied repeatedly. The roots of the second derivative lie in the hull of the first derivative’s roots, which lies in the hull of the original roots, and so on down.

Six roots, and the roots of five derivatives, each nested in the last. The roots of a degree-six polynomial and of each of its successive derivatives, with the convex hull of each set shaded. Each hull sits inside the previous one, closing down on the average of the original roots.
Fig. 5 The six roots of p (orange) and the roots of p′, p″, p‴ and p⁗ in turn, each set with its convex hull shaded. The hull areas run 8.13, 3.16, 1.39 and 0.26, then a segment of length 0.43, each inside the one before. The fifth derivative is linear, and its one root (black) is the average of the six roots.

The hulls shrink, each nested in the last, and they close down on a single point. For a polynomial of degree nn the (n1)(n-1)-th derivative is linear, with one root, and that root is the average of the original roots. The reason is the one that fixed the ellipse’s centre. Taking a derivative multiplies each coefficient by its exponent and lowers the degree by one, and the ratio of the top two coefficients — which fixes the average of the roots — is preserved: zn+azn1z^n + a z^{n-1} becomes nzn1+(n1)azn2n z^{n-1} + (n-1) a z^{n-2}, whose roots average to a/n-a/n again. Every derivative’s roots have the same centre of mass, and the last one has nowhere else to be.

That nesting says something quantitative that the single theorem does not. The derivative’s roots cannot range wider than the polynomial’s, so a bound on the roots of any polynomial is automatically a bound on all its derivatives’ roots. A polynomial whose roots all lie in the unit disc has derivatives of every order whose roots lie there too.

What the push picture explains, and what it does not

The balance argument is not only a proof; it is a way to predict where the derivative’s roots will be. Near a tight cluster of kk roots the pushes from the cluster dominate everything else, and the balance points near it behave as if the rest of the polynomial were far away — which places k1k - 1 roots of the derivative inside the cluster, exactly as the middle panel of the first figure shows. Far from all the roots, the pushes add up to something like nn pushes from the centre of mass, which never balances; so the derivative has no roots far from the roots, which is the theorem again, seen from outside.

What the push picture does not do is locate the derivative’s roots exactly. Knowing that they are weighted averages of the roots does not say which weights, because the weights depend on where the derivative’s root is — the equation is circular, and solving it is solving the derivative. The theorem constrains the answer without computing it, the same trade a shared root, found without finding it made with resultants: a guarantee about where something lies, bought without ever producing it.

No smaller region will do, and three roots are enough

Is the hull the best that can be said, or is it a crude bound with a sharper truth hiding inside it? In one sense it cannot be improved at all. Make a root double — put two roots of the polynomial at the same corner of the hull — and the derivative vanishes there too, since a repeated root of a polynomial is a root of its slope. So a root of the derivative can sit exactly on a corner of the hull, and no region strictly smaller than the hull, defined from the positions of the roots alone, contains the derivative’s roots for every polynomial with roots in those positions.

In another sense the bound is loose, and the looseness can be measured. For polynomials with real coefficients, J. L. W. V. Jensen showed in 1913 that the non-real roots of the derivative lie in the discs whose diameters join each conjugate pair of roots — a region much smaller than the hull when the pairs are few and close to the real line, and one that recovers Rolle’s interlacing when there are no pairs at all.

And there is a sharpening that costs nothing. Each root of the derivative is a weighted average of all nn roots, but a point inside the hull of many points is always inside the hull of just three of them — Carathéodory’s theorem in the plane, drawn in three points, however many there are. So every root of the derivative lies in some triangle whose corners are roots of the polynomial. The theorem does not say which triangle, for the same reason it does not say which weights; but it says that three of the nn roots always suffice to fence it in.

What passes down to every derivative

The practical force of the theorem is that any convex region containing the roots contains the roots of every derivative. Two regions matter more than the rest.

A disc. If every root lies in some disc, every root of every derivative does too. Estimates of where a polynomial’s roots lie — bounds in terms of the coefficients, which are cheap to compute — therefore bound the roots of all its derivatives at no extra cost.

A half-plane. A polynomial whose roots all have negative real part is called stable, because the linear system it describes damps every disturbance rather than amplifying one: each root contributes a term like ezte^{zt}, and a negative real part makes it decay. A half-plane is convex, so the derivative of a stable polynomial is stable. In the language of the slope of a single point, taking a slope never pushes a root out of the stable half, and repeated slopes never do either.

Both facts are the theorem used as a filter rather than a location: nothing is found, but a whole region of the plane is certified empty of roots of the derivative, by a picture of the polynomial’s roots and a ruler. The derivative’s roots can be counted inside any region by the winding of a loop around it, which is a computation; the hull tells where not to bother looking, which is not.

How far inside

The theorem says the derivative’s roots are inside the hull. It says nothing about whether they are near the roots of the polynomial, and that question turns out to be hard.

Suppose every root of a polynomial lies in the closed unit disc. Is every one of those roots within distance one of some root of the derivative? The polynomial zn1z^n - 1 shows that one cannot hope for less: its roots are on the unit circle, its derivative nzn1n z^{n-1} has every root at the centre, and every root is exactly one away from the nearest.

Sendov's conjecture on 2,880 sampled roots: none farther than 1. A scatter plot of random polynomials' roots in the unit disc, each placed by its distance from the centre and its distance to the nearest root of the derivative. Every point lies below the line at distance one.
Fig. 6 Roots of 360 random polynomials of degrees 4 to 12, every root in the unit disc: blue up to degree 8, where the conjecture is proved, orange from 9 to 12, where it is not. Each of the 2,880 roots is placed by how far out it sits and how far it is from the nearest root of the derivative. The largest distance is 0.581; the dashed line at 1 is reached by zn1z^n - 1.

The figure samples 2,880 roots and none comes close. Random polynomials are nowhere near the bound — the largest distance found is well under two thirds — which shows how special the extreme case is and how little random sampling can say about the conjecture. The polynomials that would matter, if there were any, are delicately arranged, and no sample would find them.

What these pictures cannot show

The theorem, as opposed to its instances. Every figure here checks the containment on particular polynomials, a few dozen in the constructed panels and a few hundred in the random scatter. The proof is the weighted-average argument, which is short and covers every polynomial at once; the figures illustrate it and add nothing to its certainty.

Where on the ellipse the midpoints are. The Marden figure marks the three points of tangency, and a reader can see that they sit on the sides. Why they are the midpoints rather than three other points of tangency is the content of the theorem and is not visible: the picture shows an ellipse touching a triangle, and the drawing would look the same for any inscribed ellipse.

The weights. The balance point is a weighted average of the roots, and the weights — one over the squared distance — are what make it work. None of the figures draws them. A root far from a balance point gets a small weight, a root near it a large one, and that asymmetry is why the derivative’s roots gather near clusters. The field of arrows shows the pushes and hides how their lengths are distributed.

Still open: every degree from nine to very large

The question in the last section is Sendov’s conjecture, posed by the Bulgarian mathematician Blagovest Sendov in 1958 and published in 1967 in a collection of research problems, where it was attributed to his colleague Ljubomir Iliev. It has been proved for polynomials of degree up to eight, by Johnny Brown and Guangping Xiang in 1999, and for all polynomials of sufficiently large degree, by Terence Tao in 2020. Tao’s argument does not say how large is sufficiently large, and the bound it implicitly contains is enormous.

For every degree from nine up to that unspecified bound, the conjecture is open. The orange points in the scatter are drawn from exactly that range, and they are evidence only in the weak sense that nothing random comes close. The extreme polynomials, if a counterexample existed, would have their roots on or near the unit circle and their derivative’s roots arranged to keep away from every one of them — a configuration no search has found, and no argument has yet ruled out in those degrees.

A convexity statement where an order used to be

Along the real line, the roots of the derivative sit between the roots, one per gap. In the plane there is no between, and what survives is a band: the derivative’s roots are balance points of pushes away from the roots, hence weighted averages of them, hence inside their hull. Applied repeatedly, the bands nest and close down on the centre of mass, which every derivative shares.

For three roots the band is a triangle and the balance points are the foci of the ellipse through the midpoints of its sides — a fact about cubics that turns out to be a fact about triangles, and that nobody finds obvious in either direction. And the question the theorem invites but does not answer, how close to the roots the balance points must lie, is settled below degree nine and above an astronomically large degree, and open in between.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

CentroidComplex numbersConvex hullDerivativeEllipseFocusPolynomialRootsWeighted average