Geometry

Equal diagonals in a curved quadrilateral

Two ellipses and two hyperbolas sharing the same foci cut out a four-sided region with curved sides and no symmetry to speak of. Its two diagonals are nevertheless exactly equal. The reason is a stretch that carries one ellipse onto the other and moves every pair of points so that crossed distances match — a property only confocal curves have.

Worth reading first: Two families that cross at right angles · Six points on a conic, and the line they share.

Fix two points and draw every ellipse that has them as foci, then every hyperbola. Each ellipse of the first family crosses each hyperbola of the second at a right angle, so the two families together are a coordinate system — curved, but as good as the grid of horizontal and vertical lines for naming a point.

In the ordinary grid, the rectangle between two horizontal lines and two vertical ones has equal diagonals. Nothing suggests the curved grid should do the same.

Two confocal ellipses, two confocal hyperbolas, and equal diagonals of 1.399. Two ellipses and two hyperbolas with the same foci cut out a four-sided region with curved sides. Its two diagonals, drawn as straight segments, both measure 1.3987.
Fig. 1 Two ellipses and two hyperbolas with foci at ±1.2. In the upper right they cut out a region with four curved sides, shaded. Its two diagonals, drawn straight from corner to opposite corner, both measure 1.3987 — equal to twelve decimal places. With the outer ellipse and hyperbola moved to different foci, the diagonals differ by 0.135.

The two diagonals of any quadrilateral cut out by two confocal ellipses and two confocal hyperbolas are equal. That is Ivory’s theorem, from 1809. The region has no symmetry that exchanges its diagonals — its sides are four different curves of two different kinds — and the equality holds for every choice of the four curves.

A rectangle grid, bent

The theorem makes more sense seen as a statement about a coordinate grid, and the grid makes more sense in its two limits.

Far from the foci, the ellipses become nearly circles and the hyperbolas nearly straight lines through the centre, so the confocal grid looks like polar coordinates. The region between two circles and two lines through their centre has four corners that form an isosceles trapezoid, and an isosceles trapezoid has equal diagonals by reflection in its axis. As the two foci merge, Ivory’s theorem becomes that symmetry.

Near the foci nothing of the kind is visible. The hyperbolas bend sharply round the foci, the ellipses flatten, and a region between two of each is lopsided with sides of very different curvature. The equality of the diagonals survives, but whatever reflection explained it in the polar limit is gone, and the explanation has to be something else.

Two families of conics, crossing at right angles. 2 ellipses and 2 hyperbolas with the same pair of foci. Every ellipse meets every hyperbola at a right angle, checked at all 4 crossings.
Fig. 2 The same two ellipses and two hyperbolas drawn as a coordinate grid, with every crossing marked. At each of the crossings the two curves meet at a right angle, which the figure checks from their tangent directions. Every small region of this grid, bounded by an ellipse and a hyperbola on each side, has two equal diagonals.

The stretch that carries one ellipse onto the other

The explanation is a map. Take the inner ellipse, with semi-axes a1a_1 and b1b_1, and the outer one, with a2a_2 and b2b_2. Stretch the plane horizontally by a2/a1a_2/a_1 and vertically by b2/b1b_2/b_1. That is a linear map of the kind that turns a circle into an ellipse, with the two axes as the directions it leaves alone, and it carries the inner ellipse exactly onto the outer.

It also does something less expected: it moves each point of the inner ellipse along the confocal hyperbola through that point. The corner where the inner ellipse meets a hyperbola goes to the corner where the outer ellipse meets the same hyperbola. The figures check that by measuring the difference of each image’s distances to the two foci — the quantity a hyperbola holds constant — and requiring it to equal the original’s.

Crossed distances across a stretch between confocal ellipses, both 3.119. Two confocal ellipses, two points on the inner one and their images under the stretch onto the outer one. The two crossed segments both measure 3.1186.
Fig. 3 An inner ellipse stretched onto a confocal outer one: horizontally by 1.516 and vertically by 2.060. X and Y are two points on the inner curve, and X′ and Y′ are where the stretch takes them, each along its own confocal hyperbola, drawn faint. The crossed segments X–Y′ and X′–Y both measure 3.1186, equal to twelve places, and the same equality holds for twenty-four other pairs.

That is Ivory’s lemma, and it is stronger than the theorem. For any two points X and Y on the inner ellipse, the distance from X to the image of Y equals the distance from the image of X to Y. The corners of the curved quadrilateral are a special case. The inner corners are X and Y; their images are the outer corners, each on the hyperbola its original shared; and the two diagonals are exactly the two crossed segments.

Why the foci are what make it work

The lemma has a proof of one line, and the line uses the shared foci and nothing else.

Write X=(a1coss, b1sins)X = (a_1\cos s,\ b_1 \sin s) and Y=(a1cost, b1sint)Y = (a_1 \cos t,\ b_1 \sin t). The stretch sends them to (a2coss, b2sins)(a_2 \cos s,\ b_2 \sin s) and (a2cost, b2sint)(a_2 \cos t,\ b_2 \sin t). Square the two crossed distances and subtract:

XY2XY2=(a12a22)(cos2scos2t)+(b12b22)(sin2ssin2t).|XY'|^2 - |X'Y|^2 = (a_1^2 - a_2^2)(\cos^2 s - \cos^2 t) + (b_1^2 - b_2^2)(\sin^2 s - \sin^2 t).

For confocal ellipses the focal distance cc is the same, and a2b2=c2a^2 - b^2 = c^2 for both, so a22a12=b22b12a_2^2 - a_1^2 = b_2^2 - b_1^2. Call that common value λ\lambda. The difference becomes λ[(cos2s+sin2s)(cos2t+sin2t)]-\lambda\big[(\cos^2 s + \sin^2 s) - (\cos^2 t + \sin^2 t)\big], which is λ(11)=0-\lambda(1 - 1) = 0.

The whole theorem is the identity cos2+sin2=1\cos^2 + \sin^2 = 1, used once for each point, and made available by the two ellipses growing by the same amount in both directions. Take away the shared foci and the two differences a22a12a_2^2 - a_1^2 and b22b12b_2^2 - b_1^2 are no longer equal, the bracket does not collapse, and the crossed distances differ.

The same stretch without shared foci, and crossed distances that differ. An inner ellipse and a scaled copy of it with different foci, two points and their images under the stretch. The crossed segments measure 3.0889 and 3.0421.
Fig. 4 The same stretch onto an outer ellipse of the same shape — scaled by 1.516 in both directions — whose foci are not the inner ellipse’s. The crossed segments now measure 3.0889 and 3.0421. Across twenty-four other pairs of points the largest difference between the two crossed distances is 0.467.

The control is not a small perturbation. A similar ellipse — same shape, bigger — is the most natural outer ellipse to reach for, and it is exactly the wrong one: its axes grow in proportion, which keeps b/ab/a fixed and changes cc. Confocal ellipses do the opposite, keeping cc fixed and changing the shape, and only they satisfy the lemma.

The same equality, read off the coordinates

The stretch is one way to see the theorem. The coordinates the two families provide are another, and they explain it without any map at all.

Name a point by the ellipse and the hyperbola through it: semi-axis aa for the ellipse and α\alpha for the hyperbola, with a>c>αa > c > \alpha. In the upper right its position is x=aα/cx = a\alpha/c and y=a2c2c2α2/cy = \sqrt{a^2 - c^2}\,\sqrt{c^2 - \alpha^2}/c, which is the formula the figures use to place every corner. Write bb for a2c2\sqrt{a^2 - c^2} and β\beta for c2α2\sqrt{c^2 - \alpha^2}.

Now take two points, (a1,α1)(a_1, \alpha_1) and (a2,α2)(a_2, \alpha_2), and expand the square of the distance between them. One combination appears twice and simplifies: a2α2+b2β2=c2(a2+α2c2)a^2\alpha^2 + b^2\beta^2 = c^2(a^2 + \alpha^2 - c^2). With it the squared distance is

a12+a22+α12+α222c22c2(a1a2α1α2+b1b2β1β2).a_1^2 + a_2^2 + \alpha_1^2 + \alpha_2^2 - 2c^2 - \frac{2}{c^2}\big(a_1 a_2\,\alpha_1 \alpha_2 + b_1 b_2\,\beta_1 \beta_2\big).

Every term is unchanged when the two hyperbola coordinates are swapped. The sum of squares does not care which α\alpha is which, and each product contains both αs, or both βs, once. And swapping them is exactly the passage from one diagonal to the other: the first diagonal joins (a1,α1)(a_1, \alpha_1) to (a2,α2)(a_2, \alpha_2), the second joins (a1,α2)(a_1, \alpha_2) to (a2,α1)(a_2, \alpha_1). They have the same length because the distance formula, written in these coordinates, cannot tell which ellipse a hyperbola coordinate was paired with.

The ordinary grid has the same property for the same reason. In (x1x2)2+(y1y2)2(x_1 - x_2)^2 + (y_1 - y_2)^2, swapping y1y_1 and y2y_2 changes only a sign inside a square, and that is the whole of why a rectangle’s diagonals are equal. Ivory’s theorem says the confocal grid shares that one property of the Cartesian grid, while sharing almost nothing else — its lines are curves, its cells are not congruent, and its spacing changes from place to place.

The curves at the edge of the family

The confocal family has members that are not curves in the ordinary sense, and the coordinate formula reaches them too.

As an ellipse’s semi-axis aa shrinks towards cc, its minor axis shrinks to nought and it collapses onto the segment joining the two foci. As a hyperbola’s semi-axis α\alpha shrinks to nought it straightens into the vertical line midway between the foci, and as α\alpha grows to cc it closes up onto the two rays of the axis beyond the foci. The position formula x=aα/cx = a\alpha/c, y=bβ/cy = b\beta/c makes sense at all of these values, and so does the distance formula of the previous section, which was never asked to avoid them.

So Ivory’s equality holds for figures that do not look like quadrilaterals. Put the inner ellipse at a=ca = c: its “corners” are the points (α1,0)(\alpha_1, 0) and (α2,0)(\alpha_2, 0) on the segment between the foci, and the theorem says that the distance from (α1,0)(\alpha_1, 0) to the point of the outer ellipse on the hyperbola α2\alpha_2 equals the distance from (α2,0)(\alpha_2, 0) to the point of the same ellipse on the hyperbola α1\alpha_1. Substituting into the formula gives both squared distances as a2+α12+α22c22aα1α2/ca^2 + \alpha_1^2 + \alpha_2^2 - c^2 - 2a\alpha_1\alpha_2/c, visibly unchanged by the swap.

Push one of those corners all the way to a focus, α1=c\alpha_1 = c, and the formula collapses further, to (aα2)2(a - \alpha_2)^2. The distance from a focus to the point of an ellipse on the hyperbola α\alpha is aαa - \alpha, and from the other focus it is a+αa + \alpha. Their sum is 2a2a, which is the string of the gardener’s construction; their difference is 2α2\alpha, which is the constant defining the hyperbola through the point. The focal definitions of both curves are sitting inside the distance formula, as the special case where one corner is a focus.

That is a good test of a formula that claims to explain a theorem: it should reproduce the facts the theorem grew from when pushed to the edge of its range. This one reproduces both focal definitions, and it gives Ivory’s equality for every configuration in between.

The theorem at another size

Two confocal ellipses, two confocal hyperbolas, and equal diagonals of 1.235. Two ellipses and two hyperbolas with the same foci cut out a four-sided region with curved sides. Its two diagonals, drawn as straight segments, both measure 1.2349.
Fig. 5 A different confocal quadrilateral, with foci at ±1, ellipses of semi-axis 1.3 and 2 and hyperbolas of semi-axis 0.3 and 0.8. The region is larger and more lopsided, and its diagonals both measure 1.2349. Moving the outer curves to different foci makes them differ by 0.107.

Every quadrilateral of the grid has the property, and they are not similar to one another. What they share is not a shape but a construction: the corners are pairs of points related by a stretch between confocal ellipses, and the lemma says every such pair of pairs has equal crossed distances. The equal diagonals are not a coincidence of these four curves; they are a property of the stretch, and every curved quadrilateral in the grid is one instance of it.

A string wrapped round an ellipse

The confocal family has another property of the same surprising kind, and it grows out of the gardener’s construction of an ellipse. That construction ties a loop of string round two pins and pulls it taut with a pencil, and the pencil traces an ellipse with the pins as foci, because the string holds the sum of the two distances fixed.

Replace the two pins by an ellipse. Wrap a closed loop of string, longer than the ellipse’s perimeter, around it, pull it taut with a pencil, and draw. The pencil traces a closed curve, and Charles Graves proved in 1841 that the curve is an ellipse confocal with the one the string is wrapped around. The two pins were the degenerate member of the same family: as a confocal ellipse’s minor axis shrinks to nought it collapses onto the segment between the foci, and a string round that segment is a string round two pins.

Ivory’s lemma and Graves’s theorem are both statements that something is preserved as a point moves between confocal curves, and neither has a proof by symmetry. Together with the right angles at which the two families cross, they are what make the confocal grid the natural setting for a ball in an elliptical table, whose every trajectory stays tangent to one confocal curve for ever — the property that makes an elliptical room impossible to light from some places and possible from others.

The grid is also the one the discriminant’s classification cannot see into. Every curve in it is an ellipse or a hyperbola, the sign of B24ACB^2 - 4AC says which, and nothing about that sign records the shared foci that all of these properties depend on. A classification by type throws away exactly the information a theorem like Ivory’s needs, which is why the confocal family has to be described by its foci rather than by its coefficients.

Where it came from, and where it goes

James Ivory proved the lemma in 1809, and the problem he needed it for was the gravitational attraction of a solid ellipsoid on a point outside it. The attraction of a homogeneous ellipsoid on a point inside or on its surface had been known since Maclaurin. Ivory’s lemma, in its three-dimensional form for confocal ellipsoids, lets the external problem be reduced to an internal one: the attraction on an outside point is related to the attraction of a confocal ellipsoid passing through that point, and the correspondence between the two surfaces is exactly the stretch.

The three-dimensional version holds for the same reason as the plane one. Confocal quadrics in space — ellipsoids and the two kinds of hyperboloid sharing their focal conics — form a coordinate system in which three surfaces meet at right angles at every point, and the stretch between two confocal ellipsoids preserves crossed distances by the same cancellation, with one more cos2+sin2\cos^2 + \sin^2 in it. Versions on the sphere and in the hyperbolic plane have been proved too, with the same structure behind them.

The lemma is also one of the standard tools for the elliptic billiard, where the confocal conics are the curves a trajectory keeps touching and never crosses. A ball bouncing inside an ellipse moves along chords tangent to one fixed confocal ellipse or hyperbola, and results about when such trajectories close up are proved by transporting chords from one confocal curve to another — which is what the stretch does.

Where the theorem needs its hypotheses

The four curves must share both foci. That is the whole content of the proof, and the controls show the equality failing as soon as it is dropped — by 0.135, 0.107 and 0.467 in the three figures that test it.

The diagonals are straight segments. The theorem is about the Euclidean distance between opposite corners, not about the length of any path along the grid. Measured along the curved sides, the two routes from a corner to the opposite corner have no reason to agree and do not.

The corners must pair up correctly. Each diagonal joins a corner on the inner ellipse to a corner on the outer ellipse and on the other hyperbola. The two segments joining corners on the same ellipse, or on the same hyperbola, are sides of a different kind and are not equal in general.

And the region may be anywhere in the plane. The figures draw it in the upper right for clarity, but the grid is symmetric under reflection in both axes and the theorem holds for a quadrilateral in any quadrant, or straddling an axis, as long as its four curves are two confocal ellipses and two confocal hyperbolas.

Twelve decimal places, and a proof that is one line of algebra

The diagonals are measured to twelve decimal places at two quadrilaterals and the lemma is checked at twenty-five pairs of points. That is strong evidence and not a proof; the proof is the line of algebra above, and the figures exist to show that the algebra describes what is drawn.

They also cannot show the three-dimensional version or the gravitational problem it was built for, and they do not draw the stretch as a motion. A reader sees the before and the after, with dashed lines joining each point to its image; that the image slides along a hyperbola is checked numerically and drawn as a faint curve, not animated.

Still open: whether only an ellipse has a family of caustics

The confocal conics are the curves a billiard trajectory in an ellipse keeps touching, and every trajectory touches one of them. A table with that property — every trajectory near the boundary tangent to one of a family of curves filling a region — is called integrable, and the ellipse and its degenerate case the circle are the only known examples.

George Birkhoff asked in the 1920s whether they are the only ones: whether every convex table whose billiard is integrable is an ellipse. The question is still open. Vadim Kaloshin and Alfonso Sorrentino proved in 2018 that a table close enough to an ellipse and integrable in a strong local sense must be an ellipse, and there are further partial results under extra symmetry assumptions, but the full conjecture has not been proved or refuted. If it is true, the structure behind Ivory’s theorem — curves that share foci and a stretch between them — is not one example among many but the only way a billiard table can be that orderly.

The quantity that nothing seemed to preserve

The habit is about looking for the map before looking for the symmetry.

The curved quadrilateral has no symmetry that swaps its diagonals, and a search for one would fail. The equality is explained instead by a transformation that is not a symmetry of the figure at all — a stretch that moves the whole inner ellipse outward — and by a quantity it preserves that is not a length of anything in the figure: the distance between a point and the image of another point. The diagonals are equal because they are two instances of that quantity, and the quadrilateral is simply where two such instances happen to be drawn.

When two things in a figure are equal and no symmetry explains it, it is worth asking whether both are values of a single expression under some map that the figure does not show — and then checking what that map needs, which here was one shared pair of foci and nothing else.