Geometry

One sign decides which curve

The general quadratic in two variables has six coefficients and draws a conic. Which of the four it draws is settled by a single combination of three of them, and the other three cannot change the answer however they are chosen.

Worth reading first: One cone, four curves · Two families that cross at right angles.

The curves on this ladder have been described three ways — by cutting a cone, by fixing two distances, and by fixing a distance and a line. There is a fourth, and it is the one a coordinate calculation produces:

Ax2+Bxy+Cy2+Dx+Ey+F=0.Ax^2 + Bxy + Cy^2 + Dx + Ey + F = 0.

Six numbers. Every conic is the solution set of such an equation, every such equation draws a conic or something degenerate, and which of the four it draws is decided by three of the six.

One sign decides which curve it is. 3 conics drawn from the general quadratic, each labelled with its discriminant B² − 4AC and the curve that sign names, checked against how many times the curve meets a large circle.
Fig. 1 Three curves from the same six-coefficient form. Under each is B24ACB^2 - 4AC, and the sign of that single number names the curve — checked here against an independent determination, which counts how many times the curve crosses a circle of radius thirty.

Below nought an ellipse, at nought a parabola, above nought a hyperbola. The three coefficients DD, EE and FF do not appear, and they cannot: they move the curve and change its size, and they never change its kind.

Why only the quadratic part matters

The reason is a statement about what happens far away.

For large xx and yy, the terms Ax2+Bxy+Cy2Ax^2 + Bxy + Cy^2 dominate Dx+Ey+FDx + Ey + F — one is quadratic in the size and the other linear. So the curve’s behaviour at a great distance is governed by the quadratic part alone, and the type of a conic is a statement about that behaviour: an ellipse is bounded, a parabola runs off in one direction, a hyperbola in two.

Making that precise: divide the whole equation by the square of the distance from the origin and let the distance grow. What survives is Acos2θ+Bcosθsinθ+Csin2θ=0A\cos^2\theta + B\cos\theta\sin\theta + C\sin^2\theta = 0, an equation for the directions in which the curve escapes to infinity. Multiply through by sec2θ\sec^2\theta and it becomes a quadratic in tanθ\tan\theta with discriminant B24ACB^2 - 4AC — and the number of directions is two, one or none according to its sign.

Two escape directions are a hyperbola’s asymptotes, one is a parabola’s axis, and none is an ellipse. That is the classification, and the discriminant appears as the discriminant of a quadratic in exactly the way it always does.

The last clause is worth pausing on, because the name is not a coincidence and is often taken for one. b24acb^2 - 4ac is the quantity deciding whether ax2+bx+c=0ax^2+bx+c = 0 has two real roots, one, or none — and here the quadratic whose roots are being counted is the one in tanθ\tan\theta, whose roots are the directions of escape. The conic’s discriminant is a root count, and the three cases are two roots, a repeated root, and no real roots, in that order.

That also says why the parabola is the boundary rather than a third species. A repeated root is what separates two real roots from none, and it happens on a set of measure zero in the coefficients — so a conic chosen at random is an ellipse or a hyperbola, and never a parabola. The parabola is the case that requires an exact equality, which is the same reason its eccentricity is exactly one and its cutting plane exactly parallel to the cone’s side.

The second determination

The figures do not take the sign’s word for it. For each curve they count how many times it crosses a large circle, by sampling the quadratic’s value all the way round and counting sign changes.

An ellipse is bounded, so a circle of radius thirty misses it entirely: nought crossings. A parabola has one branch escaping in one direction and coming back, so it meets the circle twice. A hyperbola has two such branches and meets it four times.

That count is a fact about the drawing and the discriminant is a fact about three coefficients, and the assertion is that the two agree for every curve drawn. A sign error in the classification would be caught, and so would a curve mislabelled in the figure’s own list.

One sign decides which curve it is. 2 conics drawn from the general quadratic, each labelled with its discriminant B² − 4AC and the curve that sign names, checked against how many times the curve meets a large circle.
Fig. 2 Two more, with linear terms that shift and tilt them. The discriminants are 44-44 and 1717; the curves are displaced from the origin and their kinds are unchanged, which is the previous section drawn.

The rotation that removes the cross term

The classification has a second reading, and it is the one that connects to the rest of the algebra on this site.

Write the quadratic part as a matrix product:

(xy)(AB/2B/2C)(xy).\begin{pmatrix} x & y \end{pmatrix} \begin{pmatrix} A & B/2 \\ B/2 & C \end{pmatrix} \begin{pmatrix} x \\ y \end{pmatrix}.

That matrix is symmetric, so it has perpendicular eigenvectorsdirections the map leaves alone — and rotating the coordinates to line up with them removes the xyxy term entirely. In the new coordinates the equation reads λ1x2+λ2y2+=0\lambda_1 x'^2 + \lambda_2 y'^2 + \cdots = 0 with the two eigenvalues as coefficients.

Now the classification is obvious. Two eigenvalues of the same sign make a bounded curve, an ellipse. Opposite signs make a hyperbola. One eigenvalue nought makes a parabola, since one squared term has vanished.

And the connection to the discriminant is the determinant: λ1λ2=ACB2/4=(B24AC)/4\lambda_1\lambda_2 = AC - B^2/4 = -(B^2-4AC)/4. So the discriminant’s sign is the opposite of the determinant’s, and the three cases are the three sign patterns of a pair of eigenvalues. The classification of conics is the classification of symmetric two-by-two matrices, and the two subjects are one.

The directions the map leaves alone. Unit vectors and their images under the map. On the two marked lines the image points the same way as the original, stretched by 3.62 and 1.38.
Fig. 3 The eigen-directions of a symmetric matrix, where the map leaves a direction alone. Rotating a conic’s coordinates onto these two axes is what removes its xyxy term, and the two eigenvalues that survive are what the discriminant is measuring the product of.

What the three coefficients see

The quadratic part is a quadratic form, and reading the classification as a statement about that object rather than about a curve says what is really being classified.

A quadratic form takes a direction and returns a number: how much the form grows along that direction. Sweep the direction round a circle and the number traces out a curve, and the three cases are the three things that curve can do.

If it is positive in every direction, the form grows whatever way the point moves, so the level set =1= 1 is a closed curve at a bounded distance in every direction. That is the ellipse, and both eigenvalues are positive.

If it is negative in some directions and positive in others, there are directions along which the form is nought — the boundary between the two behaviours — and along those directions the level set never closes. Two such directions, two asymptotes, a hyperbola.

And if it is nought in one direction and positive elsewhere, the form is degenerate: one eigenvalue is nought, one direction feels nothing quadratic at all, and along it the linear term takes over. That is the parabola, and it is the knife-edge case in the exact sense that it separates the other two.

So the classification is a statement about which values a form takes, and the technical name for that is its signature — the number of positive and negative eigenvalues. What a map does to a circle is the same reading in a different vocabulary: a symmetric map takes the unit circle to an ellipse whose axes are the eigen-directions, and the signature says whether that image is an honest ellipse, a degenerate one, or a hyperbola-shaped set of directions.

That reading also explains a fact that the coefficient version makes look accidental. The classification is unchanged by any invertible change of coordinates that preserves lengths — which is what a rotation is — and Sylvester’s law of inertia says it is unchanged by very much more: any invertible linear change at all preserves the signs of the eigenvalues, though not their values. So the type survives being stretched, sheared and skewed, and the discriminant is a shadow of a much more robust invariant.

The degenerate cases

The classification above has three outcomes and the equation has more, because a quadratic can fail to draw a curve at all.

x2y2=0x^2 - y^2 = 0 has discriminant 44, which says hyperbola, and its solution set is the two lines y=±xy = \pm x — a degenerate hyperbola, which is what a plane through a cone’s apex cuts. x2+y2=0x^2 + y^2 = 0 has discriminant 4-4, says ellipse, and its solution set is the single point at the origin. x2+y2+1=0x^2 + y^2 + 1 = 0 says ellipse and has no real solutions at all.

So the discriminant classifies the type and does not say whether the curve is non-degenerate, and a second quantity is needed for that: the determinant of the three-by-three matrix

(AB/2D/2B/2CE/2D/2E/2F),\begin{pmatrix} A & B/2 & D/2 \\ B/2 & C & E/2 \\ D/2 & E/2 & F \end{pmatrix},

which is nought exactly in the degenerate cases. Two determinants classify completely: the two-by-two one names the type and the three-by-three one says whether it is a genuine curve.

That the degenerate cases are cones cut through the apex is the geometric account of both. A plane through the apex meets the double cone in a point, a line or two lines depending on its tilt — which is the three degenerate outcomes, in the same order as the three types.

The empty case is the one the cone picture cannot supply, and it is worth saying where it comes from instead. x2+y2+1=0x^2 + y^2 + 1 = 0 has no real solutions and it does have complex ones, and in the complex plane it is a perfectly ordinary conic — an ellipse whose points all have imaginary coordinates. The three-by-three determinant is non-zero for it, correctly reporting a non-degenerate curve, and the emptiness is a fact about which of that curve’s points are real.

A classification over the reals has an outcome the geometry has no picture of, and that outcome is not a defect of the algebra but a consequence of asking a question about real solutions of an equation whose natural home is the complex numbers. It is the same situation as a quadratic with no real roots, one field away, and it is resolved the same way.

How many conics there are through given points

One consequence of the six coefficients is worth extracting, because it is the fact everything projective about conics starts from.

The equation has six coefficients and multiplying all of them by a constant gives the same curve, so a conic is determined by five numbers rather than six. Each point the curve is required to pass through imposes one linear condition on the coefficients. So five points determine a conic, generically, and any four points leave a one-parameter family.

That family is called a pencil, and the confocal family of the rung below is one — its parameter λ\lambda is the pencil’s parameter, and the two conics through a given point are the two members of the pencil that pass through it. What looked there like a quadratic equation with two roots is here the general statement that a line meets a pencil’s base locus twice.

The counting also explains why the classification is coarse. Five numbers describe a conic and one sign describes its type, so the type is throwing away four and a bit dimensions of information — which is what a classification is for, and is why two curves of the same type can look nothing alike.

Four conic sections from one cone. Circle, ellipse, parabola and hyperbola, produced by tilting a single cutting plane further and further.
Fig. 4 The four non-degenerate curves from the cutting picture. Every one of them is the solution set of a six-coefficient equation, and moving the plane through the apex collapses whichever curve was there into its degenerate version.

What the classification is worth

A single sign settling a question about a curve is useful in a specific way that is worth stating, because it is what such invariants are for.

It is computable without drawing. Three coefficients, two multiplications and a subtraction, and the answer is exact. Determining the type by plotting is a sampling procedure with all the usual hazards; determining it by arithmetic is not.

It is unchanged by moving the curve. Translating shifts DD, EE and FF and leaves AA, BB and CC alone, so the discriminant is a translation invariant. Rotating changes all three and leaves B24ACB^2-4AC unchanged, which is a small calculation and is the statement that the eigenvalues do not depend on the coordinates.

And it extends. In three variables the quadratic form is a symmetric three-by-three matrix, the classification is by the signs of three eigenvalues, and the resulting surfaces are the ellipsoids, the two kinds of hyperboloid, the two kinds of paraboloid and the cones. The same argument, one dimension up, with more sign patterns.

Three curves, one rule, one number changed. A focus and a directrix, and the curves of points whose distance to the focus is a fixed multiple of their distance to the line. Below one the curve closes, at one it is a parabola, above one it has two branches.
Fig. 5 The same three types from a completely different description — one focus, one line, one ratio. Two classifications of one family, and the correspondence between them is that the eccentricity crossing one is the discriminant crossing nought.

Where the account needs care

The discriminant’s sign convention is not universal. Some sources use ACB2/4AC - B^2/4, the determinant, whose sign is the opposite. The three cases are the same three cases either way and a reader crossing between conventions has to check which is meant.

The equation must be genuinely quadratic. If AA, BB and CC are all nought the equation is linear and draws a line, which is not a conic at all and which the discriminant reports as a parabola. The figures guard against it by requiring a quadratic term.

The circle-crossing test is a test on a drawing. Counting sign changes round a large circle is a numerical procedure and it is only reliable when the radius is large enough that the linear terms are negligible; the figures use thirty against coefficients of size a few, which is comfortable, and would be wrong for coefficients of very different magnitudes.

And a circle is an ellipse. The classification has three outcomes and the ladder’s opening essay had four curves, because the circle is the special ellipse with A=CA = C and B=0B = 0. Nothing in the discriminant distinguishes it, and nothing needs to.

Where it came from

The general equation and its classification are Euler’s, in the Introductio in analysin infinitorum of 1748, and the setting explains the shape of the treatment: Euler was building analytic geometry as a subject, and the classification of second-degree curves is its first real theorem.

What is worth noticing is how late that is. Apollonius had the curves in the third century BC, described by their sections and their focal properties; Descartes had coordinates in 1637; and the statement that one sign of one combination decides the type had to wait a further century. A description in coordinates does not immediately produce the questions coordinates are good at, and the discriminant is a question that only makes sense once the equation is the primary object.

The eigenvalue reading is later still — Cauchy in the 1820s, as part of the work on quadratic forms that made the classification a statement about matrices rather than about curves. That is the version that generalises, and it is the version that says why the classification is a classification of something rather than a list of three cases.

What the pictures cannot show

Each curve is traced by following the zero level of the quadratic over a grid of sixty-eight thousand cells, so what is drawn is a polyline approximating a curve, and its smoothness is a matter of resolution.

The window is a few units wide and the classification is about behaviour at infinity. A hyperbola looks like two curves within the window because it is two curves everywhere; an ellipse looks bounded because it is; but a parabola and a very eccentric hyperbola are indistinguishable in any finite window, and the figure’s verdict comes from the arithmetic and the circle count rather than from the drawing.

That is the honest description of what the classification is for. The type is a statement about the whole plane and every picture is a window, and the discriminant exists precisely because looking is not a method.

The ladder from here

Rungs above: the projective classification, where a change of coordinates that mixes in the constant term collapses the three cases to one and the whole distinction turns out to be about where the curve meets the line at infinity. Five points determining a conic, and Pascal’s theorem about the hexagon inscribed in one. The caustic, which is the envelope of rays a mirror of the wrong shape produces. Quadratic forms in nn variables and Sylvester’s law of inertia, which says the signs of the eigenvalues are all that survives a change of coordinates. And the pencil of conics through four points, which is the one-parameter family the confocal rung’s λ\lambda was an instance of.

Three of six, and the other three do nothing

The habit is about reading a classification for what it does not use.

Six coefficients go in and three of them are ignored. That is not an economy of the method; it is the content of the theorem — the statement that translating and moving a conic cannot change what it is, which is obvious geometrically and needs proving algebraically, and the proof is exactly the observation that DD, EE and FF do not appear.

A quantity’s arguments are a statement about what it is invariant under, and reading them off is the quickest way to know what a classification can and cannot see. The discriminant sees three coefficients, so it is blind to position, size and orientation — and correspondingly it cannot tell a small ellipse from a large one, which is right, because they are the same kind of curve.

The corollary is the test to apply to any proposed invariant: list what it does not depend on, and check that the list is exactly the transformations it is supposed to survive. An invariant depending on too much fails to be invariant; one depending on too little classifies nothing.