The square that cannot be negative
Worth reading first: The dot product is a shadow.
The dot product is a shadow, and a shadow is never longer than the thing casting it. That sentence is the geometry, it is obviously true in the plane, and it is worth nothing at all in a space of functions where there is no light, no floor and nothing to cast.
What replaces it is a quadratic.
That is the whole proof. The rest of this essay is about which hypothesis each part of it spends, because the argument is short enough that every step is visible and the shortness is what makes it portable.
The one fact the proof uses
Take any two vectors and , with not zero, and build a third from them: , where is a number free to be anything. As moves, the tip of slides along a line, and the vector’s length changes.
Its squared length is
which is a quadratic in — expanded using nothing but that the product is symmetric and linear in each slot.
Now the single fact. The left-hand side is a squared length, so it is never negative, for any whatever. The right-hand side is therefore a quadratic that never goes below zero, and a quadratic with that never goes below zero cannot have two distinct real roots — if it did, it would be negative between them. So its discriminant is at most zero:
Nothing else was used. Not the angle, which is the thing being defined; not coordinates, which the sum-of-products form would need; not the dimension, which is nowhere in the argument. Three properties of the product went in — symmetry, linearity in each argument, and that a non-zero vector has positive squared length — and the inequality came out.
Where the vertex is, and what it says
The parabola’s least value sits at
which is a number that has already appeared in this ladder. It is exactly the multiple of that the shadow of reaches. So the quadratic argument and the shadow argument are not two proofs; they are the same proof, one of them drawn.
The value at the vertex is
and reading it as a length rather than as an expression says something the inequality alone does not. It is the squared length of what is left of once its shadow on has been taken away — the part of that is perpendicular to . Cauchy–Schwarz is the statement that this leftover has a length, and the difference between the two sides of the inequality is exactly how much of points somewhere does not.
That is a more informative reading than “the product is bounded”. A bound says a quantity does not exceed something; this says what the gap is made of, and the gap turns out to be a vector, sitting in the picture, with its own length.
It also explains why the two sides of the inequality are squared. The natural statement is about the leftover, which is a length and therefore a square root of something; squaring both sides of clears the root and leaves an identity between quantities the product can produce directly. Written with the leftover named, the inequality is not an inequality at all but an equation,
whose right-hand side is a product of two things that cannot be negative. Everything else follows from reading that equation from left to right.
Equality means parallel, and the picture shows why
An inequality is only worth having if the case of equality is known, since that is where it is tight and where it is being used.
Equality holds exactly when the parabola touches the axis, which happens exactly when is the zero vector for some , which happens exactly when is a multiple of . The generator checks that as an if-and-only-if rather than in one direction — the interesting half is that a strict inequality means the two vectors genuinely point different ways, and a check that only confirmed the easy direction would pass on a figure where they did not.
Read through the vertex, the equality case says the leftover is zero, which is the same sentence: there is nothing of pointing anywhere does not.
The other extreme, where the gap is as wide as it goes
At the other end of the range the vertex sits over . Subtracting any amount of at all lengthens the result, so the nearest point of ’s line to the tip of is the origin, and .
This is the case the inequality is least often quoted for and it is the one that makes its shape clear. The two sides of are equal when the vectors are parallel and maximally apart when they are perpendicular, and the ratio between them runs continuously over in between. That ratio is what gets promoted to the definition of the cosine, and the inequality is precisely the guarantee that the promotion is legal — a ratio outside names no angle.
What the inequality is spent on
Three standard results are the same line of algebra with this inequality substituted in, and it is worth seeing all three at once because they are usually met years apart.
The triangle inequality. Expand , replace the middle term by its bound , and the right-hand side becomes . So : the direct route is never longer than a route through a third point. The whole content of “a straight line is the shortest path” between two points, in this setting, is Cauchy–Schwarz applied once.
The angle, and its cosine. Dividing by puts the ratio in and the arccosine of it is defined. In two dimensions this is a computation of an angle that exists; in a thousand dimensions, or in a space of functions, it is the definition of one, and the inequality is what licenses it.
The correlation coefficient. Take the vectors to be two lists of measurements with their averages subtracted. The dot product is then the covariance, the lengths are the standard deviations, and Cauchy–Schwarz says the correlation lies between and — with equality exactly when one list is a multiple of the other, which is to say exactly when the measurements lie on a line. The bound everybody knows for correlation and the bound everybody knows for shadows are one statement in two vocabularies, and how far from the average a thing can be is the neighbouring inequality in that language.
The pattern in all three is the same. The inequality is never the goal; it is the step where an unknown cross-term is replaced by something known, and everything after that is bookkeeping.
The proof does not care what the vectors are
Nothing above mentioned coordinates, dimension, or arrows, and that is not an accident of presentation — it is the reason this proof is the one that gets taught.
Take the “vectors” to be functions on an interval, with
The three properties hold: the integral of is the integral of ; it is linear in each slot; and is positive unless is zero. So for every , the same quadratic appears with the same coefficients, and
That is a real theorem about integrals, proved by drawing a parabola. It is the inequality that makes Fourier coefficients behave — the coefficient of each harmonic is a shadow, and this is what stops a shadow being longer than what casts it — and it is the reason a space of functions can be said to have angles in it at all.
The same substitution with sums rather than integrals gives , which is the form Cauchy proved and which is a fact about any two lists of numbers whatever.
Any rule with the three properties
The generalisation is worth taking one step further, because it shows exactly how little the proof requires.
Weight the coordinates unequally — with both weights positive — and everything survives. Symmetry is obvious, linearity is obvious, and positivity holds because a sum of positive multiples of squares is positive. So the parabola argument runs unchanged and the inequality holds for the new rule.
What changes is the geometry the rule describes. The vectors of length one are now an ellipse; the pairs the rule calls perpendicular are not the pairs that look perpendicular; and the shadow of one vector on another is computed with the weights in it. None of that disturbs the proof, because the proof never looked at the picture.
This is the ordinary situation with a good proof: it uses less than the setting it was found in provides, so it carries to settings that provide less. The shadow argument uses a drawing, and there is no drawing in a space of functions. The parabola argument uses three algebraic properties, and those are available anywhere the word “inner product” is used, because they are what the word means.
When the product is only almost positive
There is a weakening of the third property that is worth separating out, because the inequality survives it and the equality case does not.
Suppose the rule is symmetric and linear in each slot, and for every — but some non-zero vector has . Such a rule is called a semi-inner product, and it is not an exotic object: it is what arrives whenever a genuine inner product is written down before the objects that ought to be identified have been.
The parabola argument does not notice. It only ever used that is never negative, and that is still true; the leading coefficient may now be zero, in which case the quadratic is a straight line rather than a parabola, and a line that is never negative has to be horizontal — which forces and makes the inequality read . So Cauchy–Schwarz holds in full, with the degenerate case handled by the same sentence rather than by a special argument.
What is lost is the converse. Equality now means only that has zero length, and a vector of zero length need not be the zero vector, so equality no longer forces to be a multiple of .
Two familiar situations are exactly this. The integral product on functions calls two functions differing on a set of measure zero indistinguishable, so the rule is only semi-definite until such functions are identified with each other — which is what the standard construction does and why it is done. And the covariance of two measurements assigns zero length to a constant, so a constant is perfectly correlated with everything and with nothing, and the correlation coefficient of a constant with anything is not defined rather than being or .
The lesson generalises past this inequality. An inequality and its equality case have different hypotheses, and the second is nearly always the stronger of the two. A result quoted without its equality case has been quoted at half strength; a result whose equality case is quoted without checking that the hypothesis for it holds has been quoted at more than full strength.
Which property is load-bearing, and where the proof would break
The rung below asserted that positivity is what everything depends on. The proof says where, exactly, and the location is a single coefficient.
The quadratic is , and the argument is “a parabola that opens upward and never dips below the axis has no room for two roots”. Opening upward is the coefficient being positive, and that is positivity, sitting in the leading term. Take a rule that is symmetric and linear in each slot but allows — the inner product of relativity is the standard one — and the parabola opens downward. A downward parabola is negative for large in both directions, the non-negativity claim is simply false, and the discriminant argument has nothing to work with.
The failure is not marginal. For two vectors both pointing into the future the inequality reverses, and — the ratio leaves altogether and names a hyperbolic angle rather than a circular one.
A proof that fails visibly at one coefficient is worth more than a proof that fails somewhere unspecified. The geometric argument’s failure in the same setting is describable but vague: light does not work like that, shadows do not behave. The algebraic argument’s failure has an address.
Three people, three settings, sixty-seven years
Cauchy proved the sum version in 1821, in the Cours d’Analyse, as a lemma he needed and not as a result he was announcing. Bunyakovsky published the integral version in 1859, in a memoir on inequalities, and it was largely unread outside Russia. Schwarz proved the integral version again in 1888, independently, in a paper on minimal surfaces, and it is his proof — the one above, with the parabola — that spread.
The sequence is worth noticing because the mathematics did not change across it and the setting did. Cauchy’s is a statement about finite lists; Bunyakovsky’s and Schwarz’s are about functions; and the modern statement is about any inner product space at all, which is a definition that did not exist when any of the three were writing. The abstraction arrived last, after three people had proved essentially one thing about three different objects, which is the usual order.
The name has settled as Cauchy–Schwarz in English and Cauchy–Bunyakovsky–Schwarz in Russian, and the discrepancy is a fair record of who read whom.
There is a detail of Cauchy’s version worth keeping, because it is the reason his proof looks different from the one above without being a different proof. He did not write a quadratic; he wrote down Lagrange’s identity,
which is an exact equation between two polynomials, and observed that the right-hand side is a sum of squares. That is the same manoeuvre — name the thing being squared — carried out in coordinates rather than in vectors, and the terms appearing on the right are the reason the cross product exists and the reason it has the entries it has. The vector proof is Lagrange’s identity with the coordinates rubbed out, which is exactly the trade that lets it survive into spaces that have none.
What it costs, and the one place it is checked
Evaluating both sides of the inequality for two lists of numbers costs three passes: multiplications and additions for the product, the same again for each squared length, and two square roots. Linear in the length of the lists, with a constant of three, and nothing about it is interesting.
What is interesting is that the inequality is essentially never evaluated. It is a licence rather than a test — a step in an argument where an unknown cross-term is replaced by a known bound — and a program that computed both sides and compared them would be answering a question nobody asked.
With one exception, and it is the reason numerical libraries carry a line of code about this. Computing an angle means forming the ratio and taking its arccosine, and in floating-point arithmetic that ratio can come out as for two vectors that are parallel. The rounding errors in three separate sums do not cancel, the true value is , the computed value is a hair above it, and the arccosine of a number greater than one is not a number. So the ratio is clamped to before the arccosine is taken, in every library that computes an angle, and the clamp is there because the inequality is exactly true and its floating-point evaluation is not.
That is a small thing with a general shape worth naming. A theorem guarantees that a quantity lies in a range; the arithmetic that computes the quantity has no such guarantee; and the gap between them shows up as a domain error in a function called two lines later, a long way from the sum where the error was made.
What the picture cannot show
The parabola is drawn over a range of about a finger’s width wide, and the claim is about every real . Outside the plotted range the curve continues upward and is not drawn, which is the ordinary limitation; what is less ordinary is that the interesting part of the claim is entirely at the vertex, so the plot shows the whole of what matters and the truncation costs nothing.
What the figure genuinely cannot show is the case the proof is for. Every drawing here is of two arrows in a plane, where the inequality is visible without any argument at all, and the whole value of the algebraic proof is that it works where the arrows are not. A picture of the inequality for two functions in would be a picture of a parabola with no accompanying arrows — which is to say, the same picture with the part a reader trusts removed.
And the figure cannot show that the parabola is a parabola. The plotted curve is 201 sampled points joined by straight segments, and no finite sample distinguishes a quadratic from something that agrees with it at those points. That the expansion has no term is a consequence of the product being linear in each slot, which is an assumption, not a measurement.
The ladder from here
Rungs above: Gram–Schmidt, where the leftover this essay measured is used rather than merely bounded, and the orthonormal bases it produces. Projection onto a subspace rather than onto a single vector, with the normal equations derived from the picture. The adjoint of a matrix, defined by moving it across a product — and why a symmetric matrix has perpendicular eigenvectors, which is the same definition used twice. Bessel’s inequality and Parseval’s identity, which are Cauchy–Schwarz’s statements about infinitely many directions at once. Hölder’s inequality, which replaces the two squares by a and a adding reciprocally to one and has this as its case . And the cross product, which keeps the perpendicular information the dot product throws away, and exists in three dimensions and almost nowhere else.
Turning an inequality into a non-negativity
The move is the one worth carrying, because it works far outside this subject: to prove that one quantity does not exceed another, find something whose square is the difference.
Here the something is , and its squared length is exactly , so the inequality is the assertion that a certain vector has a length. Nothing is being bounded; something is being exhibited.
The same manoeuvre proves that the arithmetic mean beats the geometric one, because their difference is half of . It proves that a variance is non-negative, because a variance is an average of squares. It is why a convex function lies above its tangent — the gap is a second derivative integrated twice against something non-negative — and it is why Jensen’s inequality is provable at all.
The reverse move is the failure mode, and it is what makes inequalities feel like a bag of tricks. Presented as and verified by expanding both sides in coordinates, the result is true, opaque, and unavailable anywhere the coordinates are not. Presented as a length, it is one sentence and it travels.
What links here
Computed from the collection, not written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The nearest point of a flat thing — both name dot product, inner product, orthogonality, projection
- Circles that are diamonds and squares — both name inner product, norm
Named objects
A dashed tag is an object no other essay names yet.
Cauchy schwarzDiscriminantDot productInequalityInner productNormOrthogonalityPositive definiteProjectionTriangle inequality