Moving a map across a product
Worth reading first: The dot product is a shadow · The nearest point of a flat thing.
The normal equations of a least-squares fit are , and the essay that derived them got them from a right angle rather than from a derivative. The transpose appeared because perpendicular to every column meant a dot product with every column, and stacking those dot products into one expression is what does.
That leaves a question the derivation never had to face. The transpose is defined by an operation on a table — swap the entry in row , column with the one in row , column — and nothing about swapping table entries sounds geometric. Yet it turned up in an argument about the nearest point of a plane, which is about as geometric as arguments get.
The answer is that the transpose is not really an operation on tables. It is the one map that can be moved from one side of a product to the other without changing the number the product returns:
The map is the adjoint of , and the claim of this essay is that it depends on the product as much as on the map.
A definition that never mentions an entry
Fix and read the left-hand side as a function of alone. Applying is linear and taking a product with a fixed vector is linear, so is a linear function of — a rule that takes a vector and returns a number, respecting sums and scalings.
Every such rule on a finite-dimensional space with an inner product is a product with one particular vector. That is the step doing the work, and it is short: a linear rule is fixed by its values on a basis, and a basis on which the product is well behaved lets those values be matched by exactly one vector. Call that vector , so that for every . It depends on , and it depends on linearly, so the assignment is itself a linear map. That map is .
Nothing in that construction mentions a coordinate. It uses a map and a product, and it produces a map, and the only thing it needs of the product is that no non-zero vector is perpendicular to everything — which a genuine inner product guarantees, since a vector’s product with itself is its squared length.
Now write the ordinary dot product in coordinates. , and regrouping the same sum around gives , which is dotted with the vector whose -th entry sums down column . That vector is . So under the ordinary product the adjoint is the transpose, and the reflection across the diagonal is what moving a map across a product looks like when it is written as a table.
The argument that a symmetric matrix has perpendicular eigenvectors already used exactly this: its one essential step moved the matrix across a dot product. What it did not need to ask is whether “across” means the same thing under every product. It does not.
Measure lengths another way and the flip is wrong
An inner product is any rule that is symmetric, linear in each slot, and positive on every non-zero vector, and the ordinary one is not the only one. The rule qualifies whenever and , and its vectors of length one form an ellipse rather than a circle.
Write such a rule as with the symmetric matrix of its three numbers. Then the left side of the defining identity is and the right side, for a candidate , is . For these to agree for every pair, must equal , and so
When is the identity that is the transpose again. When it is not, it is generally a different matrix entirely, and the transpose — still perfectly well defined as a table — no longer does the job the definition asks of it.
The two arrows on the right of that figure make the point without any algebra. Under the ordinary product the adjoint and the transpose would have landed on the same arrow; under this one they point in visibly different directions, and only one of them preserves the number. Which matrix counts as the flip of was decided by the ellipse, not by .
The transpose was a statement about a basis
There is a way of seeing why the transpose worked at all, and it turns the previous section from a surprise into a consequence.
Every inner product on a finite-dimensional space is the ordinary dot product in some coordinates. Gram–Schmidt finds them: start from any basis, straighten it one direction at a time using the product in hand, and the result is a basis in which that product computes as . The tilted ellipse becomes a circle, because the coordinates have been chosen so that it is one.
In those coordinates the adjoint is the transpose of the matrix in those coordinates — the argument two sections ago goes through word for word. So the general statement is: the adjoint’s matrix is the transpose of the map’s matrix exactly when the basis is orthonormal for the product being used. The transpose was never a property of the map. It was a property of a map written in a basis that happened to suit the product, and the ordinary coordinates suit only the ordinary product.
That has a consequence people meet without recognising it. The gradient of a function is usually introduced as the list of its partial derivatives, and described as the direction of steepest ascent. But “steepest” is a comparison of rates per unit length, and length needs a product. The derivative at a point is a linear rule taking a direction to a rate; the gradient is the vector whose product with each direction gives that rate — it is the derivative moved across the product, which is an adjoint. Under it is times the list of partial derivatives, and it points somewhere else. The direction of steepest ascent is not a property of the function alone, and a method that descends along the raw partials has quietly chosen the ordinary product whether or not that product means anything for the problem.
Symmetric is a word about a pair
A map is self-adjoint under a product when it equals its own adjoint. Under the ordinary product that is the familiar symmetric matrix. Under it is the condition , which rearranges to being symmetric — a condition on the map and the product together, and not on the map.
The spectral theorem does not care which product it is handed. The argument that symmetric matrices have perpendicular eigenvectors used two facts: that the map moves across the product, and that the product is symmetric and positive. Replace “dot product” by throughout and every line survives. A self-adjoint map has real eigenvalues and eigenvectors that are perpendicular — under the product it is self-adjoint for.
The converse half is the interesting one to draw. A matrix that is symmetric as a table has page-perpendicular eigenvectors, and under a different product those same eigenvectors are generally not perpendicular at all, because the map is no longer self-adjoint.
So “these two directions are perpendicular” is incomplete as a sentence in the way “these two cities are far apart” is incomplete without a map. The right angle in the rotation that diagonalises a quadratic form is a right angle under the ordinary product because that is the product the form was being compared against, and it is the eigenvectors’ relationship to that product — not anything about the matrix alone — that the right angle records.
A chain that becomes symmetric once it is weighted
The reverse situation — a matrix that is not symmetric as a table and is self-adjoint under a well-chosen product — is not a curiosity. It is exactly what a reversible random process is.
Take a two-state chain that stays in the first state nine times in ten and in the second state eight times in ten. Its transition matrix is plainly not symmetric. Its stationary distribution puts twice as much weight on the first state as the second, and the flow between the states balances under that weighting: . That balance — the long-run rate of jumping from each state to the other being the same in both directions — is called detailed balance, and it is what a chain that runs the same backwards means.
Weight the product by the stationary distribution, , and the transition matrix becomes self-adjoint. Detailed balance is that statement: is symmetric precisely when each off-diagonal pair balances.
Everything that makes reversible chains tractable follows from the picture. The eigenvalues are real, so the approach to equilibrium decays rather than spirals. The eigenvectors are perpendicular under the weighted product, so any starting distribution splits cleanly into a stationary part and parts that shrink, each by its own eigenvalue, without interfering. The spectral theorem applies to a matrix that is not symmetric, and it applies because the question “symmetric under what?” had an answer the table could not show.
What a right angle looks like on an ellipse
The dashed lines in the last two figures deserve a paragraph, because they are how “perpendicular under the product” is made visible on a page that only knows the ordinary angle.
Two directions are perpendicular under when . The vector is, up to scale, the direction in which the length grows fastest from the tip of — it is perpendicular on the page to the unit ellipse at that tip. So says is page-perpendicular to , which says runs along the ellipse’s tangent at the tip of .
Pairs of diameters related this way have a classical name: conjugate diameters, each parallel to the tangents at the ends of the other. Apollonius had them in the third century BC as a property of the ellipse. They are exactly the pairs of directions the ellipse’s own inner product calls perpendicular, and a circle is the one ellipse on which conjugate diameters are perpendicular diameters.
That identification also explains why an ellipse has infinitely many pairs of conjugate diameters and only one pair of axes. Every direction has a perpendicular partner under the product — rotate a diameter, and its conjugate rotates too — while the axes are the one pair that is perpendicular under both the ellipse’s product and the page’s. That pair is what diagonalising the form finds.
What a map reaches is perpendicular to what its adjoint sends to nought
The single most useful consequence of the definition takes one line.
Suppose . Then for any , . Everything the map reaches is perpendicular to everything its adjoint kills. The two sets are subspaces and they are perpendicular under the product that defined the adjoint.
In finite dimensions they are more than perpendicular — they fill the space between them, because the dimension of what reaches equals the dimension of what reaches, and the rank of plus the dimension of its kernel is the whole space. So a vector is reachable exactly when it has no component along the adjoint’s kernel, and the equation has a solution exactly when for every with .
That is the Fredholm alternative in its finite-dimensional form, and it is a certificate of impossibility rather than a failed search: to show has no solution, exhibit one vector the adjoint kills on which casts a shadow. The same pattern — a solution exists, or a witness on the other side proves it cannot — is what duality in optimisation runs on, with inequalities in place of equations.
And it is where the normal equations came from. The residual of a least-squares fit, , is perpendicular to everything reaches, so it lies in what kills: . “Normal” in the normal equations means “in the kernel of the adjoint”, and appears there because the fit was measured with the ordinary product.
A weighted fit is an ordinary fit in another geometry
Measure the data with a different product and the same sentence produces a different fit. That is not a technicality; it is how a fit is made to respect measurements of different reliability.
Suppose some of six measurements are nine times as trustworthy as others — their variance is a ninth as large. The honest penalty weighs each squared residual by its reliability, , and that is a squared length under the product . The nearest point, the right angle and the normal equations all carry over, with the adjoint taken under : the equations become .
Priced in the ordinary product the ordering reverses: the dashed line is the best there is, at , and the weighted fit does worse. Neither line is the best fit. Each is the nearest point under its own product, and the question “which line is nearest to the data?” had no answer until a product was chosen — which is the same observation as the adjoint’s, arriving from the data side instead of from the map.
Integration by parts is an adjoint
The definition was built to survive the loss of coordinates, and the place it pays for that is a space with no finite basis at all.
Take functions on the interval that vanish at both ends, with the product . Differentiation is a linear map on them. Integration by parts says
and the bracket vanishes because both functions vanish at the ends. So : the adjoint of differentiation is minus differentiation. Apply it twice and the second derivative is its own adjoint.
Now the spectral theorem has something to say. The functions the second derivative merely rescales, among those vanishing at the ends, are for whole numbers , and a self-adjoint map’s eigenvectors are perpendicular. So whenever , and that is the orthogonality every Fourier sine series relies on — usually proved by a product-to-sum identity, and here obtained without computing an integral.
The boundary terms are the price, and they are the infinite-dimensional version of the ellipse. Change which functions are allowed — ask for zero slope at the ends instead of zero value — and the bracket still vanishes, the adjoint is still minus the derivative, and the eigenfunctions become cosines. Drop the condition entirely and the bracket survives, and the second derivative is not self-adjoint at all. The boundary conditions are part of the product’s geometry, in exactly the sense the weights were.
Where the adjoint needs care
Complex vectors need a conjugate. The complex inner product is conjugate-linear in one slot, so the adjoint of a complex matrix under the ordinary product is its conjugate transpose, not its transpose. A real symmetric matrix and a complex Hermitian one are the two cases of one notion.
A map between two spaces needs a product on each. For from one space to another, is measured in the second space and in the first, so the adjoint goes backwards and depends on both products. The weighted fit changed only the product on the data side; changing the product on the coefficients as well changes the adjoint again.
Infinite dimensions separate two notions finite dimensions merge. Differentiation is not defined on every function in the space — it is an unbounded map with a restricted domain — and for such maps “equals its adjoint on its domain” and “has the same domain as its adjoint” are different conditions. The first is called symmetric and the second self-adjoint, and only the second gives a spectral theorem. Quantum mechanics requires observables to be self-adjoint for exactly that reason.
And the product must be genuine. If some non-zero vector is perpendicular to everything, the construction of fails, and adjoints need not exist or need not be unique. Indefinite products, like the one in relativity, keep the non-degeneracy and lose positivity: adjoints still exist there, but the spectral theorem does not, and self-adjoint maps can have eigenvalues that are not real.
Lagrange’s multipliers, and von Neumann’s domains
The construction is older than matrices. In the 1760s Lagrange, integrating linear differential equations, multiplied an equation by an unknown function and integrated by parts to move every derivative onto the multiplier. The equation the multiplier had to satisfy is still called the adjoint equation, and it is the second half of the integration-by-parts section above, found a century before anybody wrote a matrix down.
Matrices arrived with Cayley in 1858, and the transpose with them, as an operation on arrays. For most of the following seventy years the two strands stayed separate: one about tables, one about differential operators.
What joined them was Hilbert space. Von Neumann’s work of 1929 and 1930, written to put quantum mechanics on a rigorous footing, defined the adjoint of an operator by exactly the identity at the top of this page and then discovered the problem of the previous section — that for unbounded operators the naive notion of symmetry is not strong enough, and a spectral theorem needs the domains to match. The distinction between symmetric and self-adjoint is one of the few places a physical theory forced a correction to a definition mathematicians thought they already had.
What the ellipses cannot show
Every figure here is two-dimensional, and a product on the plane is three numbers. The spaces where the adjoint earns its keep — data with thousands of coordinates, functions with no coordinates at all — have no unit curve anybody can draw.
The identity is checked on a hundred and forty-four pairs of vectors in the hero and the tilted figure, which is a sampling rather than a proof. The proof is the one line of algebra two sections in, and the pictures are evidence the algebra was carried out on the map drawn.
And the angle “under the product” never appears on the page as an angle. It is inferred from a tangent, through a construction the reader has to accept, and a reader who does not will see two arrows at seventy-two degrees. That is honest: the page has its own product and cannot display another one, and the dashed tangents are the closest a drawing gets.
Still open: what the adjoint says about sizes
The adjoint has been used here to move maps, to test for symmetry and to certify that equations have no solution. It has not yet been used to measure how much a map stretches, and that is the next thing it does.
is always self-adjoint and never negative — , a squared length — so the spectral theorem hands it perpendicular eigenvectors with non-negative eigenvalues. Their square roots are how much stretches along those directions: the axes of the ellipse a map makes of a circle, which is the singular value decomposition, and the pseudo-inverse that solves the dependent case of a fit. Bessel’s inequality and Parseval’s identity take the same projections to infinitely many directions. And the product of two vectors that the dot product discards — the perpendicular part rather than the shadow — is a plane in disguise, which exists as an arrow in three dimensions and almost nowhere else.
A definition by what an object does across a pairing
The move at the centre of this essay is to define an object not by its entries but by how it behaves inside a pairing, and it is worth recognising because it recurs.
A table of numbers is a description in coordinates, and any statement about it has to be checked against a change of basis. A relationship — “moving across the product preserves the number” — needs no such check, because it never mentioned a basis. The transpose fails that test and the adjoint passes it, and the difference between them is exactly the difference the tilted ellipse made visible.
The determinant defined as the unique function that behaves like a volume is the same move, and so is the derivative read as the flat map that fits closest rather than as a list of partial derivatives. When a formula comes out of coordinates, the question to ask is what it is the coordinate form of — and when the answer is a relationship, the formula will turn out to have been true only in the coordinates it was written in.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A walk that samples a distribution — both name eigenvector, markov chain, stationary distribution
- How long until it forgets — both name markov chain, stationary distribution
- The directions a map leaves alone — both name eigenvector, orthogonality
- The rule that forgets where it came from — both name eigenvector, markov chain
- The time spent and the share held — both name markov chain, stationary distribution
- Where the coefficients come from — both name inner product, orthogonality
Named objects
A dashed tag is an object no other essay names yet.
AdjointEigenvectorGradientInner productLeast squaresMarkov chainOrthogonalityPositive definiteResidualSpectral theoremStationary distributionSubspace