The exponential of a square
Worth reading first: The equation with only one answer · The same map in a better basis.
The equation has exactly one solution through each starting point, and it is . That is a statement about one quantity growing at a rate proportional to itself.
Almost nothing in the world is one quantity. Two chemicals feeding into each other, two populations preying on one another, a current and a voltage in a circuit, a position and a velocity — each of these is a pair whose rates depend on both members, and the natural way to write that is
with a pair and a matrix. The question this rung answers is whether the solution is still an exponential, and the answer is that it is, of exactly the same shape, once the notation is taken seriously.
Taking the series seriously
The exponential has a series,
and every operation in it — adding, multiplying by itself, dividing by a whole number — is something a square matrix can do. So define
with in the place of 1. Nothing has been assumed; a definition has been written down, and the only question is whether it converges. It does, and comfortably: the entries of grow at most like while the denominator is , which wins.
The figure sums sixty terms and asserts that the sixtieth is smaller than in every entry, which for the matrices drawn here it is by a wide margin.
Three routes to the same matrix
A definition by a series is worth having and hard to trust. The figure therefore computes by three routes that share no arithmetic and requires all three to agree to seven decimal places.
By the series, as above.
By the eigenvalues. If then , so every term of the series has a on the left and a on the right and they can be factored out:
where is the diagonal matrix of and . Exponentiating a matrix means exponentiating its eigenvalues, which is diagonalisation doing what it always does — turning a hard operation on a matrix into an easy operation on two numbers.
By following the flow. Integrate the differential equation numerically from and from , and see where each arrives after one unit of time. Those two arrival points are the columns of , because a linear map is determined by what it does to the two basis arrows.
That third route is the one that makes the picture evidence rather than decoration. The curves drawn on the page are that integration, so the paths and the matrix in the caption are one computation, and the agreement with the series is a claim that could have failed.
Doubling the time is not a new idea and it is worth drawing, because it makes the object visible as a family: there is one matrix for each , they compose, and the differential equation is the statement of how the family changes with . That family is what the next section is about.
Why it solves the equation
The claim is that solves with , and the proof is the series differentiated term by term:
which is times the original. Setting gives , so the initial condition holds. And the uniqueness argument from the previous rung transfers unchanged: if then , so is constant.
Every linear system with constant coefficients is therefore solved, completely, by one formula — and the work of understanding a particular system reduces to understanding one matrix.
That is worth pausing on as a matter of method. The series was derived for real numbers and is being applied to objects that are not numbers, and the licence for doing so is not analogy — it is that every operation appearing in it is defined for the new objects, and the convergence argument survives with sizes in place of absolute values. Wherever those two conditions hold, the exponential comes along: for complex numbers, for matrices, for bounded operators on a space of functions, for quaternions. The definition travels because it asks for so little.
The one law that fails
This is where the analogy has to be watched, and the failure is not a technicality.
For numbers, . For matrices this is false in general, and it is false for the reason everything about matrices is: and need not commute. Multiplying out the two series and comparing term by term, the quadratic terms are on one side and on the other, and those differ by .
So exactly when and commute, and the discrepancy is measured by the commutator — a quantity that is zero for numbers and is the whole subject of Lie theory for matrices. What replaces the addition law is the Baker–Campbell–Hausdorff formula, an infinite series of nested commutators, and its existence is the reason the exponential is the bridge between a group of transformations and the flat space of matrices near its identity.
The failure is not a blemish to be worked around; it is where the subject is. Two flows that do not commute — a rotation about one axis and a rotation about another — compose into something that depends on the order, and the commutator measures by how much. Doing then then undoing then undoing leaves a residue proportional to , which is why a car can be parked sideways by a sequence of forward and turning motions none of which moves it sideways. The matrix exponential is the setting in which that observation becomes an equation.
The single fact that does survive is , because commutes with itself. So is always invertible, whatever is — including matrices with zero determinant, whose exponentials are perfectly well behaved.
A rotation from no trigonometry
Here is the case that makes the whole construction feel inevitable.
Take the matrix that sends to — a quarter turn — and note . Then , , and the powers cycle with period four exactly as the powers of do. Feeding into the series and separating the terms with from the terms with :
and the two brackets are the series for cosine and sine. So
which is a rotation by . Nothing trigonometric was assumed; the two series fell out of sorting the powers of into even and odd.
This is Euler’s identity with a matrix where the imaginary unit was, and the correspondence is exact: the matrices multiply exactly as the complex numbers do. The complex plane is a subring of the two-by-two matrices, and the exponential agrees on both sides of that identification. So and are the same statement, one of them written with an and the other with a half turn drawn on a page.
Reading the behaviour off two numbers
Because the eigenvalues decide everything, the qualitative behaviour of a two-dimensional linear system is decided by the trace and the determinant — the same two numbers that make up the characteristic polynomial — and the classification is short enough to hold in the head.
If the determinant is negative the eigenvalues are real with opposite signs, and the origin is a saddle: points run out along one invariant line and in along the other, as in the second figure above. If the determinant is positive, the sign of the trace decides whether everything runs out or in, and the discriminant decides whether it does so along straight lines or by spiralling. A zero trace with positive determinant gives closed orbits and nothing else happens at all.
That is the whole of it: four behaviours, decided by two numbers and one inequality between them. The reason a subject as large as linear systems has so short a classification is that the diagonal form leaves nothing to happen beyond two independent one-dimensional stories, and the only extra possibilities come from the eigenvalues being complex or the eigenvectors being missing.
It also means the boundaries of the classification matter more than the interior. A system whose trace is exactly zero has closed orbits; nudge the trace and every orbit becomes a spiral, inwards or outwards according to the sign. A closed orbit is therefore not a robust feature of a linear system, which is why closed orbits in nature are almost always evidence of something nonlinear holding them in place.
What the trace and determinant do
Two properties survive the exponential in a clean form.
For a diagonalisable matrix this is immediate: the determinant of is the product of , which is , and the sum of the eigenvalues is the trace. The figure checks it for every matrix it draws, including the ones with complex eigenvalues where the eigen-route is not available.
The identity has a direct reading in the picture. The determinant of is the factor by which the flow multiplies areas after time , so a system whose matrix has zero trace preserves area exactly — every blob of starting points is carried to a blob of the same size, however distorted. That is why the skew matrix above, whose trace is zero, has circles for orbits and why an area-preserving flow is the normal state of affairs in mechanics rather than a coincidence.
The defective case is worth seeing because the series does not care. Diagonalisation fails and the definition does not depend on it; exists for every square matrix, and for a nilpotent one it is a finite sum. Where a defective matrix does show is in the behaviour: a repeated eigenvalue with a missing eigenvector produces solutions with a factor of in front of the exponential, which is growth no diagonal system can produce.
Where this is used, and why it is not optional
Three settings, chosen because in each the matrix exponential is the definition rather than a technique.
Mechanics. A system near equilibrium is described by a linear system, and its behaviour over time is applied to the initial state. Whether a bridge oscillates, damps or diverges is the sign of a real part; whether the oscillation is at one frequency or several is how many complex pairs there are.
Markov chains in continuous time. A process that jumps between states at constant rates has a matrix of rates whose exponential is the matrix of probabilities after time . The stationary distribution is again the eigenvector belonging to eigenvalue zero of the rate matrix, which becomes eigenvalue one after exponentiating — the same object as in the discrete case, reached by a different door.
Rotations in three dimensions. Every rotation is for a skew matrix built from its axis, which is Rodrigues’ formula and is the reason rotations are stored and interpolated as axis-and-angle rather than as nine numbers. The skew matrices are the tangent space to the rotations at the identity, and the exponential is the map from the flat thing to the curved one — the two-dimensional case above, one dimension up.
The common thread is worth stating. In each case there is a set of transformations that forms a group under composition and is curved, and a set of “generators” that is flat and easy to compute with, and the exponential is the bridge. Everything hard about the group is pushed into the failure of , which is where the curvature went.
What the picture cannot show
The flow is drawn for one unit of time, and for other is a different matrix. A picture of one moment cannot show that the family is a group — that running for time and then for time is the same as running for — although that is the property doing most of the work whenever this object is used. Drawing it would require several panels showing the same thing.
Nor does the two-dimensional picture suggest what happens with a large matrix, where the interesting question is which of many eigenvalues dominates and how long the transient behaviour lasts before it does. The rate at which a solution settles is governed by the eigenvalue with the largest real part, and the size of the excursion before it settles is governed by something else entirely — the same gap between eventual behaviour and immediate behaviour that separates eigenvalues from singular values.
And there is a practical warning the drawing gives no hint of. Computing by summing the series is a poor idea for a real matrix: for a matrix with large entries the terms grow enormously before the factorials take over, and the answer is a small number obtained by cancelling large ones. There is a well-known paper called Nineteen Dubious Ways to Compute the Exponential of a Matrix, and summing the definition is one of them.
The ladder from here
Rungs above: the logarithm of a matrix, and when it exists. The one-parameter groups that every matrix generates, and Lie’s correspondence between them and their generators. The Baker–Campbell–Hausdorff series in full. Second-order equations turned into first-order systems by carrying the derivative as an extra coordinate. Stability, read off the real parts of the eigenvalues. And the exponential of an operator on functions, where the same series generates the heat flow and the translation of a graph along its own axis.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The directions a map leaves alone — both name determinant, eigenvalue, eigenvector, matrix
- The number that says how much room is left — both name determinant, eigenvalue, matrix
- A matrix is a picture of what happens to the grid — both name determinant, matrix
- Symmetry forces a right angle — both name eigenvalue, eigenvector
- The curve that is its own slope — both name e, the number, power series
- The flat map that fits closest — both name determinant, matrix
Named objects
A dashed tag is an object no other essay names yet.
DeterminantDifferential equatione, the numberEigenvalueEigenvectorMatrixMatrix exponentialPower seriesRotationTrace