The matrix that has no logarithm
Worth reading first: Two numbers decide the flow · The exponential of a square.
For numbers, the exponential and the logarithm come as a pair. Every positive number is for exactly one real , and the logarithm is the rule that finds it. The area that names the number defined the logarithm as an area under a hyperbola precisely so that this pairing is built in from the start.
For matrices the exponential is defined by the same series, , and the exponential of a square showed that it converges for every square matrix and solves . The obvious next question is whether it can be undone. Given an invertible matrix , is there a matrix with ? And if there is, is there only one?
The answers are no and no, and both failures are informative. Some perfectly ordinary real matrices have no real logarithm at all, and some have infinitely many, sometimes a continuum of them. Which case a matrix falls into can be read off its eigenvalues, and in two dimensions the whole story fits in one picture of the complex plane.
What exponentiating does to eigenvalues
The whole criterion comes from following eigenvalues through the exponential. If , then every power of multiplies by the corresponding power of , so the series gives . The eigenvalues of are the exponentials of the eigenvalues of .
Now suppose is real. Its eigenvalues, the roots of its characteristic polynomial, are either two real numbers or a complex-conjugate pair . In the first case has eigenvalues and , and both are positive, since the exponential of a real number always is. In the second case has eigenvalues : a conjugate pair of the same size , placed symmetrically about the real axis on a circle of radius .
So the eigenvalues of a real exponential are always one of two things — two positive reals, or a conjugate pair of equal modulus. A conjugate pair can land on the negative real axis, when is an odd multiple of , but then both eigenvalues land on the same negative number, . A real matrix with a negative eigenvalue and a positive one, or with two different negative eigenvalues, is not the exponential of any real matrix. That is the matrix in the opening figure, marked D, and it is also the matrix G, whose eigenvalues are and .
The determinant says the first half of this without eigenvalues. Since , an exponential always has positive determinant, and G, with determinant , is ruled out on sight. D has positive determinant, , and passes that test; it fails only on the finer one. A matrix can reverse no orientation and still be unreachable.
The repeated negative eigenvalue, and when it is enough
The conjugate-pair case allows a negative eigenvalue twice over. That is necessary, and in two dimensions it is not sufficient, which is the difference between the matrices C and E.
C is minus the identity, . It has a logarithm, and it is one that anyone who has read multiplying is turning already knows: is rotation by half a turn, and rotation by angle is the exponential of times the matrix that turns by a right angle. So . The eigenvalues of are , whose exponentials are both .
E is : the eigenvalue repeated, but with only one eigenvector. It has no real logarithm, and the reason is the other half of the conjugate-pair story. When has a complex pair with , the two eigenvalues are different, so can be diagonalised over the complex numbers — and then so can . An exponential whose eigenvalue is repeated because of a complex pair is therefore always diagonalisable: it is times the identity, a multiple of . E is not diagonalisable, so it is not of that form, and no real matrix exponentiates to it.
The general theorem for real matrices of any size, due to Walter Culver in 1966, says exactly this in the language of Jordan blocks: a real invertible matrix has a real logarithm if and only if each Jordan block belonging to a negative eigenvalue occurs an even number of times. In two dimensions that reduces to the picture: no negative eigenvalues at all, or times the identity.
A logarithm is a motion that arrives on time
There is a way of seeing the question that makes the failures feel less arbitrary. If , then the matrices for between and form a path from the identity to , and not just any path: a steady one, in which each step is the same small transformation repeated. The flow of carries every point of the plane along, and after one unit of time it has applied . A logarithm of is a steady motion of the plane that arrives at in unit time.
Seen this way, the obstruction to a logarithm of is a statement about motion. The matrix reverses both axes and scales them by different amounts. A steady motion that reversed an axis would have to carry the points of that axis through the origin, which a linear flow cannot do, or turn them round through the plane — and turning the whole plane by a half-turn treats the two axes alike, scaling both by the same factor. No steady motion can flip two axes while stretching them unequally, because the only way to flip is to turn, and turning mixes the axes. The eigenvalue criterion is the algebraic form of that sentence.
It also explains why the determinant must be positive. A steady motion that starts at the identity never passes through a matrix of determinant zero, since is never zero; so it cannot change the sign of the determinant, and it ends where it began, at positive orientation. The determinant as the factor by which room is scaled is continuous along the motion and cannot jump from positive to negative.
Too many logarithms of a rotation
When a logarithm does exist, it is not unique, and the reason is already visible for complex numbers. The complex logarithm of is , but also , and : the exponential forgets whole turns.
For a rotation matrix the same ambiguity appears as a family of generators: for every whole number . The figure draws three of them as motions. A logarithm of is a way of reaching by a steady motion in unit time — the flow of a linear differential equation — and a quarter-turn can be reached by turning a quarter, or turning a quarter plus a whole turn, or three quarters backwards.
That discrete ambiguity is the one everyone expects. It is also usually resolved in the same way as for numbers: choose the logarithm whose eigenvalues have imaginary parts between and , the principal logarithm, which exists and is unique whenever no eigenvalue lies on the negative real axis. It is what computer libraries return, and it is what makes interpolating between two orientations in animation well defined: to move smoothly from orientation to orientation , take the principal logarithm of and run its flow.
A continuum of logarithms of minus the identity
On the negative real axis the principal logarithm breaks down, and something stranger than a discrete family appears.
is , with the right-angle turn. But was only ever used for one property: . Any real matrix with has the same exponential series, term by term, with in place of — so as well. And there are many such : every matrix of the form , for any invertible , squares to — it is the same map in a different basis, and squaring does not notice the basis.
The figure uses the stretches , giving . For each , is a rotation that has been squashed into an ellipse, and after unit time it has gone exactly halfway round — to minus the starting point. Since can be any positive number, and could be any invertible matrix, has a continuum of real logarithms. Nor do they commute with each other, although commutes with everything: the product of two of them in one order is not the product in the other.
This is the reverse of the uniqueness that makes the scalar logarithm so pleasant. For a number, or for a matrix with distinct eigenvalues, logarithms differ only by the discrete choices of branch. At every eigenvector is an eigenvector, the matrix sees no preferred axes, and a logarithm has to choose axes for itself — in as many ways as there are ellipses.
The band that no exponential reaches
The failure of the exponential to be onto has a clean picture on the matrices of determinant one, the group written .
A logarithm of a determinant-one matrix can be taken with trace zero, since . A traceless two-by-two matrix has eigenvalues , which are real if and imaginary if . So the trace of is in the first case and in the second — and it is never less than .
The curve is the whole image of the exponential, as measured by trace. To the left of the origin it rises without bound, giving the hyperbolic matrices of trace above 2; to the right it oscillates, giving the rotations of every angle; and at each of its lowest points it touches exactly, where is an odd multiple of and . The band below is never entered. Every matrix of determinant one and trace less than — a saddle-like map that also flips each axis, like — is outside the image, which is the eigenvalue criterion again: two negative eigenvalues of different sizes.
On the line itself only is reached. The matrix has trace and determinant and is no exponential, for the reason matrix E failed. So the image of the exponential in is the set of matrices with trace greater than , together with one extra point. It is not a closed set and not an open one, which is a strange shape for the image of something as smooth as a power series — and it means the group is not covered by its one-parameter subgroups, though every element is a product of exponentials, and repeated motions along one-parameter subgroups reach it.
The logarithm’s own series, and where it stops
For numbers there is a series for the logarithm too: , convergent for . The same series makes sense with a matrix in place of , and where it converges it produces a logarithm of .
The series converges exactly when every eigenvalue of lies inside the unit circle, which is to say every eigenvalue of lies within distance one of . It is the matrix version of the fact that the series for knows nothing about numbers further than one from . A rotation by has eigenvalues , at distance from , so the series runs away — while the logarithm itself exists and is simply worth of .
So the series is a local tool, as it is for numbers, and computing a logarithm of a general matrix needs more. The standard method is inverse scaling and squaring: take square roots repeatedly until the matrix is close to , apply the series there, and multiply the answer back up by the corresponding power of two, since . It is the mirror image of how the exponential itself is computed. And it works precisely when the principal logarithm exists — when no eigenvalue is on the negative real axis — because that is when the repeated square roots are defined and converge to .
Where the question is asked in earnest
Two practical settings ask exactly this question, and in both the answer “no logarithm” is a genuine finding rather than a technicality.
The first is Markov chains observed at intervals. A process that jumps between states at constant rates has a matrix of rates , and its matrix of transition probabilities over one unit of time is . Often only the transition matrix is observed — how many of the population moved from one state to another over a year — and the rates are wanted. That is a logarithm, with the extra demand that it be a valid rate matrix: non-negative off the diagonal, rows summing to zero. John Kingman posed this embedding problem in 1962, and in two states it has a clean answer — a two-state transition matrix comes from rates exactly when its determinant is positive — while in three or more it is intricate and not fully solved. A transition matrix with no valid logarithm means the observed process cannot have been a constant-rate process at all, which is a statement about the data rather than about the arithmetic.
The second is interpolation of motions, where the principal logarithm is the tool and the negative real axis is the hazard. Blending two orientations of a body, or two frames of an animation, through gives a steady motion from one to the other. When is a half-turn, the logarithm is not unique, and small changes in near it make the interpolated motion jump from turning one way to turning the other — the continuum of logarithms of is the reason. The integers among the quaternions meets the same half-turn as the point where the quaternion description of rotations has to choose a sign.
What the eigenvalue picture cannot show
The opening figure decides existence exactly, but it is a picture of eigenvalues, and eigenvalues forget things. Matrices C and E both sit on the negative axis with an eigenvalue repeated, and the picture marks them differently because of something it does not show: whether the matrix has one eigenvector or two. Move E’s eigenvalue to and the two would share a point on the plane while one has a logarithm and the other does not. The criterion is really about Jordan structure, and in dimensions higher than two a single point on the negative axis can hide Jordan blocks of several sizes, each of which must pair with another of the same size.
The pictures also show only real logarithms. Over the complex numbers every invertible matrix has a logarithm, because the complex exponential reaches every nonzero complex number, so the obstruction here is entirely about staying real. And the figures show two-by-two matrices, where the image of the exponential in the determinant-one group is drawn completely by one curve. For larger groups the image of the exponential is known in many cases and is a subtle object in general.
Still open: logarithms that must stay inside a set
Existence and uniqueness of real logarithms are settled completely by Culver’s theorem, and computing the principal logarithm is a solved numerical problem. What remains open are the versions in which the logarithm has to belong to a restricted set.
The embedding problem for Markov chains is the best known. For three states it was worked out in detail only recently, and for four or more there is no known finite test that decides whether a given transition matrix is the exponential of some rate matrix; the logarithms can be infinite in number, and it is not known in general how to search them for one that is valid. The corresponding question for groups — which elements of a given Lie group lie on some one-parameter subgroup, so that the exponential reaches them — has complete answers for many classical groups and none that covers every case at once. The same question for a single matrix but a restricted logarithm — a real logarithm that is also symmetric, or that commutes with a given matrix — is where the answers become case by case, since the eigenvalue picture decides existence and says nothing about which of a continuum of logarithms has the extra property.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A matrix that counts the returns — both name eigenvalue, trace
- A sum read from inside — both name logarithm, power series
- The curve that is its own slope — both name logarithm, power series
- The directions a map leaves alone — both name determinant, eigenvalue
- Units that form a lattice — both name determinant, logarithm
Named objects
A dashed tag is an object no other essay names yet.
DeterminantEigenvalueLogarithmMatrix exponentialPower seriesRotationTrace