Analysis

The matrix that has no logarithm

Every square matrix has an exponential, and a matrix exponential is always invertible. The converse fails, and it fails in a way that can be read straight off the eigenvalues — a real matrix with eigenvalues −1 and −2 is no exponential at all, while minus the identity is the exponential of a whole continuum of matrices that do not even commute with each other.

Worth reading first: Two numbers decide the flow · The exponential of a square.

For numbers, the exponential and the logarithm come as a pair. Every positive number is exe^x for exactly one real xx, and the logarithm is the rule that finds it. The area that names the number defined the logarithm as an area under a hyperbola precisely so that this pairing is built in from the start.

For matrices the exponential is defined by the same series, eA=I+A+A2/2!+⋯e^A = I + A + A^2/2! + \cdots, and the exponential of a square showed that it converges for every square matrix and solves x˙=Ax\dot x = Ax. The obvious next question is whether it can be undone. Given an invertible matrix MM, is there a matrix LL with eL=Me^L = M? And if there is, is there only one?

The answers are no and no, and both failures are informative. Some perfectly ordinary real matrices have no real logarithm at all, and some have infinitely many, sometimes a continuum of them. Which case a matrix falls into can be read off its eigenvalues, and in two dimensions the whole story fits in one picture of the complex plane.

Which real 2 × 2 matrices have a real logarithm, read off their eigenvalues. Eigenvalues of 7 matrices plotted in the complex plane with the negative real axis emphasised, beside a table saying for each whether a real logarithm exists and why: A yes, B yes, C yes, D no, E no, F yes, G no.
Fig. 1 The eigenvalues of seven real two-by-two matrices, lettered. A ring marks a matrix with a real logarithm, a cross one without. Every crossed matrix has a negative eigenvalue, and every ringed matrix either has none or has the negative eigenvalue −1 twice over as minus the identity. Each logarithm listed on the right was constructed and exponentiated back to the matrix it claims to undo.

What exponentiating does to eigenvalues

The whole criterion comes from following eigenvalues through the exponential. If Lv=λvLv = \lambda v, then every power of LL multiplies vv by the corresponding power of λ\lambda, so the series gives eLv=eλve^L v = e^\lambda v. The eigenvalues of eLe^L are the exponentials of the eigenvalues of LL.

Now suppose LL is real. Its eigenvalues, the roots of its characteristic polynomial, are either two real numbers or a complex-conjugate pair a±iba \pm ib. In the first case eLe^L has eigenvalues eλ1e^{\lambda_1} and eλ2e^{\lambda_2}, and both are positive, since the exponential of a real number always is. In the second case eLe^L has eigenvalues eae±ibe^a e^{\pm ib}: a conjugate pair of the same size eae^a, placed symmetrically about the real axis on a circle of radius eae^a.

So the eigenvalues of a real exponential are always one of two things — two positive reals, or a conjugate pair of equal modulus. A conjugate pair can land on the negative real axis, when bb is an odd multiple of π\pi, but then both eigenvalues land on the same negative number, −ea-e^a. A real matrix with a negative eigenvalue and a positive one, or with two different negative eigenvalues, is not the exponential of any real matrix. That is the matrix diag⁡(−1,−2)\operatorname{diag}(-1, -2) in the opening figure, marked D, and it is also the matrix G, whose eigenvalues are −2.2-2.2 and 1.41.4.

The determinant says the first half of this without eigenvalues. Since det⁡eL=etr⁡L\det e^L = e^{\operatorname{tr} L}, an exponential always has positive determinant, and G, with determinant −3.08-3.08, is ruled out on sight. D has positive determinant, 22, and passes that test; it fails only on the finer one. A matrix can reverse no orientation and still be unreachable.

The repeated negative eigenvalue, and when it is enough

The conjugate-pair case allows a negative eigenvalue twice over. That is necessary, and in two dimensions it is not sufficient, which is the difference between the matrices C and E.

C is minus the identity, −I-I. It has a logarithm, and it is one that anyone who has read multiplying is turning already knows: −I-I is rotation by half a turn, and rotation by angle θ\theta is the exponential of θ\theta times the matrix JJ that turns by a right angle. So −I=eπJ-I = e^{\pi J}. The eigenvalues of πJ\pi J are ±iπ\pm i\pi, whose exponentials are both −1-1.

E is (−1.410−1.4)\begin{pmatrix} -1.4 & 1 \\ 0 & -1.4 \end{pmatrix}: the eigenvalue −1.4-1.4 repeated, but with only one eigenvector. It has no real logarithm, and the reason is the other half of the conjugate-pair story. When LL has a complex pair a±iba \pm ib with b≠0b \ne 0, the two eigenvalues are different, so LL can be diagonalised over the complex numbers — and then so can eLe^L. An exponential whose eigenvalue is repeated because of a complex pair is therefore always diagonalisable: it is −ea-e^a times the identity, a multiple of −I-I. E is not diagonalisable, so it is not of that form, and no real matrix exponentiates to it.

The general theorem for real matrices of any size, due to Walter Culver in 1966, says exactly this in the language of Jordan blocks: a real invertible matrix has a real logarithm if and only if each Jordan block belonging to a negative eigenvalue occurs an even number of times. In two dimensions that reduces to the picture: no negative eigenvalues at all, or −r-r times the identity.

A logarithm is a motion that arrives on time

There is a way of seeing the question that makes the failures feel less arbitrary. If eL=Me^L = M, then the matrices etLe^{tL} for tt between 00 and 11 form a path from the identity to MM, and not just any path: a steady one, in which each step is the same small transformation repeated. The flow of x˙=Lx\dot x = Lx carries every point of the plane along, and after one unit of time it has applied MM. A logarithm of MM is a steady motion of the plane that arrives at MM in unit time.

The flow of a linear equation, and the matrix that runs it for one unit of time. Paths of points moving so that their velocity is [0.69, 0.92, 0, −0.69] applied to their position, with the position after time 1 marked on each; the matrix taking start to finish is e^A.
Fig. 2 The flow of the logarithm of matrix A from the opening figure, drawn for one unit of time. Each point moves along its own curve, the dots mark where it has arrived, and the matrix that takes start to arrival is A itself — the logarithm, run as a differential equation, rebuilds the matrix it was taken from.

Seen this way, the obstruction to a logarithm of diag⁡(−1,−2)\operatorname{diag}(-1, -2) is a statement about motion. The matrix reverses both axes and scales them by different amounts. A steady motion that reversed an axis would have to carry the points of that axis through the origin, which a linear flow cannot do, or turn them round through the plane — and turning the whole plane by a half-turn treats the two axes alike, scaling both by the same factor. No steady motion can flip two axes while stretching them unequally, because the only way to flip is to turn, and turning mixes the axes. The eigenvalue criterion is the algebraic form of that sentence.

It also explains why the determinant must be positive. A steady motion that starts at the identity never passes through a matrix of determinant zero, since det⁡etL=ettr⁡L\det e^{tL} = e^{t \operatorname{tr} L} is never zero; so it cannot change the sign of the determinant, and it ends where it began, at positive orientation. The determinant as the factor by which room is scaled is continuous along the motion and cannot jump from positive to negative.

Too many logarithms of a rotation

When a logarithm does exist, it is not unique, and the reason is already visible for complex numbers. The complex logarithm of ii is iπ/2i\pi/2, but also iπ/2+2πii\pi/2 + 2\pi i, and iπ/2−2πii\pi/2 - 2\pi i: the exponential forgets whole turns.

Three logarithms of a quarter-turn, winding differently to one point. Paths of the point (1, 0) under exp(tL) for three logarithms L of a 90° rotation, winding different numbers of times and all ending at the same point.
Fig. 3 Three logarithms of the quarter-turn — rotations generated to turn −270°, 90° and 450° in unit time. Each path is a point carried along by etLe^{tL} as t runs from 0 to 1; they wind different amounts, in different directions, and all end on the upper axis, because every one of these matrices exponentiates to the same quarter-turn.

For a rotation matrix the same ambiguity appears as a family of generators: (θ+2πk)J(\theta + 2\pi k)J for every whole number kk. The figure draws three of them as motions. A logarithm of MM is a way of reaching MM by a steady motion in unit time — the flow t↦etLt \mapsto e^{tL} of a linear differential equation — and a quarter-turn can be reached by turning a quarter, or turning a quarter plus a whole turn, or three quarters backwards.

That discrete ambiguity is the one everyone expects. It is also usually resolved in the same way as for numbers: choose the logarithm whose eigenvalues have imaginary parts between −π-\pi and π\pi, the principal logarithm, which exists and is unique whenever no eigenvalue lies on the negative real axis. It is what computer libraries return, and it is what makes interpolating between two orientations in animation well defined: to move smoothly from orientation AA to orientation BB, take the principal logarithm of A−1BA^{-1}B and run its flow.

A continuum of logarithms of minus the identity

On the negative real axis the principal logarithm breaks down, and something stranger than a discrete family appears.

−I-I is eπJe^{\pi J}, with JJ the right-angle turn. But JJ was only ever used for one property: J2=−IJ^2 = -I. Any real matrix KK with K2=−IK^2 = -I has the same exponential series, term by term, with KK in place of JJ — so eπK=cos⁡π I+sin⁡π K=−Ie^{\pi K} = \cos\pi\, I + \sin\pi\, K = -I as well. And there are many such KK: every matrix of the form PJP−1PJP^{-1}, for any invertible PP, squares to −I-I — it is the same map in a different basis, and squaring does not notice the basis.

A continuum of logarithms of minus the identity, as half-turns along ellipses. Half-elliptical paths from (1, 0) to (−1, 0) under exp(tπK) for 5 matrices K with K² = −I, all logarithms of −I.
Fig. 4 Five real logarithms of −I-I, each π\pi times a matrix K whose square is −I. Carrying a point along etπKe^{t\pi K} from t = 0 to t = 1 traces half of an ellipse, a different ellipse for each K, and every path ends at the opposite point. The ellipses are the circles of the rotation seen through a stretch, and there is one for every stretch.

The figure uses the stretches P=diag⁡(1,s)P = \operatorname{diag}(1, s), giving K=(0−1/ss0)K = \begin{pmatrix} 0 & -1/s \\ s & 0 \end{pmatrix}. For each ss, etπKe^{t\pi K} is a rotation that has been squashed into an ellipse, and after unit time it has gone exactly halfway round — to minus the starting point. Since ss can be any positive number, and PP could be any invertible matrix, −I-I has a continuum of real logarithms. Nor do they commute with each other, although −I-I commutes with everything: the product of two of them in one order is not the product in the other.

This is the reverse of the uniqueness that makes the scalar logarithm so pleasant. For a number, or for a matrix with distinct eigenvalues, logarithms differ only by the discrete choices of branch. At −I-I every eigenvector is an eigenvector, the matrix sees no preferred axes, and a logarithm has to choose axes for itself — in as many ways as there are ellipses.

The band that no exponential reaches

The failure of the exponential to be onto has a clean picture on the matrices of determinant one, the group written SL2(R)SL_2(\mathbb R).

A logarithm of a determinant-one matrix can be taken with trace zero, since det⁡eL=etr⁡L\det e^L = e^{\operatorname{tr}L}. A traceless two-by-two matrix LL has eigenvalues ±−det⁡L\pm\sqrt{-\det L}, which are real if det⁡L<0\det L < 0 and imaginary if det⁡L>0\det L > 0. So the trace of eLe^L is 2cosh⁡−det⁡L2\cosh\sqrt{-\det L} in the first case and 2cos⁡det⁡L2\cos\sqrt{\det L} in the second — and it is never less than −2-2.

The traces that exponentials of determinant one can have, and the band they miss. Traces of exp L for 177 random traceless 2 × 2 matrices L against det L, lying on the curve 2cos√(det L), which oscillates between 2 and −2; the region of trace below −2 is shaded as unreachable.
Fig. 5 Random traceless matrices L placed by their determinant and the trace of their exponential. Every one lands on the single curve 2cos⁡det⁡L2\cos\sqrt{\det L}, which swings between 2 and −2 and never goes below. The shaded band — trace less than −2 — holds determinant-one matrices, such as diag(−2, −1/2), that are the exponential of nothing.

The curve is the whole image of the exponential, as measured by trace. To the left of the origin it rises without bound, giving the hyperbolic matrices of trace above 2; to the right it oscillates, giving the rotations of every angle; and at each of its lowest points it touches −2-2 exactly, where det⁡L\sqrt{\det L} is an odd multiple of π\pi and eL=−Ie^L = -I. The band below −2-2 is never entered. Every matrix of determinant one and trace less than −2-2 — a saddle-like map that also flips each axis, like diag⁡(−2,−1/2)\operatorname{diag}(-2, -1/2) — is outside the image, which is the eigenvalue criterion again: two negative eigenvalues of different sizes.

On the line −2-2 itself only −I-I is reached. The matrix (−110−1)\begin{pmatrix} -1 & 1 \\ 0 & -1\end{pmatrix} has trace −2-2 and determinant 11 and is no exponential, for the reason matrix E failed. So the image of the exponential in SL2(R)SL_2(\mathbb R) is the set of matrices with trace greater than −2-2, together with one extra point. It is not a closed set and not an open one, which is a strange shape for the image of something as smooth as a power series — and it means the group is not covered by its one-parameter subgroups, though every element is a product of exponentials, and repeated motions along one-parameter subgroups reach it.

The logarithm’s own series, and where it stops

For numbers there is a series for the logarithm too: log⁡(1+x)=x−x2/2+x3/3−⋯\log(1 + x) = x - x^2/2 + x^3/3 - \cdots, convergent for ∣x∣<1|x| < 1. The same series makes sense with a matrix XX in place of xx, and where it converges it produces a logarithm of I+XI + X.

The logarithm's power series for three matrices: fast, slow, and divergent. Error of the partial sums of the logarithm series for three 2 × 2 matrices against the number of terms, on a logarithmic scale; the eigenvalues of X have sizes 0.20, 0.79, 1.93.
Fig. 6 The logarithm’s series for three matrices of the form I + X, with the error after each term on a logarithmic scale. When the eigenvalues of X are small the series closes in fast; near the unit circle it crawls; and for a turn by 149°, whose eigenvalues of X have size 1.93, it diverges — though that matrix has a perfectly good logarithm.

The series converges exactly when every eigenvalue of XX lies inside the unit circle, which is to say every eigenvalue of M=I+XM = I + X lies within distance one of 11. It is the matrix version of the fact that the series for log⁡(1+x)\log(1 + x) knows nothing about numbers further than one from 11. A rotation by 149°149° has eigenvalues e±149°ie^{\pm 149° i}, at distance 1.931.93 from 11, so the series runs away — while the logarithm itself exists and is simply 149°149° worth of JJ.

So the series is a local tool, as it is for numbers, and computing a logarithm of a general matrix needs more. The standard method is inverse scaling and squaring: take square roots repeatedly until the matrix is close to II, apply the series there, and multiply the answer back up by the corresponding power of two, since log⁡M=2klog⁡M1/2k\log M = 2^k \log M^{1/2^k}. It is the mirror image of how the exponential itself is computed. And it works precisely when the principal logarithm exists — when no eigenvalue is on the negative real axis — because that is when the repeated square roots are defined and converge to II.

Where the question is asked in earnest

Two practical settings ask exactly this question, and in both the answer “no logarithm” is a genuine finding rather than a technicality.

The first is Markov chains observed at intervals. A process that jumps between states at constant rates has a matrix of rates QQ, and its matrix of transition probabilities over one unit of time is eQe^Q. Often only the transition matrix is observed — how many of the population moved from one state to another over a year — and the rates are wanted. That is a logarithm, with the extra demand that it be a valid rate matrix: non-negative off the diagonal, rows summing to zero. John Kingman posed this embedding problem in 1962, and in two states it has a clean answer — a two-state transition matrix comes from rates exactly when its determinant is positive — while in three or more it is intricate and not fully solved. A transition matrix with no valid logarithm means the observed process cannot have been a constant-rate process at all, which is a statement about the data rather than about the arithmetic.

The second is interpolation of motions, where the principal logarithm is the tool and the negative real axis is the hazard. Blending two orientations of a body, or two frames of an animation, through etlog⁡Me^{t \log M} gives a steady motion from one to the other. When MM is a half-turn, the logarithm is not unique, and small changes in MM near it make the interpolated motion jump from turning one way to turning the other — the continuum of logarithms of −I-I is the reason. The integers among the quaternions meets the same half-turn as the point where the quaternion description of rotations has to choose a sign.

What the eigenvalue picture cannot show

The opening figure decides existence exactly, but it is a picture of eigenvalues, and eigenvalues forget things. Matrices C and E both sit on the negative axis with an eigenvalue repeated, and the picture marks them differently because of something it does not show: whether the matrix has one eigenvector or two. Move E’s eigenvalue to −1-1 and the two would share a point on the plane while one has a logarithm and the other does not. The criterion is really about Jordan structure, and in dimensions higher than two a single point on the negative axis can hide Jordan blocks of several sizes, each of which must pair with another of the same size.

The pictures also show only real logarithms. Over the complex numbers every invertible matrix has a logarithm, because the complex exponential reaches every nonzero complex number, so the obstruction here is entirely about staying real. And the figures show two-by-two matrices, where the image of the exponential in the determinant-one group is drawn completely by one curve. For larger groups the image of the exponential is known in many cases and is a subtle object in general.

Still open: logarithms that must stay inside a set

Existence and uniqueness of real logarithms are settled completely by Culver’s theorem, and computing the principal logarithm is a solved numerical problem. What remains open are the versions in which the logarithm has to belong to a restricted set.

The embedding problem for Markov chains is the best known. For three states it was worked out in detail only recently, and for four or more there is no known finite test that decides whether a given transition matrix is the exponential of some rate matrix; the logarithms can be infinite in number, and it is not known in general how to search them for one that is valid. The corresponding question for groups — which elements of a given Lie group lie on some one-parameter subgroup, so that the exponential reaches them — has complete answers for many classical groups and none that covers every case at once. The same question for a single matrix but a restricted logarithm — a real logarithm that is also symmetric, or that commutes with a given matrix — is where the answers become case by case, since the eigenvalue picture decides existence and says nothing about which of a continuum of logarithms has the extra property.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

DeterminantEigenvalueLogarithmMatrix exponentialPower seriesRotationTrace