Analysis

The slope of the mirror image

Undoing a function is reflecting its graph in the diagonal, and a reflection turns a slope into its reciprocal. That single observation supplies the derivative of every inverse — the logarithm, the roots, the inverse trigonometric functions — without differentiating any of them.
18 min read 6 figures The same thing twiceOne point away

Worth reading first: The slope of a single point · The curve that is its own slope.

Undoing a function is a geometric operation: swap the two coordinates, which reflects the graph in the line y=xy = x. Everything about the inverse follows from what that reflection does, and what it does to a slope is turn it upside down.

eˣ and its inverse, reflected in the diagonalA curve, the line y = x, and the curve reflected in it — which is the graph of the inverse function. Tangents are drawn at matched pairs of points, and the two slopes at each pair multiply to one.-101234-101234xyslope 1.73slope 0.58its inverseeˣ and its inverse, which is the same curve reflected in the dashed diagonal; the inverseis computed here by bisection, so its slope is measured rather than assumedat x = 0.55 the slope is 1.733 and the mirrored slope is 0.577; at x = 0.89 the slope is2.438 and the mirrored slope is 0.410; at x = 1.23 the slope is 3.430 and the mirroredslope is 0.292 — each pair multiplies to one
Fig. 1 A curve and its reflection in the dashed diagonal, which is the graph of its inverse. At each matched pair of points the two tangents are drawn, and the slopes multiply to one — measured, not assumed, because the inverse here is computed by bisection rather than from a formula.

A line of slope mm through the origin reflects to a line of slope 1/m1/m, since reflection swaps rise and run. A tangent line reflects to a tangent line, because reflection is a rigid motion and cannot turn a touching line into a crossing one. Put the two together and

(f1)(y)=1f(x),y=f(x),\left(f^{-1}\right)'(y) = \frac{1}{f'(x)}, \qquad y = f(x),

which is the rule, complete, with no calculation in it.

Why the figure is evidence

There is an easy way to draw this picture that proves nothing: reflect the curve, reflect its tangent line, and label the reflected slope 1/m1/m. That would be a drawing of the conclusion.

Instead the inverse in the figure is computed. For each value yy on the horizontal axis, a bisection search finds the xx with f(x)=yf(x) = y, to two hundred halvings; the curve drawn is the set of those points. The slope of the inverse at a marked point is then measured by a difference quotient on that computed function, and compared against the reciprocal of the original’s slope. The generator refuses to draw if the product of the two is not one to within a millionth.

So the reflection is a claim the picture tests rather than a construction it performs, which is the distinction this collection’s figures are built around.

The argument, in one line each way

Two derivations, and the second is the one that generalises.

By the chain rule. Differentiate f1(f(x))=xf^{-1}(f(x)) = x with respect to xx: the left side gives (f1)(f(x))f(x)(f^{-1})'(f(x)) \cdot f'(x), the right side gives 11. Rearranging is the rule. This is short and it assumes what it is proving — that the inverse is differentiable at all — since the chain rule may only be applied to a function already known to have a derivative.

By the difference quotient. The quotient for the inverse at y=f(x)y = f(x) is

f1(y+k)f1(y)k=hf(x+h)f(x),\frac{f^{-1}(y + k) - f^{-1}(y)}{k} = \frac{h}{f(x+h) - f(x)},

where hh is defined by y+k=f(x+h)y + k = f(x + h). It is the reciprocal of the original’s difference quotient, exactly, for every non-zero step. Taking the limit needs one fact: that h0h \to 0 when k0k \to 0, which is the continuity of the inverse. Given that, the limit of a reciprocal is the reciprocal of the limit, provided the limit is not zero.

The second version proves differentiability rather than assuming it, and its hypotheses are exactly the ones the rule needs: the inverse must be continuous, and the original’s derivative must be non-zero. The next section is what happens when the second fails.

The point where it fails

The reciprocal is undefined at zero, and a horizontal tangent is where the rule stops.

x³ and its inverse, reflected in the diagonalA curve, the line y = x, and the curve reflected in it — which is the graph of the inverse function. Tangents are drawn at matched pairs of points, and the two slopes at each pair multiply to one.-1.5-1-0.500.511.5-1.5-1-0.500.511.5xyslope 0.75slope 1.33its inversex³ and its inverse, which is the same curve reflected in the dashed diagonal; the inverseis computed here by bisection, so its slope is measured rather than assumedat x = 0.50 the slope is 0.750 and the mirrored slope is 1.333; at x = 0.79 the slope is1.884 and the mirrored slope is 0.531; at x = 1.08 the slope is 3.531 and the mirroredslope is 0.283 — each pair multiplies to one
Fig. 2 The cube and its inverse. At the origin the cube’s tangent is horizontal, so the cube root’s is vertical: the inverse is continuous there, has no derivative, and the reflection makes the reason obvious.

The cube has derivative zero at the origin. Its inverse, the cube root, is perfectly well defined and continuous there — and its graph has a vertical tangent, so no slope. Nothing has gone wrong with the function; the reflection has turned a horizontal line into a vertical one, and a vertical line is not the graph of a linear function.

This is the honest content of the non-zero derivative hypothesis. It is not a technical condition to be waved past: at a point where f=0f' = 0 the inverse either fails to exist locally, because the function turns around, or exists and is not differentiable. Both cases occur and both are visible in the picture.

The general shape is that dividing by a quantity that can vanish always requires a hypothesis, and the hypothesis always corresponds to something geometric. Here it is the one point where a construction breaks down, which is a recurring theme in this collection and is nearly always worth locating precisely rather than excluding by fiat.

What has to be true before any of it applies

The rule is stated for the inverse, and a function has one only under conditions that the picture assumes silently.

A function has an inverse on a stretch exactly when it is one-to-one there, and for a continuous function on an interval that is the same as being strictly monotone — strictly increasing or strictly decreasing throughout. So the first question about any inverse is where the function turns around, and the answer bounds every stretch on which the rule can be used.

A derivative that is positive throughout an interval guarantees strict increase, which is the usual way the condition is checked. The converse is not quite true: a function can be strictly increasing with a derivative that vanishes at isolated points, as the cube does at the origin, and those points are exactly the ones where the inverse exists and is not differentiable.

So there are three levels, and they are worth keeping distinct. Strictly monotone gives an inverse. Continuous and strictly monotone gives a continuous inverse. Differentiable with non-vanishing derivative gives a differentiable inverse — and each level costs one more hypothesis than the last.

Everything it gives for free

The rule’s value is that it differentiates functions that are hard to differentiate directly, by differentiating their easy inverses instead.

eˣ and its tangent linesThe exponential curve with tangent lines at several points; at each point the slope equals the height.-2-1.5-1-0.50.511.52123456xyheight 0.37slope 0.37height 1.00slope 1.00height 2.72slope 2.72height 4.95slope 4.95
Fig. 3 The exponential, which is its own slope. Reflecting it gives the logarithm, and the rule then gives the logarithm’s derivative with no work at all.

The logarithm. The exponential is its own derivative, so at the point (x,ex)(x, e^x) the slope is exe^x; the reflected point is (ex,x)(e^x, x) and the reflected slope is 1/ex1/e^x. Writing y=exy = e^x, the logarithm’s derivative at yy is 1/y1/y. That is the single most useful derivative in the subject, obtained by turning a picture on its side.

The roots. The nn-th power has derivative nxn1n x^{n-1}, so the nn-th root has derivative 1/(nxn1)1/(n x^{n-1}) at the matching point — which, rewritten in terms of y=xny = x^n, is 1ny1/n1\frac{1}{n} y^{1/n - 1}. The power rule for fractional exponents therefore needs no separate proof: it is the whole-number rule reflected.

The inverse trigonometric functions. Sine has derivative cosine; at the matched point, cosine is 1y2\sqrt{1 - y^2}; so the inverse sine has derivative 1/1y21/\sqrt{1-y^2}. The square root that appears out of nowhere in the standard formula is the Pythagorean relation between the two coordinates of a point on the unit circle, and the reflection is where it enters.

In each case the derivative of a function nobody wants to attack directly comes from the derivative of one that is easy, plus a reflection. That is a fair definition of a good technique.

Reading the same picture as a rate of exchange

There is a second way to read the reflection that makes the reciprocal feel inevitable rather than surprising.

x² and its inverse, reflected in the diagonalA curve, the line y = x, and the curve reflected in it — which is the graph of the inverse function. Tangents are drawn at matched pairs of points, and the two slopes at each pair multiply to one.00.511.522.500.511.522.5xyslope 1.40slope 0.71its inversex² and its inverse, which is the same curve reflected in the dashed diagonal; the inverseis computed here by bisection, so its slope is measured rather than assumedat x = 0.70 the slope is 1.400 and the mirrored slope is 0.714; at x = 1.33 the slope is2.659 and the mirrored slope is 0.376 — each pair multiplies to one
Fig. 4 The squaring function and the square root, in the same window. Where the square is steep the root is shallow, and the two effects are the same effect described from the two sides.

A derivative is a rate of exchange: how much output per unit of input. The inverse asks the reverse question — how much input per unit of output — and a rate of exchange read backwards is its reciprocal, exactly as a price in one currency per unit of another inverts when the currencies are swapped.

Read that way, the rule needs no geometry. What the geometry adds is the reason the reciprocal is taken at the matching point rather than at the same coordinate: the reflection sends (x,f(x))(x, f(x)) to (f(x),x)(f(x), x), so the point where the inverse’s rate is being measured has the original’s output as its input. Getting that wrong is the standard error with this rule, and it is the one thing about it that the picture makes impossible to get wrong.

sin x and its inverse, reflected in the diagonalA curve, the line y = x, and the curve reflected in it — which is the graph of the inverse function. Tangents are drawn at matched pairs of points, and the two slopes at each pair multiply to one.-1.5-1-0.500.511.5-1.5-1-0.500.511.5xyslope 0.92slope 1.09sin xits inversesin x and its inverse, which is the same curve reflected in the dashed diagonal; theinverse is computed here by bisection, so its slope is measured rather than assumedat x = 0.40 the slope is 0.921 and the mirrored slope is 1.086; at x = 1.16 the slope is0.401 and the mirrored slope is 2.496 — each pair multiplies to one
Fig. 5 Sine and its inverse on the stretch where sine increases. The restriction is not a technicality: sine is not one-to-one, and an inverse exists only after a branch has been chosen, which the window here does silently.

The third figure carries a warning worth stating plainly. Most functions worth inverting are not one-to-one, and the inverse exists only after a stretch has been selected on which the function increases. Which stretch is a convention — the inverse sine takes values between π/2-\pi/2 and π/2\pi/2 because somebody chose that — and the derivative formula holds on the chosen branch and says nothing about the others.

The same rule in more than one dimension

The reflection argument has a counterpart for maps of several variables, and it is where the rule becomes a theorem with content.

A curved map of the plane, and the flat one that fits it at a pointThe map (x + 0.55y², y − 0.45x²) carrying a small square patch of grid. Beside it, the image of the same patch under the linear map given by the matrix of partial derivatives, drawn dashed on top of the curved image.the patch(0.7, 0.6)its image, and the flat fit10.66-0.631the matrix of partial derivativesit multiplies area by 1.416the smallest patch drawn, by 1.416patchwidest gap0.70008.71e-20.35002.18e-20.17505.44e-30.08751.36e-30.04373.40e-4(x + 0.55y², y − 0.45x²) at (0.7, 0.6), on a patch 0.70 across. The straight grid on the left is carried over by the map on the right, where the dashed gridis the image under the one linear map that matches it at the pointthe widest gap between the two is 0.0871 here and 3.40e-4 on a patch 16 times smaller — a quarter at each halving, while the patch itself halves
Fig. 6 A curved map of the plane and the linear map that best fits it near a point. Inverting the map near that point corresponds to inverting the matrix — and it is possible exactly when the matrix is invertible.

If ff is a map of the plane with derivative the matrix AA at a point, and AA is invertible, then ff has a local inverse near that point and its derivative is A1A^{-1}. That is the inverse function theorem, and the one-dimensional statement is the case where AA is a single number: invertible means non-zero, and the inverse matrix is the reciprocal.

Two things change in the general case. The condition non-zero becomes non-zero determinant, which is a genuine computation rather than an inspection. And the conclusion becomes local in a way it was not before: a map of the plane can have invertible derivative everywhere and still fail to be one-to-one globally — the map sending a point at angle θ\theta to the point at angle 2θ2\theta at the same radius does exactly that, wrapping the plane twice around itself.

So on the line, an everywhere-positive derivative gives a global inverse; in the plane, an everywhere-invertible derivative gives only local ones. The reflection picture is honest about which of the two it shows: it shows a neighbourhood.

What the reflection preserves, and what it destroys

Reflecting a graph is a rigid motion of the plane, so a good deal survives it and a good deal does not, and sorting the two is worth doing once.

Preserved: tangency, continuity, and the order of contact. A line touching a curve without crossing still does after reflection, so a tangent stays a tangent. A curve with no breaks reflects to a curve with no breaks, provided the function is one-to-one on the stretch drawn. And a curve that is approximated by its tangent to second order still is.

Preserved with a twist: convexity. Reflecting turns a curve that bends upward into one that bends downward, for an increasing function — the square is convex, the square root concave. That is not obvious in advance and it falls out of differentiating the rule a second time, where a minus sign appears. Anybody using the reflection to guess a second derivative should expect the sign to flip.

Destroyed: everything about the vertical direction as such. A horizontal tangent becomes vertical, a bounded function becomes one with a bounded domain, and the point at which the function was flattest becomes the point at which its inverse is steepest. Any argument that relies on the vertical axis meaning something different from the horizontal one does not survive the reflection, and the whole trick is that a derivative is precisely such an argument — which is why it inverts rather than staying put.

What the picture cannot show

The window is square and it has to be, because a reflection in the diagonal is only a reflection when a unit is the same size on both axes. That constraint is invisible in the drawing and it is why the exponential is drawn on a range that cuts off most of its growth: on its natural domain the curve leaves the top of any square window almost immediately.

The vertical tangent in the second figure is drawn as a very steep line, because a vertical line cannot be a tangent in the sense the rest of the figure uses — it has no slope to print. The picture shows the slope growing without bound as the point approaches the origin, which is the honest form of the statement, and the word vertical is a limit rather than a description of anything drawn.

The bisection is also invisible, and it is what makes the figure an argument. What the reader sees is a curve; what produced it is two hundred halvings per plotted point, on a function the code cannot write down, and the slope printed beside it is a difference quotient taken on that computed curve. A figure that had reflected the tangent line instead would look identical and would establish nothing — which is a general hazard with pictures of inverses, since the correct drawing and the question-begging one are the same drawing.

And the matched pairs are three points on a continuum. The claim is that the slopes multiply to one at every matched pair, and three checks establish nothing; the figure’s assertion runs at each of them, and the argument that closes the gap is the difference-quotient calculation, which is four lines and has no picture.

The ladder from here

Below: the slope of a single point, which the whole rule is built on, and the exponential, which is the function the rule is most often used with. Sideways: the product rule and the derivative as a linear map, whose matrix inverse is this rule in several variables. Above: implicit differentiation, where a relation that defines no function at all is differentiated anyway, and the reflection argument becomes a statement about which variable is being solved for.

Two ways to get it wrong

Both standard errors with this rule are errors about where, and both are visible in the picture.

The first is evaluating the reciprocal at the wrong point: writing the inverse’s derivative at xx rather than at f(x)f(x). The reflection sends the point (x,f(x))(x, f(x)) to (f(x),x)(f(x), x), so the input to the inverse is the original’s output, and a formula that does not track that produces a function of the right shape evaluated in the wrong place.

The second is reflecting a stretch on which the function is not one-to-one. The result is a curve that fails the vertical line test — two output values above one input — which is not the graph of anything, and the slopes computed on it belong to two different branches.

Both are prevented by drawing the diagonal and following one point through it.

An operation, and what it does to a rate

The lasting point is that the rule was not calculated. Reflection is a rigid motion, tangency survives it, and slopes invert under it — three geometric facts, none of them about calculus, which between them determine the derivative of every inverse function there is.

The economy is worth measuring. Four derivatives that are usually proved one at a time — the logarithm, the fractional powers, the inverse sine, the inverse tangent — follow from four derivatives already known, with no limits taken and no new arguments. What replaced the work is a fact about a reflection.

That is worth generalising as a habit. Before differentiating something, it is worth asking what operation produced it and what that operation does to rates: multiplication makes relative rates add, composition makes them multiply, and reflection turns them upside down. Each of those is a statement about the operation rather than about any particular function, and each replaces a calculation with an observation.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

ApproximationContinuityDerivativeInverseLimitLogarithmReflectionTangency