Algebra

Regression is a square completed halfway

A bell in two variables, sliced at a fixed x, is a bell in y — and completing the square in y alone says where it is centred: at ρx, on a line flatter than the ellipse's own axis. That gap is regression to the mean, and the same completion says it runs backwards in time just as well as forwards.

Worth reading first: One number under every bell · The signs no completion can change.

Francis Galton measured the heights of 928 adult children and their parents in the 1880s and found something he did not expect. Children of unusually tall parents were tall, but on average less tall than their parents; children of unusually short parents were short, but less short. He called it “regression towards mediocrity” and spent some time looking for a biological mechanism that would pull each generation back towards the average.

There is none, and the clearest way to see that there is none is to notice that Galton’s data show the same thing in reverse. Parents of unusually tall children were tall, and on average less tall than their children. If regression were a force acting on each generation, it would have to act backwards in time as well. It is not a force. It is what completing the square does to a bell in two variables when the square is completed in one variable and not the other.

A tilted bell, cut into slices

Take two quantities measured in units of their own spread, so each has mean nought and spread one, and suppose they vary together with correlation ρ\rho — a number between −1-1 and 11 that says how strongly one tracks the other. The bell-shaped joint distribution of such a pair has density proportional to

exp⁡ ⁣(−x2−2ρxy+y22(1−ρ2)).\exp\!\left(-\frac{x^2 - 2\rho xy + y^2}{2(1 - \rho^2)}\right).

Its contours are ellipses, tilted along the diagonal when ρ\rho is positive, and the tilt is the whole of the correlation: at ρ=0\rho = 0 the contours are circles, and as ρ\rho approaches 11 they collapse towards the diagonal line.

Now fix xx — look only at the pairs whose first coordinate is some particular value — and ask how yy is distributed among them. That is a vertical slice through the hill, and the question is answered by treating xx as a constant and completing the square in yy alone. The exponent’s numerator y2−2ρxy+x2y^2 - 2\rho xy + x^2 completes, with xx held fixed, to

(y−ρx)2+(1−ρ2)x2,(y - \rho x)^2 + (1 - \rho^2)x^2,

so the slice is proportional to exp⁡ ⁣(−(y−ρx)22(1−ρ2))\exp\!\left(-\dfrac{(y - \rho x)^2}{2(1 - \rho^2)}\right) times a constant that depends on xx but not on yy. That is a bell in yy, centred at ρx\rho x, with spread 1−ρ2\sqrt{1 - \rho^2}.

Every slice of a tilted bell is a bell, centred off the axis. Contour ellipses of a two-variable bell with correlation 0.6, five vertical slices drawn as bells, their centres on the line y = 0.6x, and the ellipses' long axis on the diagonal y = x.
Fig. 1 The bell with correlation 0.6: three contour ellipses, and vertical slices at five values of xx, each drawn sideways as the bell it is. Each slice was integrated to find its centre and spread, and every centre lies on the line y=0.6xy = 0.6x with every spread equal to 0.8 — not on the ellipses’ long axis, the diagonal.

The figure’s five slices were integrated numerically rather than read off the formula, and each is centred at exactly 0.6x0.6x with spread exactly 0.80.8. The centres lie on a straight line of slope 0.60.6. The long axis of every ellipse, meanwhile, is the diagonal y=xy = x. Those two lines are different, and the difference between them is the phenomenon Galton found.

The line nobody expects

The natural guess, looking at a tilted ellipse, is that the typical yy for a given xx lies along the ellipse’s axis. It does not, and the figure shows why at the slice x=2x = 2. That slice crosses each ellipse at two points, and the ellipse bulges further above the diagonal on its far side — but what decides the centre of the slice is where the hill is highest along the vertical line, which is where the vertical line touches an ellipse rather than where it crosses the axis. A vertical line is tangent to an ellipse at the point where the ellipse is furthest to the side, and for a tilted ellipse that point is not on the long axis; it is on the line y=ρxy = \rho x, which runs flatter.

So a member of the population two spreads above average in xx is, on average, only 1.21.2 spreads above average in yy, when ρ=0.6\rho = 0.6. The long axis would have said two. Nothing about the individual has been pulled anywhere; the statement is about where, among everybody two spreads up in xx, the bulk of them sit in yy, and the bulk sits closer to the middle because a typical member of that group is there partly because of whatever makes xx large without making yy large.

The completed square says this without any story. The term (y−ρx)2(y - \rho x)^2 puts the centre at ρx\rho x; the term (1−ρ2)x2(1 - \rho^2)x^2, which does not involve yy at all, is the bell’s profile in xx alone, which is why it can be pulled out as a constant. That is the same split the earlier completion made when it moved a one-variable bell sideways — a part that locates the bell and a part that pays for the move — with the difference that here the move depends on the other variable.

Two regression lines, not one

The square can equally be completed in xx with yy held fixed, and the exponent is symmetric in xx and yy, so the answer is the mirror image: a horizontal slice at height yy is a bell in xx centred at ρy\rho y. Drawn in the same plane, the line of horizontal-slice centres is x=ρyx = \rho y, or y=x/ρy = x/\rho — steeper than the diagonal, because it is the first line reflected in it.

Two regression lines and the axis between them. Three panels for correlations 0.3, 0.6, 0.9: the contour ellipses, the line of vertical-slice centres, the line of horizontal-slice centres, and the long axis between them. The regression lines approach the axis as the correlation grows.
Fig. 2 Three correlations. In each panel, the line through the vertical slices’ centres (slope ρ\rho), the line through the horizontal slices’ centres (slope 1/ρ1/\rho), and the ellipses’ long axis between them. The angle between the two regression lines is 57° at correlation 0.3, 28° at 0.6 and 6.0° at 0.9; they meet the axis only at perfect correlation.

These are the two regression lines: yy regressed on xx, and xx regressed on yy. They cross at the centre, and they straddle the long axis symmetrically, one flatter and one steeper. At correlation 0.90.9 they are within six degrees of each other and of the axis; at correlation 0.30.3 they are fifty-seven degrees apart. At correlation nought they are the two coordinate axes themselves, since knowing xx then says nothing about yy and the best guess of yy is its overall average, nought, whatever xx is.

Which line is “the” line depends entirely on which variable is being predicted from which, and that is a decision about the question, not about the data. Asking for the typical yy given xx completes the square in yy; asking for the typical xx given yy completes in xx. The ellipse’s long axis answers neither question. It answers a third — which single line the points lie closest to, measured perpendicularly — and it is the line principal components analysis returns, the eigenvector direction from the completion along perpendicular axes.

Galton, both ways round

The simplest demonstration that nothing is pulling is to run Galton’s comparison in both directions on the same data.

Children of tall parents, parents of tall children. A scatter of 1000 simulated parent and child heights with correlation 0.5. Children averaged by parent's height and parents averaged by child's height both lie on lines of slope about 0.5, flatter than the diagonal.
Fig. 3 A thousand simulated parent–child pairs with heights of mean 170 cm and spread 7 cm and correlation 0.5. Filled dots average the children of parents near each height; open dots average the parents of children near each height. Both runs are flatter than the diagonal, in their own direction: the children of 186 cm parents average 178.1 cm, and the parents of 186 cm children average 178.7 cm.

The filled dots are the children’s average height among parents near each height, and they climb with slope about one half: parents eleven centimetres above average have children about five and a half above. The open dots go the other way round — the parents’ average height among children near each height, plotted with the parent’s height across and the child’s up, so that a slope of one half appears as a line steeper than the diagonal. Both runs regress. Tall parents have less tall children, and tall children have less tall parents.

The second statement is as true as the first and plainly not a fact about inheritance, since it would have parents shrinking towards the average of their own children. What both statements describe is selection. Selecting the tallest parents selects people who are tall for two kinds of reason — reasons they pass on, and reasons they do not, such as nourishment or measurement or chance — and their children inherit only the first kind. Selecting the tallest children selects in the same way from the other end. The correlation, one half here, measures how much of the variation is the shared kind, and the completed square turns that into a slope.

Tested twice, and nothing changed

The commonest way to be fooled by regression is to select on one measurement and be surprised by the next.

The same people, tested twice. Ten pairs of bars, one pair per decile of a first test: the decile's average on the first test and on a second test. Every second-test bar is shorter, by the factor 0.667 that a tick across it predicts.
Fig. 4 Twenty thousand simulated candidates, each with a fixed true skill and two test scores that each add independent noise, so the two scores correlate at two thirds. Candidates are grouped by their decile on the first test. In every decile the second test’s average is closer to the middle, and the tick across each green bar — the first average times two thirds — is where the completed square says it should land.

The candidates in the simulation do not change between the two tests. Each has a skill, and each test reports that skill plus some noise, independently drawn. The top tenth on the first test average 1.761.76 spreads above the middle and 1.191.19 on the second; the bottom tenth rise from −1.75-1.75 to −1.17-1.17. Every decile moves a third of the way to the middle, which is 1−ρ1 - \rho with ρ=23\rho = \tfrac23.

Anybody running an intervention on the top or bottom tenth would see that movement and be tempted to credit it. A remedial class for the lowest scorers would appear to work; a reward for the highest scorers would appear to make them complacent; a speed camera installed where accidents were worst last year would appear to reduce them. In each case the regression would happen with no intervention at all, because the group was selected on a measurement that was extreme partly by luck, and luck does not repeat. The only defence is the one experimental design supplies — a comparison group selected the same way and left alone, which regresses by the same amount and so can be subtracted.

The completed square makes the prediction exact rather than qualitative. The figure’s ticks are not fitted to the bars. They are the first-test averages multiplied by the correlation, and the simulated second-test averages sit on them to within 0.050.05 in every decile. Regression to the mean is not a vague tendency; it is a slope.

A book-length mistake, and a flight instructor’s

The best-documented victim of the effect is an economist. In 1933 Horace Secrist published The Triumph of Mediocrity in Business, a long statistical study of department stores, railways and banks. He ranked firms by their profit ratios in one year, grouped them, and followed each group for a decade. The most profitable groups became less profitable and the least profitable became more, year after year, and Secrist concluded that competition was dragging every business towards a mediocre middle — a finding he took to be of the first importance for economic policy.

Harold Hotelling’s review called the conclusion mathematically trivial, and the quickest way to see why is to read the same tables with the years reversed: they would show the firms diverging from mediocrity as one went back in time, which no theory of competition predicts. Secrist had selected his groups on a noisy measurement in the first year and watched the noise fail to repeat. The convergence was the slope of a regression line and nothing else, and it would have appeared for any quantity with a correlation below one between years, whatever the economics. The book had been well received in print before the arithmetic was pointed out.

Daniel Kahneman tells a smaller version from teaching flight instructors in the Israeli air force. The instructors were sure that praise made cadets worse and criticism made them better, because a cadet praised for an excellent landing usually did worse on the next and a cadet shouted at for a terrible one usually did better. Both observations were correct and neither said anything about praise or shouting. An excellent landing is excellent partly by luck, the next landing draws fresh luck, and the second landing regresses whatever the instructor says. The tested-twice figure above is that classroom with the instructor removed, and the bars move all the same.

What makes the mistake so durable is that it is made with correct data. Secrist’s tables were accurate, the instructors’ memories were accurate, and the inference in each case followed from treating the line of slice centres as though it were the ellipse’s axis. The completed square is the correction, and it has to be applied deliberately, because nothing in the raw numbers announces that it is needed.

What is left in a slice

The other half of the completion, the spread 1−ρ2\sqrt{1 - \rho^2}, has a feature that is easy to pass over. It does not depend on xx.

Every vertical slice of the tilted bell has the same width. The slice through the centre and the slice two spreads out are the same bell, merely moved. In the hero figure the five slices are copies of one another, and that is not a property of the picture’s scale; it is the statement that the leftover term (1−ρ2)(1 - \rho^2) in the completed square carries no xx. The uncertainty in predicting yy from xx is the same wherever xx is, which is the property statisticians call homoscedasticity and which bell-shaped data have exactly.

The spread left in a slice is the other leg of a right triangle. Measured spreads of vertical slices against correlation, three slices per correlation, lying on the quarter circle √(1 − ρ²); one point is joined to the origin to show the right triangle with legs ρ and √(1 − ρ²).
Fig. 5 For seven correlations, the measured spread of yy among simulated pairs with xx near −1-1, near 00 and near 11 — three dots per correlation, which land on one another because the spread in a slice does not depend on the slice. The dots trace the quarter circle 1−ρ2\sqrt{1 - \rho^2}; one point is joined to the origin to show the right triangle it makes.

And the number itself, 1−ρ2\sqrt{1 - \rho^2}, is the second leg of a right triangle whose hypotenuse is one and whose first leg is ρ\rho. The variance of yy, one, splits into the part a line through xx accounts for, ρ2\rho^2, and the part left in the slice, 1−ρ21 - \rho^2, and the split is Pythagoras. That is not a metaphor. Measured in units of their spread, random quantities are vectors, the covariance of two of them is their dot product, and the correlation is the cosine of the angle between them. Predicting yy from xx is dropping a perpendicular from yy onto the line through xx — the nearest point of a flat thing — and the regression slope is the length of the shadow. The leftover is the perpendicular, and the perpendicular is the same length wherever along the line one stands.

So completing the square in one variable, conditioning on the other, and projecting onto a line are one operation described in three vocabularies. The algebra writes (y−ρx)2+(1−ρ2)x2(y - \rho x)^2 + (1 - \rho^2)x^2; the probability says “centred at ρx\rho x with spread 1−ρ2\sqrt{1 - \rho^2}”; the geometry says “shadow of length ρ\rho, perpendicular of length 1−ρ2\sqrt{1 - \rho^2}”. This is the connection this essay set out to find, and it is the surprising one: the most widely used procedure in empirical science — fitting a regression line — is the half-completed square of a quadratic exponent, and its central number, the proportion of variance explained, is the square of a cosine.

In more variables, and the filter that runs on it

The two-variable case generalises without change. A bell in nn variables has a quadratic form in its exponent; split the variables into those being observed and those being predicted, and complete the square in the predicted ones. The centre of the conditional bell is a linear function of the observed values — the regression — and the leftover form, which in matrix language is the Schur complement of the observed block, is the conditional uncertainty, again the same wherever the observation falls.

That sentence is the whole of the Kalman filter, the procedure that tracks a moving object from noisy measurements and that guided the Apollo spacecraft. At each step it has a bell describing its belief about the state, receives a measurement, and completes the square in the unknown state given the measurement: the new centre is a weighted compromise, the new spread is the Schur complement and shrinks, and neither step involves anything but the algebra above. Combining two bells into one was the one-variable shadow of it; conditioning in many variables is the full version, and it is cheap because a completed square is.

Data that were simulated from a bell

Every dataset here was simulated from a bell, so the slices had to be bells. Real data are not bell-shaped, and two of the clean statements above depend on the shape. The centres of the slices lie on a straight line only when the joint distribution is a bell, or something with the same symmetry; for other shapes the curve of slice centres can bend. And the slices have equal spread only for bells; real data often fan out, with more scatter in yy at large xx than at small. What survives in general is weaker: the best straight-line prediction of yy from xx always has slope ρ\rho times the ratio of spreads, and selecting on extreme xx always produces less extreme yy on average whenever ρ\rho is less than one. Regression to the mean is universal; its exact form here is a property of the bell.

The figures show that regression happens; they cannot show why any particular change happened. A tested-twice decile whose average moves towards the middle is moving as the completed square predicts, and a real intervention applied to the same people might also be moving them. The pictures cannot separate the two, and the separation is precisely what a control group is for. No amount of drawing the slices more carefully replaces one.

The heights are invented. The Galton figure uses simulated families with correlation one half, which is close to what Galton measured for the average of two parents against a child, but the dots are not his data. What the figure demonstrates is a property of the model; that real heights behave this way is Galton’s finding, reported here rather than checked.

Still open: the regression nobody can see coming

Regression to the mean is completely understood as mathematics and still routinely missed in practice, and the open questions are mostly about measurement rather than algebra. How large the regression in a given study will be depends on the correlation between the selection measurement and the follow-up, which depends on how noisy the measurement is — and that is often known only after the follow-up, which is too late to have designed around it. Methods for estimating it in advance from a single measurement, using repeated measurements on a subsample or a model of the noise, are in use, and how well they work outside the bell-shaped case is studied case by case.

There is also a structural question in many variables that the completed square makes sharp. Conditioning on several measurements at once completes the square in all the unobserved variables together, and the resulting regression coefficients can have signs opposite to the simple correlations — a quantity positively correlated with an outcome can predict it negatively once another variable is held fixed. Which patterns of sign reversal are possible for a given set of correlations is answered by the signs of a Schur complement, and the law of inertia controls how many of them can be negative; but reading off, from data, which reversal is real and which is an artefact of what was measured remains the hard part of every observational study.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

Completing the squareConditional probabilityCorrelationExpectationNormal distributionProjectionPythagorean theoremQuadratic formSimulationVariance