Regression is a square completed halfway
Worth reading first: One number under every bell · The signs no completion can change.
Francis Galton measured the heights of 928 adult children and their parents in the 1880s and found something he did not expect. Children of unusually tall parents were tall, but on average less tall than their parents; children of unusually short parents were short, but less short. He called it “regression towards mediocrity” and spent some time looking for a biological mechanism that would pull each generation back towards the average.
There is none, and the clearest way to see that there is none is to notice that Galton’s data show the same thing in reverse. Parents of unusually tall children were tall, and on average less tall than their children. If regression were a force acting on each generation, it would have to act backwards in time as well. It is not a force. It is what completing the square does to a bell in two variables when the square is completed in one variable and not the other.
A tilted bell, cut into slices
Take two quantities measured in units of their own spread, so each has mean nought and spread one, and suppose they vary together with correlation — a number between and that says how strongly one tracks the other. The bell-shaped joint distribution of such a pair has density proportional to
Its contours are ellipses, tilted along the diagonal when is positive, and the tilt is the whole of the correlation: at the contours are circles, and as approaches they collapse towards the diagonal line.
Now fix — look only at the pairs whose first coordinate is some particular value — and ask how is distributed among them. That is a vertical slice through the hill, and the question is answered by treating as a constant and completing the square in alone. The exponent’s numerator completes, with held fixed, to
so the slice is proportional to times a constant that depends on but not on . That is a bell in , centred at , with spread .
The figure’s five slices were integrated numerically rather than read off the formula, and each is centred at exactly with spread exactly . The centres lie on a straight line of slope . The long axis of every ellipse, meanwhile, is the diagonal . Those two lines are different, and the difference between them is the phenomenon Galton found.
The line nobody expects
The natural guess, looking at a tilted ellipse, is that the typical for a given lies along the ellipse’s axis. It does not, and the figure shows why at the slice . That slice crosses each ellipse at two points, and the ellipse bulges further above the diagonal on its far side — but what decides the centre of the slice is where the hill is highest along the vertical line, which is where the vertical line touches an ellipse rather than where it crosses the axis. A vertical line is tangent to an ellipse at the point where the ellipse is furthest to the side, and for a tilted ellipse that point is not on the long axis; it is on the line , which runs flatter.
So a member of the population two spreads above average in is, on average, only spreads above average in , when . The long axis would have said two. Nothing about the individual has been pulled anywhere; the statement is about where, among everybody two spreads up in , the bulk of them sit in , and the bulk sits closer to the middle because a typical member of that group is there partly because of whatever makes large without making large.
The completed square says this without any story. The term puts the centre at ; the term , which does not involve at all, is the bell’s profile in alone, which is why it can be pulled out as a constant. That is the same split the earlier completion made when it moved a one-variable bell sideways — a part that locates the bell and a part that pays for the move — with the difference that here the move depends on the other variable.
Two regression lines, not one
The square can equally be completed in with held fixed, and the exponent is symmetric in and , so the answer is the mirror image: a horizontal slice at height is a bell in centred at . Drawn in the same plane, the line of horizontal-slice centres is , or — steeper than the diagonal, because it is the first line reflected in it.
These are the two regression lines: regressed on , and regressed on . They cross at the centre, and they straddle the long axis symmetrically, one flatter and one steeper. At correlation they are within six degrees of each other and of the axis; at correlation they are fifty-seven degrees apart. At correlation nought they are the two coordinate axes themselves, since knowing then says nothing about and the best guess of is its overall average, nought, whatever is.
Which line is “the” line depends entirely on which variable is being predicted from which, and that is a decision about the question, not about the data. Asking for the typical given completes the square in ; asking for the typical given completes in . The ellipse’s long axis answers neither question. It answers a third — which single line the points lie closest to, measured perpendicularly — and it is the line principal components analysis returns, the eigenvector direction from the completion along perpendicular axes.
Galton, both ways round
The simplest demonstration that nothing is pulling is to run Galton’s comparison in both directions on the same data.
The filled dots are the children’s average height among parents near each height, and they climb with slope about one half: parents eleven centimetres above average have children about five and a half above. The open dots go the other way round — the parents’ average height among children near each height, plotted with the parent’s height across and the child’s up, so that a slope of one half appears as a line steeper than the diagonal. Both runs regress. Tall parents have less tall children, and tall children have less tall parents.
The second statement is as true as the first and plainly not a fact about inheritance, since it would have parents shrinking towards the average of their own children. What both statements describe is selection. Selecting the tallest parents selects people who are tall for two kinds of reason — reasons they pass on, and reasons they do not, such as nourishment or measurement or chance — and their children inherit only the first kind. Selecting the tallest children selects in the same way from the other end. The correlation, one half here, measures how much of the variation is the shared kind, and the completed square turns that into a slope.
Tested twice, and nothing changed
The commonest way to be fooled by regression is to select on one measurement and be surprised by the next.
The candidates in the simulation do not change between the two tests. Each has a skill, and each test reports that skill plus some noise, independently drawn. The top tenth on the first test average spreads above the middle and on the second; the bottom tenth rise from to . Every decile moves a third of the way to the middle, which is with .
Anybody running an intervention on the top or bottom tenth would see that movement and be tempted to credit it. A remedial class for the lowest scorers would appear to work; a reward for the highest scorers would appear to make them complacent; a speed camera installed where accidents were worst last year would appear to reduce them. In each case the regression would happen with no intervention at all, because the group was selected on a measurement that was extreme partly by luck, and luck does not repeat. The only defence is the one experimental design supplies — a comparison group selected the same way and left alone, which regresses by the same amount and so can be subtracted.
The completed square makes the prediction exact rather than qualitative. The figure’s ticks are not fitted to the bars. They are the first-test averages multiplied by the correlation, and the simulated second-test averages sit on them to within in every decile. Regression to the mean is not a vague tendency; it is a slope.
A book-length mistake, and a flight instructor’s
The best-documented victim of the effect is an economist. In 1933 Horace Secrist published The Triumph of Mediocrity in Business, a long statistical study of department stores, railways and banks. He ranked firms by their profit ratios in one year, grouped them, and followed each group for a decade. The most profitable groups became less profitable and the least profitable became more, year after year, and Secrist concluded that competition was dragging every business towards a mediocre middle — a finding he took to be of the first importance for economic policy.
Harold Hotelling’s review called the conclusion mathematically trivial, and the quickest way to see why is to read the same tables with the years reversed: they would show the firms diverging from mediocrity as one went back in time, which no theory of competition predicts. Secrist had selected his groups on a noisy measurement in the first year and watched the noise fail to repeat. The convergence was the slope of a regression line and nothing else, and it would have appeared for any quantity with a correlation below one between years, whatever the economics. The book had been well received in print before the arithmetic was pointed out.
Daniel Kahneman tells a smaller version from teaching flight instructors in the Israeli air force. The instructors were sure that praise made cadets worse and criticism made them better, because a cadet praised for an excellent landing usually did worse on the next and a cadet shouted at for a terrible one usually did better. Both observations were correct and neither said anything about praise or shouting. An excellent landing is excellent partly by luck, the next landing draws fresh luck, and the second landing regresses whatever the instructor says. The tested-twice figure above is that classroom with the instructor removed, and the bars move all the same.
What makes the mistake so durable is that it is made with correct data. Secrist’s tables were accurate, the instructors’ memories were accurate, and the inference in each case followed from treating the line of slice centres as though it were the ellipse’s axis. The completed square is the correction, and it has to be applied deliberately, because nothing in the raw numbers announces that it is needed.
What is left in a slice
The other half of the completion, the spread , has a feature that is easy to pass over. It does not depend on .
Every vertical slice of the tilted bell has the same width. The slice through the centre and the slice two spreads out are the same bell, merely moved. In the hero figure the five slices are copies of one another, and that is not a property of the picture’s scale; it is the statement that the leftover term in the completed square carries no . The uncertainty in predicting from is the same wherever is, which is the property statisticians call homoscedasticity and which bell-shaped data have exactly.
And the number itself, , is the second leg of a right triangle whose hypotenuse is one and whose first leg is . The variance of , one, splits into the part a line through accounts for, , and the part left in the slice, , and the split is Pythagoras. That is not a metaphor. Measured in units of their spread, random quantities are vectors, the covariance of two of them is their dot product, and the correlation is the cosine of the angle between them. Predicting from is dropping a perpendicular from onto the line through — the nearest point of a flat thing — and the regression slope is the length of the shadow. The leftover is the perpendicular, and the perpendicular is the same length wherever along the line one stands.
So completing the square in one variable, conditioning on the other, and projecting onto a line are one operation described in three vocabularies. The algebra writes ; the probability says “centred at with spread ”; the geometry says “shadow of length , perpendicular of length ”. This is the connection this essay set out to find, and it is the surprising one: the most widely used procedure in empirical science — fitting a regression line — is the half-completed square of a quadratic exponent, and its central number, the proportion of variance explained, is the square of a cosine.
In more variables, and the filter that runs on it
The two-variable case generalises without change. A bell in variables has a quadratic form in its exponent; split the variables into those being observed and those being predicted, and complete the square in the predicted ones. The centre of the conditional bell is a linear function of the observed values — the regression — and the leftover form, which in matrix language is the Schur complement of the observed block, is the conditional uncertainty, again the same wherever the observation falls.
That sentence is the whole of the Kalman filter, the procedure that tracks a moving object from noisy measurements and that guided the Apollo spacecraft. At each step it has a bell describing its belief about the state, receives a measurement, and completes the square in the unknown state given the measurement: the new centre is a weighted compromise, the new spread is the Schur complement and shrinks, and neither step involves anything but the algebra above. Combining two bells into one was the one-variable shadow of it; conditioning in many variables is the full version, and it is cheap because a completed square is.
Data that were simulated from a bell
Every dataset here was simulated from a bell, so the slices had to be bells. Real data are not bell-shaped, and two of the clean statements above depend on the shape. The centres of the slices lie on a straight line only when the joint distribution is a bell, or something with the same symmetry; for other shapes the curve of slice centres can bend. And the slices have equal spread only for bells; real data often fan out, with more scatter in at large than at small. What survives in general is weaker: the best straight-line prediction of from always has slope times the ratio of spreads, and selecting on extreme always produces less extreme on average whenever is less than one. Regression to the mean is universal; its exact form here is a property of the bell.
The figures show that regression happens; they cannot show why any particular change happened. A tested-twice decile whose average moves towards the middle is moving as the completed square predicts, and a real intervention applied to the same people might also be moving them. The pictures cannot separate the two, and the separation is precisely what a control group is for. No amount of drawing the slices more carefully replaces one.
The heights are invented. The Galton figure uses simulated families with correlation one half, which is close to what Galton measured for the average of two parents against a child, but the dots are not his data. What the figure demonstrates is a property of the model; that real heights behave this way is Galton’s finding, reported here rather than checked.
Still open: the regression nobody can see coming
Regression to the mean is completely understood as mathematics and still routinely missed in practice, and the open questions are mostly about measurement rather than algebra. How large the regression in a given study will be depends on the correlation between the selection measurement and the follow-up, which depends on how noisy the measurement is — and that is often known only after the follow-up, which is too late to have designed around it. Methods for estimating it in advance from a single measurement, using repeated measurements on a subsample or a model of the noise, are in use, and how well they work outside the bell-shaped case is studied case by case.
There is also a structural question in many variables that the completed square makes sharp. Conditioning on several measurements at once completes the square in all the unobserved variables together, and the resulting regression coefficients can have signs opposite to the simple correlations — a quantity positively correlated with an outcome can predict it negatively once another variable is held fixed. Which patterns of sign reversal are possible for a given set of correlations is answered by the signs of a Schur complement, and the law of inertia controls how many of them can be negative; but reading off, from data, which reversal is real and which is an artefact of what was measured remains the hard part of every observational study.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- An average that never settles — both name expectation, normal distribution, variance
- Counting a population by its repeats — both name expectation, simulation, variance
- How far from the average a thing can be — both name expectation, normal distribution, variance
- How fast the bell arrives — both name expectation, normal distribution, variance
- The average settles and the wobble does not — both name expectation, normal distribution, variance
- A coin that lets the first player win — both name conditional probability, expectation
Named objects
A dashed tag is an object no other essay names yet.
Completing the squareConditional probabilityCorrelationExpectationNormal distributionProjectionPythagorean theoremQuadratic formSimulationVariance