Regression to the Mean
Regression to the mean describes the fact that particularly extreme measurements are usually followed by more ordinary ones. The reason is not some mysterious balancing force, but the element of chance contained in almost every measurement.
Almost every result we measure has two components. One component is genuine skill or a genuine underlying state. The other component is luck, day-to-day form, or simply measurement inaccuracy. An extremely high value arises especially often when both components happen to align favorably. The next time, the skill is still there, but the luck usually isn’t anymore. That’s why the second value typically lies closer to the average. This exact pattern is called regression to the mean.
The fallacy behind many success stories
This effect matters because it is constantly mistaken for a real causal effect. One example: a class writes a math test, and the worst five students get tutoring. On the next test, they do better. This looks like proof that the tutoring worked, but it is largely just the normal pull back toward the mean. Without a comparison group, the two cannot be separated.
The same thing happens in business and in medicine. A company changes its CEO after its worst quarter, and things improve afterward. A patient takes a remedy exactly when the pain is at its worst, and afterward it gets better. In both cases, an improvement would have been likely even without any intervention. Anyone who overlooks this systematically overestimates the effectiveness of interventions.
The effect also has an unpleasant flip side. Praise after an outstanding performance is often followed by a weaker result, while criticism after a slip-up is often followed by a better one. Many people wrongly conclude from this that criticism works and praise is harmful. The psychologist Daniel Kahneman described exactly this case involving flight instructors.
Why chance and skill diverge here
What matters is how strongly chance is involved in a given value. If something is measured with near-perfect reliability, such as height with a good measuring tape, there is hardly any regression to the mean. If, on the other hand, a value is almost pure chance, as with a roll of the dice, it returns almost completely to the average. Most real-world measurements lie somewhere in between: grades, revenues, sports results, model evaluations.
Statisticians describe this using the correlation between two measurements, that is, how strongly they are related. The weaker this relationship, the more strongly the second value slides toward the mean. At a correlation of 0.5, the expected second value lies only half as far from the average as the first one did. Importantly, this applies to the expected value, that is, the average across many cases. Individual exceptions are always possible.
A common mistake is to suspect a balancing force at work here. Chance has no memory and owes no one compensation. It’s simply that average results occur far more often than extreme ones. Anyone who specifically picks out the extremes from a group will almost inevitably end up with a mix of lucky and unlucky cases.
From language model leaderboards to stock portfolios
In the world of AI, this effect shows up in leaderboards for models. A model leads a benchmark, a standardized test procedure, by a narrow margin. In the next test run with different tasks, it often slips back down. Part of its lead consisted of test questions it happened to be well suited for. That’s why credible evaluations report margins of variation instead of just a single number.
In financial news, the term comes up with funds and stocks. The best-performing fund of a given year strikingly often ends up in the middle of the pack the following year. Advertising based on past top rankings is therefore not very informative. The warning that past returns are no guarantee of future performance is aimed at exactly this.
The effect is also clearly visible in everyday life. A soccer player scores three goals in one game and then goes quiet for weeks afterward. A video unexpectedly goes viral, and the following ones perform normally. Recognizing such patterns as regression to the mean saves a lot of unnecessary explanations. The best safeguard remains a control group: a comparison group with the same starting conditions, but without the intervention.