Regression to the mean, and why the second test looks worse
One of the most useful ideas in measurement, and the one whose absence produces the most confident wrong conclusions.
A student scores exceptionally on a test. Extra coaching is arranged. They are tested again and the score is lower. The natural conclusion is that the coaching failed, or that the first result was a fluke, or that the student has stopped trying. There is a fourth explanation, it requires no assumptions about anyone's effort, and it is usually the largest part of what happened.
The idea
Any observed score has two components: the person's true standing, and the error of that occasion — the questions that happened to suit them, the sleep, the room, the guesses. When someone scores extremely, it is more likely than not that the occasion's error was pushing in the same direction as their ability. On a second occasion the error is redrawn, and is unlikely to be as favourable again. The score moves back towards the average.
This happens at both ends. Very low scores tend to rise on retesting for the same reason. Nothing about the person has changed. Only the error has been resampled.
The more extreme the first score, and the less reliable the measure, the larger the movement back. It is a property of the measurement, not of the person. (Kahneman, 2011, ch. 17)
Why it produces false conclusions about interventions
Consider the standard arrangement: identify the weakest students, give them an intervention, retest, observe improvement. Some of that improvement is regression to the mean, and without a comparison group of equally weak students who received nothing, there is no way to say how much. The design guarantees an apparent effect whether or not the intervention does anything.
This is the single most common reason that programmes appear to work in-house and fail to replicate elsewhere. It also runs the other way and is even less noticed: praise a student after an unusually good performance and the next one will typically be worse, which teaches the observer that praise is counterproductive. Criticise after an unusually bad one and the next will typically be better. Both lessons are artefacts.
What to do about it
Compare with a control. The only sound way to know whether something worked is to observe what happened to comparable people who did not receive it.
Do not select on an extreme score alone. Where selection must be made, use more than one occasion or more than one source. Both reduce the influence of any single occasion's error.
Expect the movement and say so in advance. Telling a family before a retest that an exceptional score usually comes down slightly, and that this is expected, prevents a disappointment that was built into the measurement.
Use the more reliable instrument. The effect shrinks as reliability rises (AERA, APA, & NCME, 2014), which is one more practical reason that reliability is worth paying for.
The general form
The habit of mind is worth carrying beyond testing. Whenever something is selected because it was extreme — the best-performing school, the worst month, the star quarter — expect the next observation to be less extreme, and be suspicious of any explanation offered for the change before that expectation has been accounted for.
References
Every work below was checked against a primary or catalogue record. Where a volume or page range could not be confirmed it is left out rather than guessed.
- American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association.
- Kahneman, D. (2011). Thinking, fast and slow. Farrar, Straus and Giroux.