Reliability and validity, without the jargon
A test can be highly reliable and completely useless. The two words are not synonyms and the difference decides whether a report means anything.
Two words appear in every technical manual and are used interchangeably almost everywhere else. They describe different properties, and a test can have one in abundance while lacking the other entirely.
Reliability is consistency
A reliable instrument gives the same answer when the same thing is measured again. If a student sits an aptitude test on Monday and again a fortnight later, and the two results are close, the test has good test–retest reliability. If the questions within a test all appear to be measuring the same underlying thing, the test has good internal consistency (AERA, APA, & NCME, 2014).
Reliability is a property of the measurement, not of the person. It says nothing whatsoever about whether the thing being measured is worth measuring.
Validity is aboutness
A valid instrument measures what it claims to measure, and — the part more often forgotten — supports the conclusions people actually draw from it (Messick, 1995). Validity is not a single property that a test either has or lacks. It is a claim about a particular use. (Cronbach & Meehl, 1955)
A test may be perfectly valid as a measure of verbal reasoning and entirely invalid as a basis for predicting who will succeed in a sales role, even though people will use it for the second purpose because they have it to hand. The right question is never “is this test valid?” but “is it valid for this decision, about this person, in this setting?”
Reliability asks whether the instrument is steady. Validity asks whether the instrument is pointed at the right thing. A steady instrument pointed at the wrong thing is worse than an unsteady one, because it inspires confidence.
Why a reliable test can be useless
Consider a questionnaire that asks people how tall they feel. Administered twice, it would produce similar answers each time — people are consistent about how they feel. Its reliability could be excellent. Its validity as a measure of height would be zero.
This is not a fanciful example. A great many popular assessments in circulation have respectable reliability and have never established that their categories predict anything outside the questionnaire itself. Consistency is easy to achieve. It is achieved by asking the same question in several ways. Whether the answers correspond to anything in the world is a separate and much harder question, and it is the one that requires evidence.
Reliability limits validity
The relationship between the two runs one way. An instrument cannot be more valid than it is reliable — if a measurement wobbles badly, it cannot correspond closely to anything stable. So reliability is necessary. It is simply not sufficient, and treating it as sufficient is the most common technical error in how assessments are chosen.
What to look for in a manual
- Reliability coefficients, stated with the method used and the sample they came from
- Validity evidence of more than one kind, and specifically evidence for the use you have in mind
- The date and composition of the norm sample
- A statement of what the instrument does not measure and should not be used for
The last item is the most informative. A manual that names the limits of its own instrument is describing a serious piece of work. A manual with no limits section is describing a product.
References
Every work below was checked against a primary or catalogue record. Where a volume or page range could not be confirmed it is left out rather than guessed.
- American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association.
- Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52, 281–302.
- Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons’ responses and performances as scientific inquiry into score meaning. American Psychologist, 50, 741–749.