Sample, population, and why the norms matter more than the score
A test result is a comparison. If you do not know the comparison group, you do not have a result.
Every norm-referenced score answers the question "compared with whom?" — and answers it silently. The comparison group was fixed years earlier, by whoever standardised the instrument, and it is carried into every report the instrument ever produces. Most people reading a score never learn what it was.
What a norm sample is
To produce norms, an instrument is administered to a group assembled to represent the population it will be used with. Their scores form the distribution against which every later test-taker is placed. The whole interpretive apparatus — percentiles, standard scores, bands, descriptive labels — rests on that group.
This means the quality of an instrument's interpretation is limited by the quality of that sample, no matter how well the questions themselves are written. A carefully constructed test with an unrepresentative norm sample will produce precise, consistent, misleading placements.
Three ways norms fail
They are too old. Populations change — in schooling, nutrition, familiarity with the test format, and exposure to the kind of reasoning being tested. A comparison group assembled decades ago may place a student inaccurately for reasons that have nothing to do with the student.
They are from elsewhere. An instrument normed on one country's students and applied to another's compares a person with a group they do not belong to. Where the material involves language, cultural reference or schooling conventions, the mismatch is not small.
They are not what they claim. A sample of a thousand collected from four urban private schools is a sample of a thousand. It is not a national norm, and describing it as one is a claim about representativeness that the collection method does not support.
Sample size is the number most often advertised and the least informative. A large unrepresentative sample gives a precise answer to the wrong question. (AERA, APA, & NCME, 2014)
The practical consequence
If a student's profile is being used to compare their own strengths against one another — which is the most common and most defensible use in career work — every score in the profile must come from the same norm group. Scores normed on different samples cannot be laid side by side, however similar the numbers look. Reports that assemble subtests from different sources without saying so make this error invisibly.
Questions worth asking
- How many people, of what ages, from where, and collected in which year?
- How were they recruited? Convenience sampling and stratified sampling produce very different claims.
- Are separate norms provided where they matter — by age band, by language of administration?
- Do all the scores in this report share one norm group?
A supplier who can answer these has done the work. One who responds by describing the sample as large has changed the subject, and the change of subject is the information you were looking for.
References
Every work below was checked against a primary or catalogue record. Where a volume or page range could not be confirmed it is left out rather than guessed.
- American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association.