Before you buy an AI-based assessment, ask these questions
The technology is new. The questions that decide whether an instrument is sound are not.
Assessment products increasingly describe themselves as AI-driven, adaptive, or powered by machine learning. Some are substantial pieces of work. Others are a conventional questionnaire with a new adjective. A school or organisation buying one is rarely in a position to inspect the method, but it is in a position to ask questions, and the answers are informative even when the method is opaque.
The questions are the old ones
Nothing about the technology changes what makes an instrument sound. The claims a report makes about a person still have to be supported by evidence, and the evidence is of the same kinds it has always been (AERA, APA, & NCME, 2014).
What exactly does it claim to measure, in a sentence? If the answer is a list of eight fashionable constructs, ask which of them the instrument has evidence for, separately. Breadth of claim and depth of evidence are usually in inverse proportion.
What is the norm group? How many people, of what ages, in what country, collected when. This question is answerable by any serious supplier and is frequently answered vaguely, which is itself the answer.
What is the reliability, and by which method? A number with a method attached and a sample described. A number alone is not evidence.
What validity evidence exists for the use I intend? Not validity in general. Validity for deciding this, about people like these. (Society for Industrial and Organizational Psychology, 2018)
What does the report say that the data cannot support? Read a sample report critically. Look for confident sentences about personality, suitability or future performance and ask what measurement stands behind each one.
Questions specific to the technology
Some questions do arise from the method, and they are worth adding.
Adaptive on what basis? If the instrument selects later items from earlier answers, that is a design with a long and respectable history, and the supplier should be able to name the model it uses and how it was calibrated. If it cannot, "adaptive" is a description of the interface.
What was the training data? A model trained on one population and applied to another may behave differently, and in assessment that difference has a name and a legal weight. Ask whether the instrument has been examined for differential functioning across groups, and what was found.
Can the result be explained? If a student asks why their report says what it says, is there an answer? An instrument whose output cannot be explained to the person it describes is difficult to defend and impossible to appeal.
Where does the data go? Who holds it, for how long, under which jurisdiction, and what happens if the supplier is acquired. For minors this is not an administrative detail.
A supplier who answers these questions plainly is showing you their work. A supplier who answers them with adjectives has answered them.
The simplest test
Ask what the instrument is not suitable for. Every real instrument has a boundary and a careful supplier knows exactly where it is. An assessment that is presented as appropriate for every purpose, every age and every decision has not been examined closely enough by the people selling it, and it will not be examined more closely by the people using it.
References
Every work below was checked against a primary or catalogue record. Where a volume or page range could not be confirmed it is left out rather than guessed.
- American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association.
- Society for Industrial and Organizational Psychology. (2018). Principles for the validation and use of personnel selection procedures (5th ed.). Industrial and Organizational Psychology, 11(S1), 1–97.