As employers increasingly adopt AI to evaluate interview responses, understanding what makes these systems reliable becomes important. The study offers one of the clearest roadmaps at this point in time for building AI interview scoring systems that are both scientifically sound and fair.
The researchers compared large language models with both trained human interviewers and earlier machine-learning systems that required thousands of examples to learn how to score interviews.
Their conclusion was encouraging: The latest Gen AI LLMs can produce interview scores with reliability and validity that are comparable to, and sometimes better than, traditional machine-learning models and individual human raters. However, this wasn't automatic. Performance depended heavily on how the AI was designed and instructed.
Think of an LLM as a newly hired interviewer. Simply asking it to "score this interview" is like saying, "You can figure it out on your own." Instead, the AI performs much better when it's given the equivalent of interviewer training: clear definitions of what it's evaluating, behavioral examples, structured rating scales, and explicit scoring instructions.
The study identified several practices that consistently improved AI scoring.
Use newer, larger language models. Just as experienced interviewers generally make better judgments than novices, newer generations of LLMs handled the complex task of evaluating interview responses more effectively than older or smaller models.
Give the AI detailed information about the competency being measured. Rather than merely telling the model to score "adaptability" or "conscientiousness," providing definitions, behavioral descriptions, and behaviorally anchored rating scales (BARS) substantially improved scoring accuracy. The richer the guidance, the more the AI focused on the intended competency instead of superficial aspects of the response.
Evaluate one interviewee at a time. Humans can be influenced by comparing one candidate to another, and the researchers found evidence that AI systems can also become distracted when evaluating multiple people simultaneously. Scoring candidates independently helps reduce unwanted comparison effects.
Use multiple AI "judges." Just as organizations often use interview panels instead of a single interviewer, combining ratings from multiple AI evaluations, even different language models, can reduce random errors and improve consistency. The researchers describe this as creating an ensemble of raters, similar to benefiting from the "wisdom of the crowd."
One especially interesting finding is that LLMs differ from earlier AI scoring methods. Traditional machine-learning systems usually require thousands of previously scored interviews before they can evaluate new candidates. The latest LLMs, by contrast, can often perform well with little or no task-specific training simply by receiving well-designed prompts. This dramatically lowers the barrier to creating AI-assisted interview scoring systems.
What This Means for AI Interviews
For interviewees, the practical lesson is reassuring: Well-designed LLM-based systems can evaluate whether your story demonstrates the competency being assessed, not merely whether you used keywords.
For employers, the study delivers a different message. Buying access to a powerful language model is only the beginning. The quality of an AI interview scorer depends on thoughtful design: selecting an appropriate model, carefully engineering prompts, providing detailed competency definitions, using structured rating scales, and validating the results against established psychometric standards. While it is possible to use a frontier model instead of a proprietary model, cautions must be taken.
The researchers aemphasize LLMs should not be treated as infallible. High-stakes employment decisions still require careful validation, ongoing monitoring for bias and drift, and evidence that the system measures what it claims to measure. In other words, AI may become an outstanding interview evaluator, but only when it is built with the same scientific rigor expected of human assessment systems.
Currently, most employers will use a validated, proprietary model, but the future may hold a new answer where vendors are not needed as much.
To Reference the Paper
Title: Scoring Employment Interviews with Large Language Models: Evaluation Design Components, Validity Investigations, and Best Practice Recommendations
Authors: Kayden Stockdale, Louis Hickman, & Siyi Liu
Journal: Journal of Applied Psychology
Publication Year: 2026
DOI: 10.1037/apl0001396
Better stories are possible — and they start with preparation.
AI Interview Training Course