Testing standards for AI-based scores in automated essay scoring

R Rudolf Debelak M Matthias Ziegler

Abstract

Recent developments in the field of artificial intelligence and machine learning allow the wide application of large language models for the evaluation of written text and other non-numerical data. When applied in the context of psychological and educational assessments, such models can be used for assigning scores to essays and other types of responses. In contrast to classical tests, essays do not consist of test items, which leads to specific challenges in the evaluation of testing standards for scores obtained from AI models that differ from those observed for classical ability tests and personality questionnaires. To address these challenges, we discuss the evaluation of validity, fairness, and reliability for scores obtained from models of artificial intelligence in the context of automated essay scoring. We review existing methods, propose new methods, and further illustrate the reviewed methods with an empirical example based on the Hewlett Foundation data set on automated essay scoring. By applying the proposed framework to an evaluation based on a DistilBERT model, we find the model to be robust with sufficiently high internal consistency (Spearman-Brown coefficients in the range from .77 to .92). We further found empirical evidence for the validity of the evaluation model, but also indications for violations of fairness when comparing the human and AI scores across different topics. This study provides a standardized, replicable toolkit for researchers and practitioners to evaluate the psychometric quality of AI-based assessments.

Article Details

Journal PLoS ONE
Volume / Issue Vol. 21, Issue 7
Published July 31, 2026
Pages e0354680
ISSN 1932-6203
Publisher Public Library of Science

Journal Info

PLoS ONE

Public Library of Science

ISSN: 1932-6203 Open Access Health Sciences

Authors (2)

R

Rudolf Debelak

M

Matthias Ziegler