Examining the joint impact of missing data mechanisms and item parameter drift on the accuracy of item response theory-based test equating

A Ayse Bilicioglu Gunes S Serife Zeybekoglu Yesil

Abstract

Maintaining score comparability across different test administrations is essential in large-scale educational and psychological assessment. Two major threats to equating accuracy are item parameter drift (IPD), reflecting changes in anchor item characteristics over time, and missing item responses, which commonly occur in operational testing. Although each factor has been studied separately, their joint impact on equating accuracy has not been systematically evaluated. A Monte Carlo simulation was conducted using a nonequivalent groups with anchor test design. Data were generated under a three-parameter logistic model with 2,000 examinees per form and 500 replications per condition. The Stocking–Lord method was used to estimate equating constants across 25 conditions, varying IPD rate (0%, 10%, 20%), IPD magnitude (0, 0.25, 0.50), missing data mechanism (MCAR vs. MAR), and missing data rate (0%, 10%, 20%). Missing responses were addressed using multiple imputation by chained equations. Equating accuracy was evaluated using bias and root mean square error. Results indicated that the B constant was highly sensitive to the missing data mechanism: MCAR conditions produced near-zero bias, whereas MAR conditions introduced substantial positive bias, particularly at higher missing rates. IPD alone did not result in meaningful bias but increased estimation variability. When IPD and MAR co-occurred, equating error exceeded the sum of their individual effects, indicating an interaction rather than a purely additive relationship. Although multiple imputation reduced error under MCAR, it did not fully eliminate bias under MAR conditions. In contrast, the scale constant remained stable across all conditions. Overall, the missing data mechanism had a stronger impact on equating accuracy than either the rate of missingness or the degree of item drift alone. When both factors were present, their combined influence led to increased error that could not be fully corrected through imputation. These findings underscore the importance of carefully evaluating missing data mechanisms, selecting appropriate imputation strategies, such as MICE, and monitoring anchor item stability to ensure accurate score equating, particularly in psychological assessment contexts where even small errors may affect individual-level decisions.

Article Details

Journal PLoS ONE
Volume / Issue Vol. 21, Issue 7
Published July 29, 2026
Pages e0353665
ISSN 1932-6203
Publisher Public Library of Science

Journal Info

PLoS ONE

Public Library of Science

ISSN: 1932-6203 Open Access Health Sciences

Authors (2)

A

Ayse Bilicioglu Gunes

S

Serife Zeybekoglu Yesil