Examining the joint impact of missing data mechanisms and item parameter drift on the accuracy of item response theory-based test equating
Abstract
Maintaining score comparability across different test administrations is essential in large-scale educational and psychological assessment. Two major threats to equating accuracy are item parameter drift (IPD), reflecting changes in anchor item characteristics over time, and missing item responses, which commonly occur in operational testing. Although each factor has been studied separately, their joint impact on equating accuracy has not been systematically evaluated. A Monte Carlo simulation was conducted using a nonequivalent groups with anchor test design. Data were generated under a three-parameter logistic model with 2,000 examinees per form and 500 replications per condition. The Stocking–Lord method was used to estimate equating constants across 25 conditions, varying IPD rate (0%, 10%, 20%), IPD magnitude (0, 0.25, 0.50), missing data mechanism (MCAR vs. MAR), and missing data rate (0%, 10%, 20%). Missing responses were addressed using multiple imputation by chained equations. Equating accuracy was evaluated using bias and root mean square error. Results indicated that the B constant was highly sensitive to the missing data mechanism: MCAR conditions produced near-zero bias, whereas MAR conditions introduced substantial positive bias, particularly at higher missing rates. IPD alone did not result in meaningful bias but increased estimation variability. When IPD and MAR co-occurred, equating error exceeded the sum of their individual effects, indicating an interaction rather than a purely additive relationship. Although multiple imputation reduced error under MCAR, it did not fully eliminate bias under MAR conditions. In contrast, the scale constant remained stable across all conditions. Overall, the missing data mechanism had a stronger impact on equating accuracy than either the rate of missingness or the degree of item drift alone. When both factors were present, their combined influence led to increased error that could not be fully corrected through imputation. These findings underscore the importance of carefully evaluating missing data mechanisms, selecting appropriate imputation strategies, such as MICE, and monitoring anchor item stability to ensure accurate score equating, particularly in psychological assessment contexts where even small errors may affect individual-level decisions.
Article Details
Authors (2)
Ayse Bilicioglu Gunes
Serife Zeybekoglu Yesil