Pronunciation assessment in foreign language learning: Reliability and scoring bias in human–generative AI evaluation
Abstract
This study examines the reliability and scoring bias of generative AI (Gen-AI)-based pronunciation assessment compared with human raters, addressing whether AI-generated scores can be trusted in real educational settings. Sixty students participated in a 12-week program. A total of 180 pronunciation samples were evaluated across eight subcomponents (individual phonemes, stress, rhythm, intonation, linking, reduction, fluency, and clarity) by three standardized human raters and Gen-AI using the same 7-point rubric. Quantitative analyses (intraclass correlation coefficients, paired-samples t-tests, and Pearson correlations) assessed reliability and bias, while semi-structured interviews with raters provided explanatory qualitative insights. Gen-AI demonstrated moderate reliability with human raters across most components, showing the highest agreement in fluency and the weakest in individual phonemes. However, Gen-AI consistently assigned significantly higher scores than human raters across all subcomponents. Qualitative findings revealed that discrepancies originated from Gen-AI’s limited discriminative power, systematic flaws in handling missing data, decontextualized scoring approach, lack of sensitivity to L1 interference, and inability to interpret pragmatic context. While Gen-AI cannot fully replace human expertise, it can function as a complementary tool for formative assessment and autonomous practice. A hybrid assessment model integrating Gen-AI’s efficiency with human raters’ contextual and interpretive insights is recommended for effective foreign language pronunciation education.
Article Details
Authors (1)
Whyunyoung Choi