GPT-4 shows comparable performance to human examiners in ranking open-text answers
Abstract
Abstract Can GPT-4 replace human examiners? To address this question we explore the performance of GPT-4 as an examiner of answers to open-text questions. We formulate questions and sample solutions in the field of macroeconomics and collect answers from cohorts of undergraduate students. We then conduct a fair competition between GPT-4 and human experts, employing their expertise to assess the quality of the answers. We observe that the substitution of GPT-4 for a human examiner does not decrease inter-rater reliability on tasks that rank the quality of answers. We run checks on potential biases (whether GPT-4 prefers AI-generated or lengthy answers). We find no consistent evidence of such biases. Our findings are robust to tilting the competition to one side’s advantage, by using inferior or advanced prompting strategies. Our results are more attenuated on tasks where GPT-4 assigns points to student answers. Here, GPT-4 shows a bias towards longer answers. Overall, our study cautiously supports the utilization of GPT-4 as an assistant for automated grading systems, particularly those where answers are ranked according to their quality.
Article Details
Authors (5)
Abdullah Al Zubaer
Michael Granitzer
Stephan Geschwind
Johann Graf Lambsdorff
Deborah Voss