How valuable are the questions and answers generated by large language models in oral and maxillofacial surgery?

K Kyuhyung Kim (Department of Brain Sciences, Daegu Gyeongbuk Institute of Science & Technology (DGIST)) S Sae Byeol Mun Y Young Jae Kim B Bong Chul Kim K Kwang Gi Kim

Abstract

Introduction In this study, we aim to evaluate the ability of large language models (LLM) to generate questions and answers in oral and maxillofacial surgery. Methods ChatGPT4, ChatGPT4o, and Claude3-Opus were evaluated in this study. Each LLM was instructed to generate 50 questions about oral and maxillofacial surgery. Three LLMs were asked to answer the generated 150 questions. Results All 150 questions generated by the three LLMs were related to oral and maxillofacial surgery. Each model exhibited a correct answer rate of over 90%. None of the three models were able to answer correctly all the questions they generated themselves. The correct answer rate was 97.0% for questions with figures, significantly higher than the 88.9% rate for questions without figures. The analysis of problem-solving by the three LLMs showed that each model generally inferred answers with high accuracy, and there were few logical errors that could be considered controversial. Additionally, all three scored above 88% for the fidelity of their explanations. Conclusion This study demonstrates that while LLMs like ChatGPT4, ChatGPT4o, and Claude3-Opus exhibit robust capabilities in generating and solving oral and maxillofacial surgery questions, their performance is not without limitations. None of the models were able to answer correctly all the questions they generated themselves, highlighting persistent challenges such as AI hallucinations and contextual understanding gaps. The results also emphasize the importance of multimodal inputs, as questions with annotated images achieved higher accuracy rates compared to text-only prompts. Despite these shortcomings, the LLMs showed significant promise in problem-solving, logical consistency, and response fidelity, particularly in structured or numerical contexts.

Article Details

Journal PLoS ONE
Volume / Issue Vol. 20, Issue 5
Published May 28, 2025
Pages e0322529
ISSN 1932-6203
Publisher Public Library of Science

Journal Info

PLoS ONE

Public Library of Science

ISSN: 1932-6203 Open Access Health Sciences

Authors (5)

K

Kyuhyung Kim

Department of Brain Sciences, Daegu Gyeongbuk Institute of Science & Technology (DGIST)

S

Sae Byeol Mun

Y

Young Jae Kim

B

Bong Chul Kim

K

Kwang Gi Kim