Evaluating artificial intelligence's (AI) clinical reasoning with skin cancer multiple-choice questions.
Abstract
e21598 Background: While large language models (LLMs) like GPT-4 (OpenAI) and Claude Opus (Anthropic) perform well on general medical exams, limited data exist on their application in oncology, particularly in multiple-choice question (MCQ) scenarios. This study benchmarks the accuracy and assesses the clinical reasoning of LLMs on cutaneous oncology MCQs from the American Society of Clinical Oncology (ASCO) question bank. Methods: Skin oncology MCQs from the ASCO question bank were tested using GPT-4 (OpenAI) and Claude Opus (Anthropic). Questions were entered into OpenAI and Anthropic API playgrounds without additional prompting, with identical settings for both models (Temperature = 0, Token = Max). Each question was tested three times per model to assess response consistency and precision. Initial responses were collected, and chain-of-thought (COT) prompting was applied to encourage stepwise reasoning mimicking an oncologist’s thought process. The percentage of revised answers after COT prompting was analyzed. Accuracy before and after COT prompting was compared. Incorrect responses were reviewed by board-certified oncologists specializing in skin cancer at National Cancer Institute (NCI)-designated centers. Each oncologist scored the responses on clarity, clinical reasoning, bias, and confidence in the models’ utility in clinical practice. Qualitative feedback was collected via free-text comments and analyzed descriptively. Results: Of 66 skin oncology MCQs, Claude correctly answered 55 questions (83% accuracy, 95% CI: 72%–91%), while ChatGPT correctly answered 53 (80% accuracy, 95% CI: 68%–89%). Both models performed comparably, with overlapping confidence intervals. A chi-square test showed no significant difference between their performances (p = 0.65). After COT prompting, both tools remained at their respective accuracy level. Thematic analysis of incorrect answers revealed that LLM responses frequently deviated from clinical guidelines and failed to reflect evidence-based oncology practice, often neglecting key trial data like DREAMSeq and frontline therapeutic strategies. The most frequent errors were Treatment Planning Errors, where recommendations misaligned with NCCN guidelines, and Interpretation Errors, involving misinterpretation of clinical trials or patient data. Anchoring bias emerged as a recurring issue, with models fixating on details (e.g., BRAF mutation status) at the expense of broader reasoning. Conclusions: This study demonstrates that while large language models like GPT-4 and Claude Opus show promise in the field of oncology, they still face significant challenges in clinical reasoning and adherence to evidence-based practices. Continued refinement and training on specialized medical data are essential before these models can be reliably integrated into clinical workflows.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (6)
Suha Soni
Department of Medicine, Division of Cardiology, The University of Texas Health Science Center, San Antonio, TX
Roupen Odabashian
Department of Hematology and Oncology, Karmanos Cancer Institute, Wayne State University, Detroit, MI
Gregory Dyson
Karmanos Cancer Institute/Department of Oncology, Detroit, MI
James William Smithy
Memorial Sloan Kettering Cancer Center, New York, NY
Padmapriya Muthu
Mayo Clinic Arizona, Scottsdale, Arizona, United States
Yusra F. Shao
Barbara Ann Karmanos Cancer Institute, Wayne State University, Detroit, MI