Artificial intelligence comparison of Google MedGemma and ChatGPT-5 in expert-reviewed gastrointestinal and genitourinary oncology cases.
Abstract
e15509 Background: Large language models (LLMs) are increasingly explored for oncology decision support, yet comparative evidence evaluating medical-specific versus general-purpose LLMs using expert-reviewed clinical cases is limited. We compared a medical-specific LLM (Google MedGemma 27B) with a general-purpose LLM (ChatGPT-5) using gastrointestinal (GI) and genitourinary (GU) oncology vignettes. Methods: This cross-sectional study evaluated 32 de-identified oncology vignettes (26 GI, 6 GU) created by a U.S. board-certified medical oncologist and informed by real-world cases. Vignettes included text-only information on the clinical context, description of pathology, tumor staging, biomarkers, and relevant comorbidities. Both LLMs were prompted to provide National Comprehensive Cancer Network(NCCN) guideline-based treatment recommendations, monitoring strategies, maintenance therapy considerations, and references, along with a self-reported confidence score (0–100). Board-certified U.S. medical oncologists independently reviewed outputs using a structured instrument; two disease specific expert oncologists evaluated each case. Outcomes included correctness and comprehensiveness (1–10 scales), hallucinations (factually fabricated or unsupported statements), NCCN traceability, potential patient harm, and clinician trust. Associations between accuracy and comprehensiveness were assessed using correlation coefficients. Inter-rater reliability was assessed using a two-way random-effects intraclass correlation coefficient (ICC). Results: ChatGPT-5 demonstrated higher mean correctness (8.65 vs 4.68) and comprehensiveness (8.68 vs 5.22) scores compared with MedGemma 27B. Accuracy and comprehensiveness were strongly correlated overall (Pearson r = 0.93, p < 0.001), with significant model-specific correlations (ChatGPT-5 r = 0.86, p < 0.001; MedGemma r = 0.77, p < 0.001). Hallucinations occurred less frequently with ChatGPT-5 (12.3% vs 65%), and NCCN-traceable recommendations were more common (75.4% vs 36.7%). Responses judged as potentially harmful occurred in 18.5% of ChatGPT-5 outputs versus 70% of MedGemma outputs. Clinician trust favored ChatGPT-5 (84.6% vs 15%). Inter-rater reliability improved with rater averaging (ICC(A,k): correctness 0.72; comprehensiveness 0.64). Conclusions: In expert-reviewed GI and GU oncology cases, ChatGPT-5 demonstrated higher clinical accuracy, comprehensiveness, and guideline alignment than a medical-specific LLM; however, a substantial rate of hallucinations and potentially harmful recommendations persisted across models. These findings underscore that, despite performance differences, current LLMs require rigorous clinician oversight and should not be used as standalone decision-support tools in oncology.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (14)
Ali Zidan
University of Toronto, Toronto, ON, Canada
Mousa El-Sururi
University of Toronto, Toronto, ON, Canada
Avi Belbase
Tufts University, Somerville, MA
Babar Bashir
Sidney Kimmel Comprehensive Cancer Center at Thomas Jefferson University Hospital, Philadelphia, PA
Meredith Pelster
Sarah Cannon Research Institute, Nashville
Youssef Bouferraa
Cleveland Clinic Taussig Cancer Institute, Cleveland, OH
Alok A. Khorana
Taussig Cancer Institute, Cleveland, OH
Daniel Lin
Childrens Hospital Colorado, Aurora, Colorado, United States
Steven J. Cohen
Asplundh Cancer Pavilion, Sidney Kimmel Medical College, Willow Grove, PA
Hao Xie
Beijing National Laboratory for Condensed Matter Physics, Institute of Physics, Chinese Academy of Sciences
Cody Eslinger
Department of Hematology and Oncology, Mayo Clinic Arizona, Phoenix, AZ
Benjamin Garmezy
Sarah Cannon Research Institute, Nashville, TN
Benjamin L. Maughan
University of Utah, Salt Lake City, UT
Roupen Odabashian
Department of Hematology and Oncology, Karmanos Cancer Institute, Wayne State University, Detroit, MI