Artificial intelligence comparison of Google MedGemma and ChatGPT-5 in expert-reviewed gastrointestinal and genitourinary oncology cases.

A Ali Zidan (University of Toronto, Toronto, ON, Canada) M Mousa El-Sururi (University of Toronto, Toronto, ON, Canada) A Avi Belbase (Tufts University, Somerville, MA) B Babar Bashir (Sidney Kimmel Comprehensive Cancer Center at Thomas Jefferson University Hospital, Philadelphia, PA) M Meredith Pelster (Sarah Cannon Research Institute, Nashville) Y Youssef Bouferraa (Cleveland Clinic Taussig Cancer Institute, Cleveland, OH) A Alok A. Khorana (Taussig Cancer Institute, Cleveland, OH) D Daniel Lin (Childrens Hospital Colorado, Aurora, Colorado, United States) S Steven J. Cohen (Asplundh Cancer Pavilion, Sidney Kimmel Medical College, Willow Grove, PA) H Hao Xie (Beijing National Laboratory for Condensed Matter Physics, Institute of Physics, Chinese Academy of Sciences) C Cody Eslinger (Department of Hematology and Oncology, Mayo Clinic Arizona, Phoenix, AZ) B Benjamin Garmezy (Sarah Cannon Research Institute, Nashville, TN) B Benjamin L. Maughan (University of Utah, Salt Lake City, UT) R Roupen Odabashian (Department of Hematology and Oncology, Karmanos Cancer Institute, Wayne State University, Detroit, MI)

Abstract

e15509 Background: Large language models (LLMs) are increasingly explored for oncology decision support, yet comparative evidence evaluating medical-specific versus general-purpose LLMs using expert-reviewed clinical cases is limited. We compared a medical-specific LLM (Google MedGemma 27B) with a general-purpose LLM (ChatGPT-5) using gastrointestinal (GI) and genitourinary (GU) oncology vignettes. Methods: This cross-sectional study evaluated 32 de-identified oncology vignettes (26 GI, 6 GU) created by a U.S. board-certified medical oncologist and informed by real-world cases. Vignettes included text-only information on the clinical context, description of pathology, tumor staging, biomarkers, and relevant comorbidities. Both LLMs were prompted to provide National Comprehensive Cancer Network(NCCN) guideline-based treatment recommendations, monitoring strategies, maintenance therapy considerations, and references, along with a self-reported confidence score (0–100). Board-certified U.S. medical oncologists independently reviewed outputs using a structured instrument; two disease specific expert oncologists evaluated each case. Outcomes included correctness and comprehensiveness (1–10 scales), hallucinations (factually fabricated or unsupported statements), NCCN traceability, potential patient harm, and clinician trust. Associations between accuracy and comprehensiveness were assessed using correlation coefficients. Inter-rater reliability was assessed using a two-way random-effects intraclass correlation coefficient (ICC). Results: ChatGPT-5 demonstrated higher mean correctness (8.65 vs 4.68) and comprehensiveness (8.68 vs 5.22) scores compared with MedGemma 27B. Accuracy and comprehensiveness were strongly correlated overall (Pearson r = 0.93, p < 0.001), with significant model-specific correlations (ChatGPT-5 r = 0.86, p < 0.001; MedGemma r = 0.77, p < 0.001). Hallucinations occurred less frequently with ChatGPT-5 (12.3% vs 65%), and NCCN-traceable recommendations were more common (75.4% vs 36.7%). Responses judged as potentially harmful occurred in 18.5% of ChatGPT-5 outputs versus 70% of MedGemma outputs. Clinician trust favored ChatGPT-5 (84.6% vs 15%). Inter-rater reliability improved with rater averaging (ICC(A,k): correctness 0.72; comprehensiveness 0.64). Conclusions: In expert-reviewed GI and GU oncology cases, ChatGPT-5 demonstrated higher clinical accuracy, comprehensiveness, and guideline alignment than a medical-specific LLM; however, a substantial rate of hallucinations and potentially harmful recommendations persisted across models. These findings underscore that, despite performance differences, current LLMs require rigorous clinician oversight and should not be used as standalone decision-support tools in oncology.

Article Details

Volume / Issue Vol. 44, Issue 16_suppl
Published June 01, 2026
ISSN 0732-183X
Publisher Lippincott Williams & Wilkins

Journal Info

Journal of Clinical Oncology

Lippincott Williams & Wilkins

ISSN: 0732-183X Health Sciences

Authors (14)

A

Ali Zidan

University of Toronto, Toronto, ON, Canada

M

Mousa El-Sururi

University of Toronto, Toronto, ON, Canada

A

Avi Belbase

Tufts University, Somerville, MA

B

Babar Bashir

Sidney Kimmel Comprehensive Cancer Center at Thomas Jefferson University Hospital, Philadelphia, PA

M

Meredith Pelster

Sarah Cannon Research Institute, Nashville

Y

Youssef Bouferraa

Cleveland Clinic Taussig Cancer Institute, Cleveland, OH

A

Alok A. Khorana

Taussig Cancer Institute, Cleveland, OH

D

Daniel Lin

Childrens Hospital Colorado, Aurora, Colorado, United States

S

Steven J. Cohen

Asplundh Cancer Pavilion, Sidney Kimmel Medical College, Willow Grove, PA

H

Hao Xie

Beijing National Laboratory for Condensed Matter Physics, Institute of Physics, Chinese Academy of Sciences

C

Cody Eslinger

Department of Hematology and Oncology, Mayo Clinic Arizona, Phoenix, AZ

B

Benjamin Garmezy

Sarah Cannon Research Institute, Nashville, TN

B

Benjamin L. Maughan

University of Utah, Salt Lake City, UT

R

Roupen Odabashian

Department of Hematology and Oncology, Karmanos Cancer Institute, Wayne State University, Detroit, MI