Sources matter: Learning to evaluate large language models for oncology treatment plans and changes observed by anchoring models to an oncology-specific database.
Abstract
1617 Background: Large language models (LLMs) are increasingly evaluated for oncology decision support, yet guideline concordance and hallucinations remain variable across cancer types and tasks. Retrieval-augmented generation (RAG) may improve safety by constraining the evidence base used during generation. Prior oncology LLM studies report discordance with guideline-based care, but few describe a reproducible health-system “vetting mechanism” that couples guideline-only retrieval with clinician-scored safety rubrics across multiple oncology domains. Methods: We assembled 316 de-identified oncology cases in a tumor-board template: medical oncology (breast, GI, GU, leukemia; n=216), gynecologic oncology (n=50), and neuro-oncology/CNS metastatic cancer (n=50). Each case was run on three systems: baseline GPT-4 (O1), NCCN-guideline anchored RAG (O2; guideline-constrained retrieval), and a literature-anchored system (OpenEvidence; O3). Oncologists independently scored outputs using a modified Generative Performance Score (mGPS; −1 to +1) combining guideline concordance with hallucination penalties; readability/rationality was scored on a 1–5 Likert scale. Inter-rater reliability was assessed (κ =.63). This workflow was designed as a pre-deployment model credentialing step for health-system governance (source policy → case bank → clinician-scored acceptance criteria). Results: Across all 316 cases (case-averaged), O2 achieved the highest mean total mGPS (0.75) compared with O1 (0.43) and O3 (0.59). In the medical oncology cohort (n=216), O2 outperformed O1 and O3 (mean mGPS 0.72 vs 0.35 and 0.54), with higher guideline concordance (0.85 vs 0.68 and 0.75) and lower hallucination penalties (−0.09 vs −0.25 and −0.19); readability was also highest for O2 (4.17 vs 3.22 and 3.86). In gynecologic oncology (n=50), O2 again performed best (mGPS 0.827) relative to O3 (0.697) and O1 (0.647). In the CNS cohort (n=50), mean composite mGPS favored O2 (0.82) over O3 (0.68) and O1 (0.55), while readability remained stable across models (3.55–3.73). Severe hallucination/unsafe recommendation rate: O1 [26%], O2 [5%], O3 [17%] (pre-specified definition; e.g., major fabricated citation or guideline-discordant harmful recommendation). Conclusions: In a 316-case, multi-domain clinician-scored evaluation, limiting retrieval to oncology guideline sources improved safety-aligned performance relative to unconstrained generation or literature-only anchoring. A guideline-constrained RAG configuration paired with clinician-scored rubrics and tail-risk reporting can serve as a practical health-system vetting mechanism for governance, re-validation, and change control of LLM decision support, including in resource-limited settings where reliable guideline-consistent recommendations are essential.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (13)
Connor Yost
1Creighton University school of medicine, Phoenix, United States
Peter Palumbo
4Geisel School of Medicine at Dartmouth, Hanover, United States
Nikita Tripathi
1Creighton University school of medicine, Phoenix, United States
Bradley Callas
Creighton University School of Medicine, Phoenix, AZ
Yan Leyfman
5NewYork-Presbyterian Hospital, Hematology, New York, United States
Sarah Elizabeth Monick
Mayo Clinic Arizona, Phoenix, AZ
Matthew Reeves Sullivan
Dartmouth-Hitchcock Medical Center, Lebanon, NH
Yoshie Umemura
Ivy Brain Tumor Center at Barrow Neurological Institute, Phoenix, AZ
Vamsee Torri
Samuel A. Funt
Memorial Sloan Kettering Cancer Center, New York, NY
Ryan Huu-Tuan Nguyen
University of Illinois College of Medicine at Chicago, Division of Hematology and Oncology, Chicago, IL
Irbaz Bin Riaz
Irbaz Bin Riaz, MD, PhD; R. Bryan Rumble, MSc; Thomas A. Hope, MD; Giuseppe Procopio, MD; and Neha Vapiwala, MD; Mayo Clinic, Phoenix, AZ; American Society of Clinical Oncology, Alexandria, VA; University of California, San Francisco, San Francisco, CA; Fondazione IRCCS Istituto Nazionale dei Tumori di Milano, Milan, Italy; and University of Pennsylvania Abramson Cancer Center, Philadelphia, PA
Arturo Loaiza-Bonilla
3St. Luke's Cancer Center, Oncology Hematology, Easton, United States