Sources matter: Learning to evaluate large language models for oncology treatment plans and changes observed by anchoring models to an oncology-specific database.

C Connor Yost (1Creighton University school of medicine, Phoenix, United States) P Peter Palumbo (4Geisel School of Medicine at Dartmouth, Hanover, United States) N Nikita Tripathi (1Creighton University school of medicine, Phoenix, United States) B Bradley Callas (Creighton University School of Medicine, Phoenix, AZ) Y Yan Leyfman (5NewYork-Presbyterian Hospital, Hematology, New York, United States) S Sarah Elizabeth Monick (Mayo Clinic Arizona, Phoenix, AZ) M Matthew Reeves Sullivan (Dartmouth-Hitchcock Medical Center, Lebanon, NH) Y Yoshie Umemura (Ivy Brain Tumor Center at Barrow Neurological Institute, Phoenix, AZ) V Vamsee Torri S Samuel A. Funt (Memorial Sloan Kettering Cancer Center, New York, NY) R Ryan Huu-Tuan Nguyen (University of Illinois College of Medicine at Chicago, Division of Hematology and Oncology, Chicago, IL) I Irbaz Bin Riaz (Irbaz Bin Riaz, MD, PhD; R. Bryan Rumble, MSc; Thomas A. Hope, MD; Giuseppe Procopio, MD; and Neha Vapiwala, MD; Mayo Clinic, Phoenix, AZ; American Society of Clinical Oncology, Alexandria, VA; University of California, San Francisco, San Francisco, CA; Fondazione IRCCS Istituto Nazionale dei Tumori di Milano, Milan, Italy; and University of Pennsylvania Abramson Cancer Center, Philadelphia, PA) A Arturo Loaiza-Bonilla (3St. Luke's Cancer Center, Oncology Hematology, Easton, United States)

Abstract

1617 Background: Large language models (LLMs) are increasingly evaluated for oncology decision support, yet guideline concordance and hallucinations remain variable across cancer types and tasks. Retrieval-augmented generation (RAG) may improve safety by constraining the evidence base used during generation. Prior oncology LLM studies report discordance with guideline-based care, but few describe a reproducible health-system “vetting mechanism” that couples guideline-only retrieval with clinician-scored safety rubrics across multiple oncology domains. Methods: We assembled 316 de-identified oncology cases in a tumor-board template: medical oncology (breast, GI, GU, leukemia; n=216), gynecologic oncology (n=50), and neuro-oncology/CNS metastatic cancer (n=50). Each case was run on three systems: baseline GPT-4 (O1), NCCN-guideline anchored RAG (O2; guideline-constrained retrieval), and a literature-anchored system (OpenEvidence; O3). Oncologists independently scored outputs using a modified Generative Performance Score (mGPS; −1 to +1) combining guideline concordance with hallucination penalties; readability/rationality was scored on a 1–5 Likert scale. Inter-rater reliability was assessed (κ =.63). This workflow was designed as a pre-deployment model credentialing step for health-system governance (source policy → case bank → clinician-scored acceptance criteria). Results: Across all 316 cases (case-averaged), O2 achieved the highest mean total mGPS (0.75) compared with O1 (0.43) and O3 (0.59). In the medical oncology cohort (n=216), O2 outperformed O1 and O3 (mean mGPS 0.72 vs 0.35 and 0.54), with higher guideline concordance (0.85 vs 0.68 and 0.75) and lower hallucination penalties (−0.09 vs −0.25 and −0.19); readability was also highest for O2 (4.17 vs 3.22 and 3.86). In gynecologic oncology (n=50), O2 again performed best (mGPS 0.827) relative to O3 (0.697) and O1 (0.647). In the CNS cohort (n=50), mean composite mGPS favored O2 (0.82) over O3 (0.68) and O1 (0.55), while readability remained stable across models (3.55–3.73). Severe hallucination/unsafe recommendation rate: O1 [26%], O2 [5%], O3 [17%] (pre-specified definition; e.g., major fabricated citation or guideline-discordant harmful recommendation). Conclusions: In a 316-case, multi-domain clinician-scored evaluation, limiting retrieval to oncology guideline sources improved safety-aligned performance relative to unconstrained generation or literature-only anchoring. A guideline-constrained RAG configuration paired with clinician-scored rubrics and tail-risk reporting can serve as a practical health-system vetting mechanism for governance, re-validation, and change control of LLM decision support, including in resource-limited settings where reliable guideline-consistent recommendations are essential.

Article Details

Volume / Issue Vol. 44, Issue 16_suppl
Published June 01, 2026
Pages 1617-1617
ISSN 0732-183X
Publisher Lippincott Williams & Wilkins

Journal Info

Journal of Clinical Oncology

Lippincott Williams & Wilkins

ISSN: 0732-183X Health Sciences

Authors (13)

C

Connor Yost

1Creighton University school of medicine, Phoenix, United States

P

Peter Palumbo

4Geisel School of Medicine at Dartmouth, Hanover, United States

N

Nikita Tripathi

1Creighton University school of medicine, Phoenix, United States

B

Bradley Callas

Creighton University School of Medicine, Phoenix, AZ

Y

Yan Leyfman

5NewYork-Presbyterian Hospital, Hematology, New York, United States

S

Sarah Elizabeth Monick

Mayo Clinic Arizona, Phoenix, AZ

M

Matthew Reeves Sullivan

Dartmouth-Hitchcock Medical Center, Lebanon, NH

Y

Yoshie Umemura

Ivy Brain Tumor Center at Barrow Neurological Institute, Phoenix, AZ

V

Vamsee Torri

S

Samuel A. Funt

Memorial Sloan Kettering Cancer Center, New York, NY

R

Ryan Huu-Tuan Nguyen

University of Illinois College of Medicine at Chicago, Division of Hematology and Oncology, Chicago, IL

I

Irbaz Bin Riaz

Irbaz Bin Riaz, MD, PhD; R. Bryan Rumble, MSc; Thomas A. Hope, MD; Giuseppe Procopio, MD; and Neha Vapiwala, MD; Mayo Clinic, Phoenix, AZ; American Society of Clinical Oncology, Alexandria, VA; University of California, San Francisco, San Francisco, CA; Fondazione IRCCS Istituto Nazionale dei Tumori di Milano, Milan, Italy; and University of Pennsylvania Abramson Cancer Center, Philadelphia, PA

A

Arturo Loaiza-Bonilla

3St. Luke's Cancer Center, Oncology Hematology, Easton, United States