Using large language models for grading CTCAE toxicity after radiation therapy for prostate cancer.

R Renthony Wilson (Mayo Clinic Rochester, Rochester, MN) F Federico Mastroleo M Mariana Borras Osorio (Department of Radiation Oncology, Mayo Clinic Rochester, Rochester, MN) S Shiv Patel (Mayo Clinic Rochester, Rochester, MN) S Sarah Peterson M Mi Zhou S Satomi Shiraishi (Mayo Clinic Rochester, Rochester, MN) A Andrew Y.K. Foong (Mayo Clinic Rochester, Rochester, MN) D David M. Routman (Mayo Clinic Rochester, Rochester, MN) M Mark Raymond Waddle (Department of Radiation Oncology, Mayo Clinic Rochester, Rochester, MN)

Abstract

357 Background: Prostate cancer (PCa) is the most incident cancer in adult males in the United States, and the second highest in the world. CTCAE is the standardized method to grade the severity of treatment-related adverse events (AEs), but are tedious to collect and subject to inter-observer variability. This study aimed to evaluate the performance of LLMs in automatically extracting and grading CTCAE toxicities from clinical notes and patient-reported outcomes of PCa patients from a clinical trial. Methods: A total of 55 patients from NCT02874014 trial, undergoing proton therapy (PT) for localized PCa, were included. The ground truth of CTCAE graded toxicities was obtained from the trial’s results. Clinical notes within 3 days, and the most recent patient questionnaires within 90 days of toxicity assessment were retrieved from the electronic health record. A comprehensive prompt was developed to extract and CTCAE grade toxicities from the retrieved records. LLMs used for this analysis were OpenAI GPT-4o and Gemini 2.0 Flash. Phase 1 evaluated both LLMs' performance against the ground truth. To ensure the metrics reflected the models’ performance based only on the information available in the text provided, cases where both LLMs contradicted the ground truth were manually reviewed. Phase 2 recalculated LLMs' performance metrics using this adapted ground truth. A conservative stepwise ensemble model was developed to leverage the complementary strengths of both LLMs. Results: In the conservative stepwise ensemble model, Gemini 2.0 Flash served as the initial screening platform to identify AEs in step 1. All cases with identified toxicity events were then reviewed and confirmed using OpenAI GPT-4o in step 2. This effectively reduced false positive rates while preserving high sensitivity. The ensemble model achieved strong performance metrics: binary accuracy (0.954), sensitivity (0.961), specificity (0.953), F1 score (0.836), and grade accuracy (0.932). Conclusions: This study demonstrates the feasibility of the use of LLMs for extraction and CTCAE toxicity grading in PCa patients treated with RT. The ensemble model achieved strong performance metrics. Further validation is needed for other cancer diagnosis and treatment modalities. Phase Model Binary Accuracy Precision Sensitivity Specificity F1 Score Grade Accuracy 1 Gemini 0.946 0.672 1.000 0.939 0.804 0.919 1 OpenAI 0.940 0.785 0.637 0.978 0.703 0.895 2 Gemini 0.956 0.735 1.000 0.950 0.848 0.927 2 OpenAI 0.954 0.917 0.679 0.991 0.780 0.903

Article Details

Volume / Issue Vol. 44, Issue 7_suppl
Published March 01, 2026
Pages 357-357
ISSN 0732-183X
Publisher Lippincott Williams & Wilkins

Journal Info

Journal of Clinical Oncology

Lippincott Williams & Wilkins

ISSN: 0732-183X Health Sciences

Authors (10)

R

Renthony Wilson

Mayo Clinic Rochester, Rochester, MN

F

Federico Mastroleo

M

Mariana Borras Osorio

Department of Radiation Oncology, Mayo Clinic Rochester, Rochester, MN

S

Shiv Patel

Mayo Clinic Rochester, Rochester, MN

S

Sarah Peterson

M

Mi Zhou

S

Satomi Shiraishi

Mayo Clinic Rochester, Rochester, MN

A

Andrew Y.K. Foong

Mayo Clinic Rochester, Rochester, MN

D

David M. Routman

Mayo Clinic Rochester, Rochester, MN

M

Mark Raymond Waddle

Department of Radiation Oncology, Mayo Clinic Rochester, Rochester, MN