Using large language models for grading CTCAE toxicity after radiation therapy for prostate cancer.
Abstract
357 Background: Prostate cancer (PCa) is the most incident cancer in adult males in the United States, and the second highest in the world. CTCAE is the standardized method to grade the severity of treatment-related adverse events (AEs), but are tedious to collect and subject to inter-observer variability. This study aimed to evaluate the performance of LLMs in automatically extracting and grading CTCAE toxicities from clinical notes and patient-reported outcomes of PCa patients from a clinical trial. Methods: A total of 55 patients from NCT02874014 trial, undergoing proton therapy (PT) for localized PCa, were included. The ground truth of CTCAE graded toxicities was obtained from the trial’s results. Clinical notes within 3 days, and the most recent patient questionnaires within 90 days of toxicity assessment were retrieved from the electronic health record. A comprehensive prompt was developed to extract and CTCAE grade toxicities from the retrieved records. LLMs used for this analysis were OpenAI GPT-4o and Gemini 2.0 Flash. Phase 1 evaluated both LLMs' performance against the ground truth. To ensure the metrics reflected the models’ performance based only on the information available in the text provided, cases where both LLMs contradicted the ground truth were manually reviewed. Phase 2 recalculated LLMs' performance metrics using this adapted ground truth. A conservative stepwise ensemble model was developed to leverage the complementary strengths of both LLMs. Results: In the conservative stepwise ensemble model, Gemini 2.0 Flash served as the initial screening platform to identify AEs in step 1. All cases with identified toxicity events were then reviewed and confirmed using OpenAI GPT-4o in step 2. This effectively reduced false positive rates while preserving high sensitivity. The ensemble model achieved strong performance metrics: binary accuracy (0.954), sensitivity (0.961), specificity (0.953), F1 score (0.836), and grade accuracy (0.932). Conclusions: This study demonstrates the feasibility of the use of LLMs for extraction and CTCAE toxicity grading in PCa patients treated with RT. The ensemble model achieved strong performance metrics. Further validation is needed for other cancer diagnosis and treatment modalities. Phase Model Binary Accuracy Precision Sensitivity Specificity F1 Score Grade Accuracy 1 Gemini 0.946 0.672 1.000 0.939 0.804 0.919 1 OpenAI 0.940 0.785 0.637 0.978 0.703 0.895 2 Gemini 0.956 0.735 1.000 0.950 0.848 0.927 2 OpenAI 0.954 0.917 0.679 0.991 0.780 0.903
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (10)
Renthony Wilson
Mayo Clinic Rochester, Rochester, MN
Federico Mastroleo
Mariana Borras Osorio
Department of Radiation Oncology, Mayo Clinic Rochester, Rochester, MN
Shiv Patel
Mayo Clinic Rochester, Rochester, MN
Sarah Peterson
Mi Zhou
Satomi Shiraishi
Mayo Clinic Rochester, Rochester, MN
Andrew Y.K. Foong
Mayo Clinic Rochester, Rochester, MN
David M. Routman
Mayo Clinic Rochester, Rochester, MN
Mark Raymond Waddle
Department of Radiation Oncology, Mayo Clinic Rochester, Rochester, MN