Evaluation of a large language model–generated after-visit summary from outpatient hematology/oncology progress notes.
Abstract
e13692 Background: Oncology progress notes are technical and may be difficult for patients to interpret. We evaluated readability, patient-education quality, and clinician-rated completeness, accuracy, and safety of a large language model (LLM)-generated after-visit summary (AVS) derived from outpatient hematology/oncology progress notes. Methods: Single-center retrospective study of 150 de-identified progress notes from adult outpatient hematology/oncology encounters (IRB #PRO00041351; waiver of consent granted for retrospective de-identified review). AVS were generated in a HIPAA-compliant environment using Google AI Studio Gemini 3 Pro Preview (temperature = 1; no post-processing) and were not used for clinical care. Flesch Reading Ease (FRE) and PEMAT understandability/actionability (0-100) were computed for each note and AI AVS; paired comparisons used Wilcoxon signed-rank tests. One independent oncology clinician reviewed each AI AVS versus its source note using a 15-item Oncology AI-AVS Completeness & Accuracy Survey (completeness = % items present; mean item accuracy 0-2) and assigned safety risk (low/moderate/high). Results: Patients had median age 64.5 years (IQR 54.2-72.0); 66.0% female; 64.7% stage IV. AI AVS had higher FRE than progress notes (mean 64.9 vs 41.3; mean paired diff +23.6, 95% CI 22.6-24.6; p < 0.001). PEMAT understandability (85.7 vs 32.2; diff +53.5; p < 0.001) and actionability (60.8 vs 16.9; diff +43.9; p < 0.001) were higher; understandability improved in 150/150 pairs and actionability in 149/150. Clinician review: completeness mean 95.4% (SD 4.4) and mean item accuracy 1.93/2 (SD 0.07). Safety risk was low in 135 (90.0%) and moderate in 15 (10.0%); none high. Conclusions: An LLM-generated AVS derived from outpatient oncology progress notes was associated with improved readability and PEMAT scores while maintaining high clinician-rated completeness and accuracy; a minority were rated moderate risk, supporting continued clinician oversight and prospective evaluation. Outcomes (N=150). Metric Note AI AVS Diff/P FRE mean (SD) 41.3 (8.3) 64.9 (6.5) +23.6; < 0.001 PEMAT understandability % mean (SD) 32.2 (12.1) 85.7 (10.2) +53.5; < 0.001 PEMAT actionability % mean (SD) 16.9 (8.9) 60.8 (16.0) +43.9; < .0001 Completeness % mean (SD) -- 95.4 (4.4) -- Accuracy (0-2) mean (SD) -- 1.93 (0.07) -- Safety risk n (%) -- Low 135 (90); Moderate 15 (10); High 0 --
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (6)
Akhil Mehta
Houston Methodist Hospital Neal Cancer Center, Houston, TX
Milan Sheth
Houston Methodist Hospital Neal Cancer Center, Houston, TX
Minhal Zaidi
Houston Methodist Hospital, Houston, TX
Colin Chan
Houston Methodist Hospital, Houston, TX
Faizan Azim
Houston Methodist Cancer Center, Houston, TX
Ryan Blair Kieser
Houston Methodist Neal Cancer Center, Houston, TX