Evaluation of a large language model–generated after-visit summary from outpatient hematology/oncology progress notes.

A Akhil Mehta (Houston Methodist Hospital Neal Cancer Center, Houston, TX) M Milan Sheth (Houston Methodist Hospital Neal Cancer Center, Houston, TX) M Minhal Zaidi (Houston Methodist Hospital, Houston, TX) C Colin Chan (Houston Methodist Hospital, Houston, TX) F Faizan Azim (Houston Methodist Cancer Center, Houston, TX) R Ryan Blair Kieser (Houston Methodist Neal Cancer Center, Houston, TX)

Abstract

e13692 Background: Oncology progress notes are technical and may be difficult for patients to interpret. We evaluated readability, patient-education quality, and clinician-rated completeness, accuracy, and safety of a large language model (LLM)-generated after-visit summary (AVS) derived from outpatient hematology/oncology progress notes. Methods: Single-center retrospective study of 150 de-identified progress notes from adult outpatient hematology/oncology encounters (IRB #PRO00041351; waiver of consent granted for retrospective de-identified review). AVS were generated in a HIPAA-compliant environment using Google AI Studio Gemini 3 Pro Preview (temperature = 1; no post-processing) and were not used for clinical care. Flesch Reading Ease (FRE) and PEMAT understandability/actionability (0-100) were computed for each note and AI AVS; paired comparisons used Wilcoxon signed-rank tests. One independent oncology clinician reviewed each AI AVS versus its source note using a 15-item Oncology AI-AVS Completeness & Accuracy Survey (completeness = % items present; mean item accuracy 0-2) and assigned safety risk (low/moderate/high). Results: Patients had median age 64.5 years (IQR 54.2-72.0); 66.0% female; 64.7% stage IV. AI AVS had higher FRE than progress notes (mean 64.9 vs 41.3; mean paired diff +23.6, 95% CI 22.6-24.6; p < 0.001). PEMAT understandability (85.7 vs 32.2; diff +53.5; p < 0.001) and actionability (60.8 vs 16.9; diff +43.9; p < 0.001) were higher; understandability improved in 150/150 pairs and actionability in 149/150. Clinician review: completeness mean 95.4% (SD 4.4) and mean item accuracy 1.93/2 (SD 0.07). Safety risk was low in 135 (90.0%) and moderate in 15 (10.0%); none high. Conclusions: An LLM-generated AVS derived from outpatient oncology progress notes was associated with improved readability and PEMAT scores while maintaining high clinician-rated completeness and accuracy; a minority were rated moderate risk, supporting continued clinician oversight and prospective evaluation. Outcomes (N=150). Metric Note AI AVS Diff/P FRE mean (SD) 41.3 (8.3) 64.9 (6.5) +23.6; < 0.001 PEMAT understandability % mean (SD) 32.2 (12.1) 85.7 (10.2) +53.5; < 0.001 PEMAT actionability % mean (SD) 16.9 (8.9) 60.8 (16.0) +43.9; < .0001 Completeness % mean (SD) -- 95.4 (4.4) -- Accuracy (0-2) mean (SD) -- 1.93 (0.07) -- Safety risk n (%) -- Low 135 (90); Moderate 15 (10); High 0 --

Article Details

Volume / Issue Vol. 44, Issue 16_suppl
Published June 01, 2026
ISSN 0732-183X
Publisher Lippincott Williams & Wilkins

Journal Info

Journal of Clinical Oncology

Lippincott Williams & Wilkins

ISSN: 0732-183X Health Sciences

Authors (6)

A

Akhil Mehta

Houston Methodist Hospital Neal Cancer Center, Houston, TX

M

Milan Sheth

Houston Methodist Hospital Neal Cancer Center, Houston, TX

M

Minhal Zaidi

Houston Methodist Hospital, Houston, TX

C

Colin Chan

Houston Methodist Hospital, Houston, TX

F

Faizan Azim

Houston Methodist Cancer Center, Houston, TX

R

Ryan Blair Kieser

Houston Methodist Neal Cancer Center, Houston, TX