Using large language models to extract diagnosis, recurrence, and death dates from breast cancer electronic medical records.

J James Dickerson (Stanford Hospital and Clinics, Stanford, CA) M Marni Blair McClure (Stanford Cancer Institute, Stanford, CA) M Margaret Shaw (Stanford University, Stanford, CA) J Jennifer Lee Caswell-Jin (Stanford Cancer Institute, Stanford, CA)

Abstract

e13005 Background: Breast cancer research needs accurate extraction of clinical event dates (e.g., recurrence dates) from the electronic medical record (EMR). Manual chart review is time-consuming and error-prone. We evaluated whether unadapted, HIPAA-compliant, commercially available large language models (LLMs) could extract event dates with accuracy comparable to research coordinators, benchmarked against medical oncologists, in patients with complex breast cancer histories. Methods: We randomly selected 99 patients from an institutional cohort. Ground truth was established by oncologist review of all EMR data, except CareEverywhere. Abstracted events included: original diagnosis date; first locoregional recurrence (LRR); first distant metastatic date; vital status; and an exact or estimated date of death. In parallel, research coordinators abstracted original diagnosis and metastatic dates. For the LLMs, a patient’s full chart was ingested into a patient-specific vector database. Relevant evidence chunks were retrieved using hybrid lexical + semantic retrieval. Chunks were sent to the LLMs with structured, sequential prompts that enforced definitions and prioritized pathology provenance. Estimation was used only if no explicit date was found. Identical evidence and prompts were used with Gemini 2.5, GPT-4, and GPT-5. We report only the best-performing model for each event due to space constraints. Results: All 99 patients had a documented year of original diagnosis. Thirty-one had LRR, 78 developed distant disease, and 13 presented with de novo metastatic disease. As of the censor date (01/10/2026), 83 patients died, with 49 having an exact date of death. Agreement for vital status for the best performing model (GPT-5) was 99%. Time to first recurrence for stage 1-3 patients was 29 [17–49] months (median [IQR]); the best LLM estimate (GPT-5) of recurrence differed by 3.8 ± 12.7 months (mean ± SD) with 87% of events ≤ 90 days. Cohort overall survival was 67 [43–99] months. The best LLM (GPT-5) estimate of survival differed by 2.1 ± 3.5 months, with 71% ≤ 90 days. When the EMR had the exact date of death, concordance was high (98% ≤ 90 days); when estimation was used, concordance was low (28% ≤ 90 days). Original diagnosis date within ± 90 days was 98% for research coordinators and GPT-5. For LRR, the best LLM (Gemini) had an F1 score of 0.80; among correctly identified events, 91% were ≤ 90 days. For 91 metastatic events, the best LLM (GPT-4) had an F1 of 0.96, with 72% ≤ 90 days. Research coordinators had an F1 of 0.96 with 76% ≤ 90 days. Conclusions: Commercial LLMs, when provided the full chart and a provenance-aware prompting framework, can reliably extract explicit dates and clinical details, and are similar to trained research coordinators. However, for complex details in patients with recurrent breast cancer histories spanning years and often institutions, agreement with clinicians remains below thresholds needed for fully automated extraction.

Article Details

Volume / Issue Vol. 44, Issue 16_suppl
Published June 01, 2026
ISSN 0732-183X
Publisher Lippincott Williams & Wilkins

Journal Info

Journal of Clinical Oncology

Lippincott Williams & Wilkins

ISSN: 0732-183X Health Sciences

Authors (4)

J

James Dickerson

Stanford Hospital and Clinics, Stanford, CA

M

Marni Blair McClure

Stanford Cancer Institute, Stanford, CA

M

Margaret Shaw

Stanford University, Stanford, CA

J

Jennifer Lee Caswell-Jin

Stanford Cancer Institute, Stanford, CA