The accuracy and efficiency of large language models for chart review in cancer genetics.

J James Dickerson (Stanford Hospital and Clinics, Stanford, CA) M Margaret Shaw (Stanford University, Stanford, CA) M Mina Satoyoshi (Stanford University, Stanford, CA) S Sonia Rios-Ventura (Stanford School of Medicine, Stanford, CA) K Kerry Kingham (Stanford University School of Medicine, Department of Pediatrics, Stanford, CA) A Allison W. Kurian (Stanford Cancer Institute, Stanford University School of Medicine, Stanford, CA) J Jennifer Lee Caswell-Jin (Stanford Cancer Institute, Stanford, CA)

Abstract

e22603 Background: Constructing databases is crucial for answering clinical questions but is time-consuming and error-prone. Our institution has maintained a REDCap database of cancer genetics encounters since 2002, manually curated by research assistants. We explored automating some data entry using a HIPAA-compliant, commercially available large language model (LLM). Methods: We randomly selected 100 patients from our database since 2017; a board-certified oncologist reviewed each chart to establish a gold standard. We examined variable abstraction for (1) whether genetic testing was ordered, (2) whether genetic testing results were obtained, (3) whether a variant was identified and, if so, the (4) gene and (5) variant status (benign, uncertain significance, or pathogenic). For the LLM input, we provided every Epic note and letter from January 2017 to January 2025 from the Cancer Genetics group (n = 308) for the 100 patients. For patients with multiple notes, we took (1) concordant values from ≥ 2 notes or (2) a non-benign variant as the true LLM result. We made two API calls per note using Stanford Healthcare Secure GPT with OpenAI’s gpt-4o model. The code is available at https://github.com/MrJimb0/ASCO2025 . We calculated summary statistics for time, token use, accuracy, and sensitivity/specificity, with the oncologist chart review as the reference. Results: The LLM accurately categorized 88% of the 100 patients compared to 87% by research assistants in REDCap. LLM errors that occurred in more than one patient were from information being outside of the provided notes (n = 4), information being in an image never converted to text (n = 2), and incorrectly interpreting a familial variant as being the patients’ (n = 2). In contrast, errors in REDCap were from new results returning after the date the research assistant did data entry (n = 7) and typos (n = 5). 29% of the cohort had a pathogenic variant. The LLM had a sensitivity of 83% and specificity of 96% for pathogenic variant detection, compared to 76% and 100% for REDCap. The LLM processed an average of 9,801 input tokens and 372 output tokens per patient, processing each patient in approximately 24 seconds. For a research assistant, the average time was 6 minutes per patient. Assuming 2,500 patients in a year, typical for this clinic, the LLM would take 16.5 hours of work at around $72 compared to 250 hours, or $7,500 of effort, for a research assistant. Conclusions: Compared to abstraction by a research assistant, the LLM was quicker and had similar sensitivity and specificity for these five variables. We obtained these results without hyperparameter tuning, vectorization, note standardization, model retraining, or the development of a foundational model. These results suggest that commercial LLMs with limited prompt engineering and post-LLM processing can support chart review in cancer genetics, potentially reducing costs and improving the efficiency of database construction.

Article Details

Volume / Issue Vol. 43, Issue 16_suppl
Published June 01, 2025
ISSN 0732-183X
Publisher Lippincott Williams & Wilkins

Journal Info

Journal of Clinical Oncology

Lippincott Williams & Wilkins

ISSN: 0732-183X Health Sciences

Authors (7)

J

James Dickerson

Stanford Hospital and Clinics, Stanford, CA

M

Margaret Shaw

Stanford University, Stanford, CA

M

Mina Satoyoshi

Stanford University, Stanford, CA

S

Sonia Rios-Ventura

Stanford School of Medicine, Stanford, CA

K

Kerry Kingham

Stanford University School of Medicine, Department of Pediatrics, Stanford, CA

A

Allison W. Kurian

Stanford Cancer Institute, Stanford University School of Medicine, Stanford, CA

J

Jennifer Lee Caswell-Jin

Stanford Cancer Institute, Stanford, CA