Use of large language models to extract cancer diagnosis, histology, grade, and staging from unstructured electronic health records.

G Gayathri Namasivayam (Ontada, Boston, MA) Z Zhaohui Su (1Ontada, Boston, United States) A Akhil Bhat (McKesson Specialty Health, The Woodlands, TX) W Wendy Haydon (Ontada, Boston, MA) N Nicholas J. Robert (Ontada, Boston, MA) J Janet L. Espirito (Ontada, Boston, MA)

Abstract

11176 Background: The oncology ecosystem contains millions of unstructured documents within Electronic Health Record (EHR) systems, including clinical notes and pathology reports with vital patient information like cancer diagnoses, histology, grade, and staging. These documents are often text or scanned PDFs. Extracting clinical details from these sources can improve EHR completeness and accuracy. To address this challenge, we developed an information extraction pipeline that leverages the recent advances in artificial intelligence (AI). Methods: Large Language Models (LLMs) and prompt engineering was used to extract information from both clinical notes and pathology reports. For pathology reports, Optical Character Recognition (OCR) was applied to convert scanned images into text, which was then processed by the LLM using a tailored prompt designed to extract the relevant cancer diagnosis and staging details. For clinical notes, the text was directly passed into the LLM along with an optimized prompt to extract the same clinical information. The information extraction pipeline was validated using a dataset of 829 pathology reports and 569 progress notes from the EHR system across 40 cancer types, including 26 solid tumors and 14 hematologic malignancies. Clinical specialists manually extracted cancer diagnoses, histology, grade, and staging, which were compared to the automated output. An F1 score using a combined measure of precision and recall was calculated to assess the model’s accuracy. Results: The pipeline retrieved cancer type, histology, and grade from pathology reports with an F1 score over 0.85. Similarly, it extracted cancer type and staging information from progress notes with an F1 score over 0.85, demonstrating high accuracy and reliability. Additional performance metrics are shown in the table. Conclusions: This LLM-based extraction pipeline accurately identified cancer diagnoses, histology, grade, and staging information from unstructured text and scanned documents within the EHR achieving an F1 score ≥0.85. The validation results suggest that this approach could be scaled to improve the completeness and utility of EHR data, supporting the availability of more robust information for scientific research, clinical care, and other uses of health data. Performance metrics of LLM-based extraction pipeline. Information Source Precision Recall Specificity NPV* Accuracy F1 > Cancer Type Pathology (PDF) 0.86 0.89 0.99 0.98 0.89 0.87 Histology Pathology (PDF) 0.83 0.88 0.90 0.88 0.90 0.85 Grade Pathology (PDF) 0.8 0.95 0.96 0.96 0.88 0.87 Cancer Type Progress Notes 0.8 0.95 0.96 0.96 0.88 0.87 Staging TNM (T) Progress Notes 0.96 0.87 0.99 0.97 0.97 0.91 Staging TNM (N) Progress Notes 0.92 0.85 0.99 0.97 0.96 0.88 Staging TNM (M) Progress Notes 0.87 0.83 0.98 0.97 0.96 0.85 *NPV= negative predictive value.

Article Details

Volume / Issue Vol. 43, Issue 16_suppl
Published June 01, 2025
Pages 11176-11176
ISSN 0732-183X
Publisher Lippincott Williams & Wilkins

Journal Info

Journal of Clinical Oncology

Lippincott Williams & Wilkins

ISSN: 0732-183X Health Sciences

Authors (6)

G

Gayathri Namasivayam

Ontada, Boston, MA

Z

Zhaohui Su

1Ontada, Boston, United States

A

Akhil Bhat

McKesson Specialty Health, The Woodlands, TX

W

Wendy Haydon

Ontada, Boston, MA

N

Nicholas J. Robert

Ontada, Boston, MA

J

Janet L. Espirito

Ontada, Boston, MA