Automated clinical data extraction from oncology trials using large language models.

R Ryumei Nakada (Department of Statistics, Rutgers University, New Brunswick, NJ) F Federico Ferrari M Michelle Ngo (Merck & Co., Inc., Rahway, NJ) X Xiang Peng (Department of Neurobiology, School of Basic Medicine, Tongji Medical College, Huazhong University of Science and Technology) S Shuyan Wan (Merck & Co., Inc., Rahway, NJ) J Junshui Ma (Merck & Co., Inc., Rahway, NJ) T Thomas Jemielita (Merck & Co., Inc., Rahway, NJ) Y Yulia Sidi (Merck & Co., Inc., Rahway, NJ)

Abstract

e13694 Background: Efficient access to structured clinical data is critical for decision making in oncology trials. However, manual extraction from unstructured sources, such as Clinical Study Reports (CSRs) and abstracts, is time-intensive, error-prone, and often incomplete. Recent advancements in large language models (LLMs) have shown promise in automating data extraction, offering enhanced accuracy and scalability. This study evaluates an LLM driven pipeline for extracting essential clinical variables, such as tumor indications, biomarkers, and outcomes, from oncology trial documents. Methods: A novel pipeline leveraging LLMs was developed to automate data extraction from unstructured documents. The system was trained to identify and extract key clinical variables, including Objective Response Rate (ORR), Progression-Free Survival (PFS), and Overall Survival (OS), from 30 internal CSRs and 36 oncology abstracts. Accuracy was benchmarked against manual extraction by calculating a 0-1 loss metric. Results: The pipeline achieved an overall accuracy of 94.9% across 14 numerical variables for CSRs and 97.0% across 47 variables for abstracts, closely matching the accuracy achieved by manual extraction (96.2%). Among CSRs, ORR-related variables demonstrated the highest accuracy (97.3%), while the accuracy for PFS related variables was slightly lower (92.3%) due to document formatting inconsistencies. The average processing time was 3–5 minutes per CSR and 30 seconds per abstract, with costs ranging from $1 to $8 per document. This represents a significant improvement in efficiency compared to manual human extraction, which takes about one hour per CSR document on average. Conclusions: This LLM-powered pipeline provides a scalable, fast and cost-effective solution for extracting structured data from unstructured clinical documents in oncology. By achieving accuracy comparable to manual extraction and significantly reducing processing time, the pipeline has the potential to transform data workflows in oncology trials, facilitating timely and informed clinical decision-making. Future developments will address challenges in table parsing and further enhance accuracy for complex data structures.

Article Details

Volume / Issue Vol. 43, Issue 16_suppl
Published June 01, 2025
ISSN 0732-183X
Publisher Lippincott Williams & Wilkins

Journal Info

Journal of Clinical Oncology

Lippincott Williams & Wilkins

ISSN: 0732-183X Health Sciences

Authors (8)

R

Ryumei Nakada

Department of Statistics, Rutgers University, New Brunswick, NJ

F

Federico Ferrari

M

Michelle Ngo

Merck & Co., Inc., Rahway, NJ

X

Xiang Peng

Department of Neurobiology, School of Basic Medicine, Tongji Medical College, Huazhong University of Science and Technology

S

Shuyan Wan

Merck & Co., Inc., Rahway, NJ

J

Junshui Ma

Merck & Co., Inc., Rahway, NJ

T

Thomas Jemielita

Merck & Co., Inc., Rahway, NJ

Y

Yulia Sidi

Merck & Co., Inc., Rahway, NJ