Automated clinical data extraction from oncology trials using large language models.
Abstract
e13694 Background: Efficient access to structured clinical data is critical for decision making in oncology trials. However, manual extraction from unstructured sources, such as Clinical Study Reports (CSRs) and abstracts, is time-intensive, error-prone, and often incomplete. Recent advancements in large language models (LLMs) have shown promise in automating data extraction, offering enhanced accuracy and scalability. This study evaluates an LLM driven pipeline for extracting essential clinical variables, such as tumor indications, biomarkers, and outcomes, from oncology trial documents. Methods: A novel pipeline leveraging LLMs was developed to automate data extraction from unstructured documents. The system was trained to identify and extract key clinical variables, including Objective Response Rate (ORR), Progression-Free Survival (PFS), and Overall Survival (OS), from 30 internal CSRs and 36 oncology abstracts. Accuracy was benchmarked against manual extraction by calculating a 0-1 loss metric. Results: The pipeline achieved an overall accuracy of 94.9% across 14 numerical variables for CSRs and 97.0% across 47 variables for abstracts, closely matching the accuracy achieved by manual extraction (96.2%). Among CSRs, ORR-related variables demonstrated the highest accuracy (97.3%), while the accuracy for PFS related variables was slightly lower (92.3%) due to document formatting inconsistencies. The average processing time was 3–5 minutes per CSR and 30 seconds per abstract, with costs ranging from $1 to $8 per document. This represents a significant improvement in efficiency compared to manual human extraction, which takes about one hour per CSR document on average. Conclusions: This LLM-powered pipeline provides a scalable, fast and cost-effective solution for extracting structured data from unstructured clinical documents in oncology. By achieving accuracy comparable to manual extraction and significantly reducing processing time, the pipeline has the potential to transform data workflows in oncology trials, facilitating timely and informed clinical decision-making. Future developments will address challenges in table parsing and further enhance accuracy for complex data structures.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (8)
Ryumei Nakada
Department of Statistics, Rutgers University, New Brunswick, NJ
Federico Ferrari
Michelle Ngo
Merck & Co., Inc., Rahway, NJ
Xiang Peng
Department of Neurobiology, School of Basic Medicine, Tongji Medical College, Huazhong University of Science and Technology
Shuyan Wan
Merck & Co., Inc., Rahway, NJ
Junshui Ma
Merck & Co., Inc., Rahway, NJ
Thomas Jemielita
Merck & Co., Inc., Rahway, NJ
Yulia Sidi
Merck & Co., Inc., Rahway, NJ