From text to insight: Methods for validation of NLP-extracted real-world oncology data using a human-in-the-loop electronic review tool.
Abstract
e23291 Background: Natural Language Processing (NLP) for extraction of oncology-specific data-phenotypes from free-text in electronic medical records (EMR) is an efficient method for generating real-world data (RWD) for use in oncology research. NLP-extracted RWD enables rapid research outcomes to be examined across numerous patient sub-populations. However, validation of rules-based NLP algorithms to ensure accuracy in a complex therapeutic area like oncology is time-intensive. This time-intensive validation can delay deployment of NLP algorithms for extraction of research-grade RWD, or in some cases, this important validation is not performed, putting well informed decision-making at risk. Building on prior validation of NLP algorithms identifying AJCC Cancer stage, TNM stage, and Tumor grade validated by clinical chart abstractors, this study aims to increase efficiency of oncology-specific RWD-generation by defining methods for NLP-algorithm validation using an “expert-in-the-loop” review tool. Methods: An electronic review tool with an interactive user-interface was used to facilitate validation of NLP-extracted measures of Cancer stage, TNM stage, and Tumor grade. The review tool was designed to allow “expert-in-the-loop" chart abstraction by comparing non-identifiable source text against NLP-extracted measures, with clinical oncology subject matter experts defining true positive and true negative benchmarks. This tool was used to annotate non-identifiable free-text notes (n = 300 [TNM stage, detail]; n = 600 [grade]) from a single hospital in the U.S. Clinical oncology subject matter experts compared annotated results to the NLP-extracted measures, making iterative improvements to the NLP algorithms until a target positive predictive value (PPV) of > 90% was reached and maintained. For each NLP-extracted measure, we assessed true positives (TP), false positives (FP), true negatives (TN), false negatives (FN), negative predictive value (NPV), PPV, and F1 score. Results: The PPV for Cancer stage was 100% (n = 47 TP, n = 0 FP), NPV 94% (n = 240 TN, n = 16 FN), and F1 85%. The PPV for TNM stage was 98% (n = 175 TP, n = 4 FP), NPV 83% (n = 209 TN, n = 43 FN), and F1 88%. The PPV for Tumor grade was 96% (n = 366 TP, n = 16 FP), NPV 80% (n = 220 TN, n = 54 FN), and F1 91%. Conclusions: The NLP algorithms used to extract measures of Cancer stage, TNM stage, and Tumor grade from EMR free-text performed exceptionally well as assessed by PPV, NPV, and F1, with refinement of algorithms and efficient validation supported by an “expert-in-the-loop” review tool. These methods illustrate the effectiveness of validated NLP algorithms in quickly gathering real-world oncology data for timely evidence generation across clinically meaningful sub-populations.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (7)
Elizabeth Eldridge
IQVIA, Boston, MA
Elizabeth Trowbridge
IQVIA, South Bend, IN
Brian Berns
IQVIA, Bethesda, MD
Julien Heidt
IQVIA, Durham, NC
Nicholas Stemkowski
IQVIA, Charleston, SC
Celeste Adams
IQVIA, Denver, CO
Christina D. Mack
IQVIA, Durham, NC