A comparison of data-processing methods to support patient matching to oncology clinical trials.

M May Pini (Tempus AI, Inc., Chicago, IL) K Kunal Nagpal (Tempus AI, Inc., Chicago, IL) M Mark Vance (Tempus AI, Inc., Chicago, IL) C Caity Moran Rose (Tempus AI, Inc., Chicago, IL) M Michelle Huang (Tempus AI, Inc., Chicago, IL) M Michael Singer (Tempus AI, Inc., Chicago, IL) G Gianna Klonk (Tempus AI, Inc., Chicago, IL) R Ryan Godart (Tempus AI, Inc., Chicago, IL) T Tian Kang (Tempus AI, Inc., Chicago, IL) D Dan Sun A Arpita Saha (Tempus AI, Inc., Chicago, IL) C Chelsea Kendall Osterman (Tempus AI, Inc., Chicago, IL)

Abstract

e13651 Background: To pre-screen and match patients to trials at scale, technology tools are needed to support human workflows. These tools require numerous data inputs, including clinical concepts such as cancer diagnosis, stage, and histology. The more complete these data are, the better the accuracy of matches produced, and the more efficient human reviewers can be. While some clinical data are available from clinical system structured fields, most are within unstructured documentation. Methods including human abstraction, natural language processing (NLP) models, and large language model (LLM) agents can improve data completeness and accuracy beyond structured EHR and laboratory information management system (LIMS) data. We compared these methods to obtain diagnosis and stage data for Tempus Link, a tool that supports patient-trial matching at scale for a national oncology trials network. Methods: We randomly sampled patients with one of 4 previously abstracted cancer diagnoses: lung (LC), prostate (PC), colorectal (CRC), and breast cancer (BC). For LC we also examined histology (NSCLC vs SCLC). For each method (EHR, LIMS, NLP predicted, LLM agent), diagnosis, stage, and histology values were compared to the abstracted value, and accuracy was calculated as the percent of patients with a correct value among patients with an abstracted value available. Since these data are inputs into a tool used by RN screeners who confirm matches before notifying a site, diagnosis was focused on accuracy in broad cancer diagnostic categories (e.g. neoplasm of lung). In addition, we assessed completeness as presence of a usable value across all patients in a cohort. Results: Completeness for EHR, NLP, and LLM was higher for diagnosis compared to stage, ranging from 91.7 - 100% vs 20 - 95.8%. LIMS completeness was lower for all variables across cohorts, ranging from 0 - 87.8%. Conclusions: Variability exists across data sources and processing methods in generating data inputs for trial matching. All methods have high completeness and accuracy for diagnosis, while there are significant gains when applying NLP and LLMs for stage and histology. To balance accuracy, matching efficiency, and cost, the use of enhanced data processing methods are required for certain variables, but may not be needed for others. Accuracy by cohort, variable, and data processing method. Cohort EHR LIMS NLP LLM LCn=48 # diagnosis correct 48 (100%) 33 (68.8%) 48 (100%) 45 (93.8%) # NSCLC correctn=48 1 (2.1%) 27 (56.3%) 48 (100%) 39 (81.3%) # stage correctn=41 16 (39%) 0 (0%) 27 (65.9%) 30 (73.2%) PCn=48 # diagnosis correct 48 (100%) 43 (89.6%) 48 (100%) 44 (91.7%) # stage correctn=45 22 (48.9%) 0 (0%) 15 (33.3%) 38 (84.4%) CRCn=49 # diagnosis correct 47 (95.9%) 42 (85.7%) 47 (95.9%) 46 (93.9%) # stage correctn=46 21 (45.7%) 0 (0%) 36 (78.3%) 42 (91.3%) BCn=50 # diagnosis correct 50 (100%) 9 (18%) 47 (94%) 50 (100%) # stage correctn=36 5 (13.9%) 0 (0%) 23 (63.9%) 35 (97.2%)

Article Details

Volume / Issue Vol. 44, Issue 16_suppl
Published June 01, 2026
ISSN 0732-183X
Publisher Lippincott Williams & Wilkins

Journal Info

Journal of Clinical Oncology

Lippincott Williams & Wilkins

ISSN: 0732-183X Health Sciences

Authors (12)

M

May Pini

Tempus AI, Inc., Chicago, IL

K

Kunal Nagpal

Tempus AI, Inc., Chicago, IL

M

Mark Vance

Tempus AI, Inc., Chicago, IL

C

Caity Moran Rose

Tempus AI, Inc., Chicago, IL

M

Michelle Huang

Tempus AI, Inc., Chicago, IL

M

Michael Singer

Tempus AI, Inc., Chicago, IL

G

Gianna Klonk

Tempus AI, Inc., Chicago, IL

R

Ryan Godart

Tempus AI, Inc., Chicago, IL

T

Tian Kang

Tempus AI, Inc., Chicago, IL

D

Dan Sun

A

Arpita Saha

Tempus AI, Inc., Chicago, IL

C

Chelsea Kendall Osterman

Tempus AI, Inc., Chicago, IL