Conversational AI-assisted exploratory data analysis (CA-EDA): A comparative evaluation of commercial large language models for clinical trial analysis.

Y Yaacov Lawrence (Chaim Sheba Medical Center, Ramat Gan, Israel) Y Yakir Malyanker (Sheba Medical Center, Ramat Gan, Israel) O Ofer Margalit (Sheba Medical Center, Ramat Gan, Israel) A Aharon M. Lawrence (Azrieli Faculty of Medicine, Bar-Ilan University, Safed, Israel) A Adam P. Dicker A Ayelet Geva (Sheba Medical Center, Ramat Gan, Israel) A Aviad Buskila (Sheba Medical Center, Ramat Gan, Israel) R Raanan Rephael Berger (Chaim Sheba Medical Center, Ramat Gan, Israel)

Abstract

1636 Background: Conventional analysis of clinical trial data requires statistical expertise, time and effort. Previous AI approaches have used structured, pre-processed datasets on specialized platforms. We hypothesized that widely available commercial large language models (LLMs) would perform rapid analysis of raw data. We evaluated three commercial LLM platforms fed identical data, and subsequently validated the outcomes with a formal statistical analysis. We name this approach "Conversational AI-assisted Exploratory Data Analysis" (CA-EDA). Methods: We used data from a completed phase 2 trial (NCT03323489; Lancet Oncol 2024; 25:1070-9); n=125; 2165 variables) that evaluated the efficacy of celiac plexus radiosurgery in controlling retroperitoneal pain syndrome amongst pancreatic cancer patients. Raw individual patient data was uploaded to three commercial LLMs: Claude Opus 4.1 (Anthropic), ChatGPT 5.2 (OpenAI) in deep research mode, and Gemini 3 Pro (Google). Each LLM received identical inputs: the raw Excel dataset, the codebook, the protocol, and the published manuscript. The prompt requested identification of baseline predictors of pain response and development of a clinical score. Outputs were compared for statistical analyses performed, predictors identified, visualizations generated, and clinical utility. Findings were formally verified using Stata IC/16.1. Results: Analysis completion time ranged from 10 to 30 minutes across all three platforms. Claude performed a comprehensive statistical analysis, generated a 9-panel visualization dashboard, identified prior exposure to neurotoxic chemotherapy as a novel predictor of response (OR 0.20, p<0.001), and created a 4-variable clinical score predictive of response (AUC 0.71). ChatGPT performed a statistical analysis but missed neurotoxic chemotherapy, creating a 3-variable score. Gemini produced an 8-page narrative with biological insights, but did not perform any statistical analysis. Analysis using Stata confirmed the association between prior exposure to neurotoxic chemotherapy and response (35% with prior exposure vs 75% without, p<0.001). Conclusions: Claude Opus 4.1 identified a clinically significant predictor that had not previously been recognized, and developed a helpful clinical score. CA-EDA always requires formal biostatistical verification due to concerns about LLMs’ tendency to hallucinate, may analyse only sampled data and potentially implement incorrect Python code. Despite these concerns, here we found CA-EDA of clinical trial data to be rapid, accurate and cost-effective. Financial support for clinical trial: Gateway for Cancer Research, The Israel Cancer Association.

Article Details

Volume / Issue Vol. 44, Issue 16_suppl
Published June 01, 2026
Pages 1636-1636
ISSN 0732-183X
Publisher Lippincott Williams & Wilkins

Journal Info

Journal of Clinical Oncology

Lippincott Williams & Wilkins

ISSN: 0732-183X Health Sciences

Authors (8)

Y

Yaacov Lawrence

Chaim Sheba Medical Center, Ramat Gan, Israel

Y

Yakir Malyanker

Sheba Medical Center, Ramat Gan, Israel

O

Ofer Margalit

Sheba Medical Center, Ramat Gan, Israel

A

Aharon M. Lawrence

Azrieli Faculty of Medicine, Bar-Ilan University, Safed, Israel

A

Adam P. Dicker

A

Ayelet Geva

Sheba Medical Center, Ramat Gan, Israel

A

Aviad Buskila

Sheba Medical Center, Ramat Gan, Israel

R

Raanan Rephael Berger

Chaim Sheba Medical Center, Ramat Gan, Israel