Transforming evidence into answers: An agentic framework to support living clinical practice guidelines.
Abstract
e17004 Background: Living clinical practice guidelines offer the opportunity to keep recommendations aligned with the latest evidence, but they also present a fundamental challenge: as evidence accumulates, panelists cannot sustainably query and synthesize a growing body of literature. In developing ASCO's living metastatic castration-resistant prostate cancer (mCRPC) guideline, we encountered this challenge firsthand. To address it, we developed and evaluated an interactive agentic framework that enables guideline panelists to query evidence in natural language and retrieve structured facts, tables, and figures—with the goal of extending this tool as a clinician-facing companion to enhance guideline accessibility. Methods: The mCRPC systematic review dataset supporting ASCO's clinical practice guideline was used, comprising 212 structured columns capturing trial characteristics, patient populations, and outcomes extracted from 188 references across 88 clinical trials. A large language model (LLM)-based agentic framework was developed employing a coding agent that translates natural language queries into executable scripts for structured evidence retrieval. Claude (Haiku 4.5, Sonnet 4.5, and Opus 4.5) served as the core LLM agents. We constructed a benchmark of 50 fact-based queries with gold-standard answers targeting trial identification, treatment modalities, and endpoint characteristics (e.g., "Which trials evaluated systemic monotherapy?"). Performance was evaluated using exact-match accuracy for final answers, and precision, recall, and F1 score for trial-level retrieval. Results: Across 50 queries, exact-match accuracy ranged from 82% to 86%. Haiku 4.5 and Sonnet 4.5 each achieved 82% (42/50), while Opus 4.5 achieved 86% (43/50). Precision ranged from 0.93–0.97, recall from 0.94–0.96, and F1 scores from 0.91–0.94. Conclusions: This agentic framework enables scalable, auditable, natural language interrogation of guideline evidence, addressing a critical bottleneck in living guideline maintenance. For panelists, it supports rapid, reproducible evidence surveillance during recommendation development. For clinicians, it offers a foundation for an interactive companion tool that transforms static guidelines into queryable resources—enabling personalized evidence retrieval at the point of care. LLM EM Precision Recall F1 claude-haiku-4-5 0.82(42/50) 0.95 0.94 0.91 claude-sonnet-4-5 0.82(42/50) 0.97 0.96 0.93 claude-opus-4-5 0.86(43/50) 0.93 0.96 0.94
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (9)
Syed Arsalan Ahmed Naqvi
Mayo Clinic, Phoenix, AZ
Muhammad Ali Khan
Fouad Nahhat
1Mayo Clinic, Division of Hematology/Oncology, Department of Internal Medicine, Phoenix, United States
Muhammad Uzair Sarfraz
1Mayo Clinic, Division of Hematology/Oncology, Department of Internal Medicine, Phoenix, United States
Priya Kumar
Bryan Bryan Rumble
ASCO, Alexandria, VA
Tom Oliver
ASCO, Alexandria, VA
Irbaz Bin Riaz
Irbaz Bin Riaz, MD, PhD; R. Bryan Rumble, MSc; Thomas A. Hope, MD; Giuseppe Procopio, MD; and Neha Vapiwala, MD; Mayo Clinic, Phoenix, AZ; American Society of Clinical Oncology, Alexandria, VA; University of California, San Francisco, San Francisco, CA; Fondazione IRCCS Istituto Nazionale dei Tumori di Milano, Milan, Italy; and University of Pennsylvania Abramson Cancer Center, Philadelphia, PA
Huan He
National Engineering Laboratory for Druggable Gene and Protein Screening, College of Life Science, Northeast Normal University