Engineering and evaluating intelligent information retrieval systems for hepatocellular carcinoma clinical question answering.

M Michael Sheen (The Warren Alpert Medical School of Brown University, Providence, RI) A Andrew J. Yang (The Warren Alpert Medical School of Brown University, Providence, RI) S Saajan Patel (The Warren Alpert Medical School of Brown University, Providence, RI) J Jonathan Z. Liu (The Warren Alpert Medical School of Brown University, Providence, RI) T Taemin Kim (Department of Chemistry) A Aaron Seto (The Warren Alpert Medical School of Brown University, Providence, RI) K Khaldoun Almhanna

Abstract

e16009 Background: Large language models (LLMs) help navigate complex clinical guidelines but often suffer from hallucinations. Retrieval-augmented generation (RAG) systems aim to mitigate this, yet systematic evaluations in medical contexts are rare. This study benchmarks four AI architectures against the AASLD hepatocellular carcinoma guidelines to address this gap. Methods: Four AI architectures were evaluated: a baseline LLM; a custom RAG system; a custom multimodal RAG system integrating figures; and a custom agentic multimodal RAG system with autonomous multi-step retrieval. All systems used GPT-5.1. Two test sets were used: 39 guideline-derived question–answer pairs spanning epidemiology, diagnostics, and therapeutics, and 10 complex clinical vignettes requiring integration of multiple guideline components. Each query was run five times and scored using an LLM-as-a-judge three-point scale: 0.0 (incorrect), 0.5 (partially correct), and 1.0 (correct). Mann–Whitney U tests compared systems with the baseline LLM. Results: For guideline questions, performance increased with architectural complexity. Agentic multimodal RAG achieved the highest mean score (0.81±0.31, SD), with 28/39 questions (71.8%) scoring 1.0 and 3/39 (7.7%) scoring 0.0. Multimodal RAG scored 0.74±0.36, traditional RAG 0.62±0.41, and the base LLM 0.51±0.40. All RAG systems significantly outperformed the base LLM (p < 0.001 for multimodal and agentic; p = 0.004 for traditional). Agentic multimodal RAG showed the lowest variability across guideline questions, indicating more consistent performance than other architectures. In contrast, for clinical vignettes (n = 10), the base LLM performed best (0.90±0.24), with 8/10 questions scoring 1.0. All RAG systems significantly underperformed compared to base LLM (p≤0.005): traditional RAG scored 0.63±0.46 (6/10 perfect), agentic multimodal RAG scored 0.73±0.35 (5/10 perfect), and multimodal RAG scored only 0.42±0.44, with 5/10 questions (50%) scoring 0.0. These findings suggest complex clinical reasoning relies more on parametric knowledge and general reasoning rather than document retrieval. Conclusions: These findings reveal a critical dichotomy in medical AI performance. While agentic and multimodal RAG architectures excel at factual extraction and visual data interpretation for specific guideline queries, they falter in complex clinical vignettes, where the base model’s holistic reasoning proves superior. This suggests that retrieval mechanisms can inadvertently fragment context or introduce noise during synthesis tasks. Comparative performance metrics across four AI system architectures. System Guideline Questions (n=39) Mean (SD) Vignettes (n=10) Mean (SD) Base LLM 0.51 (0.40) 0.90 (0.24) RAG Pipeline 0.62 (0.41) 0.63 (0.46) Multimodal RAG 0.74 (0.36) 0.42 (0.44) Agentic Multimodal RAG 0.81 (0.31) 0.73 (0.35)

Article Details

Volume / Issue Vol. 44, Issue 16_suppl
Published June 01, 2026
ISSN 0732-183X
Publisher Lippincott Williams & Wilkins

Journal Info

Journal of Clinical Oncology

Lippincott Williams & Wilkins

ISSN: 0732-183X Health Sciences

Authors (7)

M

Michael Sheen

The Warren Alpert Medical School of Brown University, Providence, RI

A

Andrew J. Yang

The Warren Alpert Medical School of Brown University, Providence, RI

S

Saajan Patel

The Warren Alpert Medical School of Brown University, Providence, RI

J

Jonathan Z. Liu

The Warren Alpert Medical School of Brown University, Providence, RI

T

Taemin Kim

Department of Chemistry

A

Aaron Seto

The Warren Alpert Medical School of Brown University, Providence, RI

K

Khaldoun Almhanna