Utilization of large language models to facilitate guideline-directed therapy of genitourinary cancer screening and primary workup.

M Matthew Driscoll (Loyola University Chicago Stritch School of Medicine, Maywood, IL) K Keenan Moore (Loyola University Chicago Stritch School of Medicine, Maywood, IL) M Milan K Patel (Loyola University Medical Center, Maywood, IL) K Kristin Baldea (Loyola University Medical Center, Maywood, IL) A Ahmer Farooq (Loyola University Medical Center, Maywood, IL) A Ahmad El-Arabi (Loyola University Medical Center, Maywood, IL)

Abstract

e13705 Background: The rise of large language models (LLMs) offers a unique opportunity to streamline healthcare and simplify medical decision-making. These AI tools are increasingly used in clinical medicine, supporting physicians in tasks ranging from diagnosis to treatment planning. Here, the authors investigate the ability of four common LLMs to direct genitourinary cancer screening and primary workup in alignment with the American Urological Association’s (AUA) Clinical Practice Guidelines (CPGs). Methods: Four LLMs - ChatGPT 4o, Claude 3.5 Sonnet, Google Gemini 1.5 Flash, and OpenEvidence - were selected for evaluation on the basis of popularity, accessibility, and target audience. AUA guidelines for Microhematuria and Early Detection of Prostate Cancer were reworded into question prompts and submitted to each LLM. All queries were conducted within a 72-hour time frame to minimize variability due to LLM adaptation. Two board certified urologists blindly reviewed and assigned responses a rating of concordant, discordant, or indeterminate with AUA CPGs. The responses were evaluated using chi-square tests with the assumption that the LLMs were equally as likely to generate concordant, discordant, or indeterminate responses. Interrater agreement was assessed using a Cohen’s kappa. Results: Of the total 456 unique assessments of LLM performance made, 95% of the assessments were concordant with the AUA CPGs. Each of the four LLMs performed significantly better than chance (p < 0.00001, α < 0.05). Of the four LLMs, ChatGPT produced the highest rate of responses noted as concordant by both reviewers at 98.6%, compared to Gemini which produced the lowest at 84.2%. ChatGPT was the only of the four LLMs to not produce a response that was rated as discordant by both reviewers whereas the other three LLMs each produced just one response that was unanimously discordant. Interrater agreement between reviewers was strong, in line with the moderate range (Cohen’s kappa = 0.506). Lastly, there was a statistically significant difference in the concordance with AUA CPGs between the four LLMs (p=0.0001). However, there was no statistically significant difference with concordance between ChatGPT, OpenEvidence, and Claude responses themselves (p = 0.06). Conclusions: These results suggest that LLMs are significantly making recommendations in line with AUA CPGs. These findings indicate that LLMs could play a valuable role in assisting physicians with the screening and initial evaluation of common genitourinary cancers, particularly those that present in outpatient or primary care settings. It is important to recognize that LLMs are continuously evolving and improving with each query. While this ongoing learning enhances their potential to become more accurate and reliable over time, it also underscores the need for continuous oversight to ensure their responses remain aligned with CPGs.

Article Details

Volume / Issue Vol. 43, Issue 16_suppl
Published June 01, 2025
ISSN 0732-183X
Publisher Lippincott Williams & Wilkins

Journal Info

Journal of Clinical Oncology

Lippincott Williams & Wilkins

ISSN: 0732-183X Health Sciences

Authors (6)

M

Matthew Driscoll

Loyola University Chicago Stritch School of Medicine, Maywood, IL

K

Keenan Moore

Loyola University Chicago Stritch School of Medicine, Maywood, IL

M

Milan K Patel

Loyola University Medical Center, Maywood, IL

K

Kristin Baldea

Loyola University Medical Center, Maywood, IL

A

Ahmer Farooq

Loyola University Medical Center, Maywood, IL

A

Ahmad El-Arabi

Loyola University Medical Center, Maywood, IL