Differential performance of large language models in advanced cardiac life support assessment: A comprehensive multi-dimensional analysis of accuracy, consistency, and visual recognition capabilities

M Murat Genç B Bensu Bulut M Medine Akkan Öz A Ayşenur Gür M Mehmet Yortanlı B Betül Çiğdem Yortanlı O Oguz Sariyildiz R Ramiz Yazici H Hüseyin Mutlu Z Zekeriya Uykan

Abstract

Background Large Language Models (LLMs) have been increasingly adopted in healthcare settings, yet comparative evaluations of their performance in standardized medical assessments remain limited. This study aims to evaluate the accuracy and consistency of four LLMs in answering Advanced Cardiac Life Support (ACLS) questions. Methods In this observational study, 50 ACLS questions were categorized as knowledge-based (n = 29), visual content (n = 12), or case-based (n = 9). Each question was posed to ChatGPT-4o, Gemini 2.0, Claude 3.5, and DeepSeek R1 on three separate occasions to assess consistency. Performance was evaluated using three accuracy metrics: overall accuracy (all three responses correct), strict accuracy (at least two responses correct), and ideal accuracy (at least one response correct). Results ChatGPT-4o demonstrated superior performance with 100% accuracy across all categories and perfect consistency (Fleiss’ Kappa = 1.0). Claude 3.5 achieved 92.0% overall accuracy with excellent consistency (Fleiss’ Kappa = 0.89). Gemini 2.0 showed 86.0% overall accuracy with moderate consistency (Fleiss’ Kappa = 0.58). DeepSeek R1 performed lowest at 70.0% overall accuracy with moderate consistency (Fleiss’ Kappa = 0.58) and failed completely on visual content questions (0%). All models achieved 100% accuracy on knowledge-based questions. Performance differences were statistically significant across models (p < 0.001). Conclusion LLMs demonstrate variable capabilities in ACLS knowledge assessment, with ChatGPT-4o showing exceptional performance. While these models show promise as supplementary tools in resuscitation education and clinical decision support, significant variations in visual recognition capabilities and response consistency highlight the importance of critical evaluation before clinical implementation.

Article Details

Journal PLoS ONE
Volume / Issue Vol. 21, Issue 4
Published April 29, 2026
Pages e0347611
ISSN 1932-6203
Publisher Public Library of Science

Journal Info

PLoS ONE

Public Library of Science

ISSN: 1932-6203 Open Access Health Sciences

Authors (10)

M

Murat Genç

B

Bensu Bulut

M

Medine Akkan Öz

A

Ayşenur Gür

M

Mehmet Yortanlı

B

Betül Çiğdem Yortanlı

O

Oguz Sariyildiz

R

Ramiz Yazici

H

Hüseyin Mutlu

Z

Zekeriya Uykan