Human–AI collectives most accurately diagnose clinical vignettes

N Nikolas Zöller (Center for Adaptive Rationality) J Julian Berger (Center for Adaptive Rationality) I Irving Lin (The Human Diagnosis Project) N Nathan Fu (The Human Diagnosis Project) J Jayanth Komarneni (The Human Diagnosis Project) G Gioele Barabucci (Department of Digital Humanities) K Kyle Laskowski (The Human Diagnosis Project) V Victor Shia (Harvey Mudd College) B Benjamin Harack (Department of Politics and International Relations) E Eugene A. Chu (Kaiser Permanente) V Vito Trianni (Laboratory of Autonomous Robotics and Artificial Life & Collective Intelligence in Natural and Artificial Systems Lab) R Ralf H. J. M. Kurvers (Center for Adaptive Rationality) S Stefan M. Herzog (Center for Adaptive Rationality)

Abstract

AI systems, particularly large language models (LLMs), are increasingly being employed in high-stakes decisions that impact both individuals and society at large, often without adequate safeguards to ensure safety, quality, and equity. Yet LLMs hallucinate, lack common sense, and are biased—shortcomings that may reflect LLMs’ inherent limitations and thus may not be remedied by more sophisticated architectures, more data, or more human feedback. Relying solely on LLMs for complex, high-stakes decisions is therefore problematic. Here, we present a hybrid collective intelligence system that mitigates these risks by leveraging the complementary strengths of human experience and the vast information processed by LLMs. We apply our method to open-ended medical diagnostics, combining 40,762 differential diagnoses made by physicians with the diagnoses of five state-of-the art LLMs across 2,133 text-based medical case vignettes. We show that hybrid collectives of physicians and LLMs outperform both single physicians and physician collectives, as well as single LLMs and LLM ensembles. This result holds across a range of medical specialties and professional experience and can be attributed to humans’ and LLMs’ complementary contributions that lead to different kinds of errors. Our approach highlights the potential for collective human and machine intelligence to improve accuracy in complex, open-ended domains like medical diagnostics.

Article Details

Volume / Issue Vol. 122, Issue 24
Published June 17, 2025
ISSN 0027-8424
Publisher National Academy of Sciences

Authors (13)

N

Nikolas Zöller

Center for Adaptive Rationality

J

Julian Berger

Center for Adaptive Rationality

I

Irving Lin

The Human Diagnosis Project

N

Nathan Fu

The Human Diagnosis Project

J

Jayanth Komarneni

The Human Diagnosis Project

G

Gioele Barabucci

Department of Digital Humanities

K

Kyle Laskowski

The Human Diagnosis Project

V

Victor Shia

Harvey Mudd College

B

Benjamin Harack

Department of Politics and International Relations

E

Eugene A. Chu

Kaiser Permanente

V

Vito Trianni

Laboratory of Autonomous Robotics and Artificial Life & Collective Intelligence in Natural and Artificial Systems Lab

R

Ralf H. J. M. Kurvers

Center for Adaptive Rationality

S

Stefan M. Herzog

Center for Adaptive Rationality