Expert evaluation of LLM world models: A high-T <sub> <i>c</i> </sub> superconductivity case study

H Haoyu Guo (Department of Physics) M Maria Tikhanovskaya (Google) P Paul Raccuglia A Alexey Vlaskin (Google) C Chris Co D Daniel J. Liebling (Google) S Scott Ellsworth M Matthew Abraham E Elizabeth Dorfman (Google) N N. P. Armitage (William H. Miller III Department of Physics and Astronomy) C Chunhan Feng A Antoine Georges (Center for Computational Quantum Physics, Flatiron Institute) O Olivier Gingras (Center for Computational Quantum Physics) D Dominik Kiese (Center for Computational Quantum Physics) S Steven A. Kivelson (Geballe Laboratory for Advanced Materials) V Vadim Oganesyan (Physics Program and Initiative for the Theoretical Sciences) B B. J. Ramshaw S Subir Sachdev T T. Senthil (Department of Physics) J J. M. Tranquada (Condensed Matter Physics and Materials Science Division) M Michael P. Brenner S Subhashini Venugopalan E Eun-Ah Kim (Department of Physics)

Abstract

Large Language Models (LLMs) show great promise as a powerful tool for scientific literature exploration. However, their effectiveness in providing scientifically accurate and comprehensive answers to complex questions within specialized domains remains an active area of research. Using the field of high-temperature cuprates as an exemplar, we evaluate the ability of LLM systems to understand the literature at the level of an expert. We construct an expert-curated database of 1,726 scientific papers that covers the history of the field, and a set of 67 expert-formulated questions that probe deep understanding of the literature. We then evaluate six different LLM-based systems for answering these questions, including both commercially available closed models and a custom retrieval-augmented generation (RAG) system capable of retrieving images alongside text. Experts then evaluate the answers of these systems against a rubric that assesses balanced perspectives, factual comprehensiveness, succinctness, and evidentiary support. Among the six systems, two using RAG on curated literature outperformed existing closed models across key metrics, particularly in providing comprehensive and well-supported answers. We discuss promising aspects of LLM performances as well as critical short-comings of all the models. The set of expert-formulated questions and the rubric will be valuable for assessing expert level performance of LLM based reasoning systems.

Article Details

Volume / Issue Vol. 123, Issue 11
Published March 17, 2026
ISSN 0027-8424
Publisher National Academy of Sciences

Authors (23)

H

Haoyu Guo

Department of Physics

M

Maria Tikhanovskaya

Google

P

Paul Raccuglia

A

Alexey Vlaskin

Google

C

Chris Co

D

Daniel J. Liebling

Google

S

Scott Ellsworth

M

Matthew Abraham

E

Elizabeth Dorfman

Google

N

N. P. Armitage

William H. Miller III Department of Physics and Astronomy

C

Chunhan Feng

A

Antoine Georges

Center for Computational Quantum Physics, Flatiron Institute

O

Olivier Gingras

Center for Computational Quantum Physics

D

Dominik Kiese

Center for Computational Quantum Physics

S

Steven A. Kivelson

Geballe Laboratory for Advanced Materials

V

Vadim Oganesyan

Physics Program and Initiative for the Theoretical Sciences

B

B. J. Ramshaw

S

Subir Sachdev

T

T. Senthil

Department of Physics

J

J. M. Tranquada

Condensed Matter Physics and Materials Science Division

M

Michael P. Brenner

S

Subhashini Venugopalan

E

Eun-Ah Kim

Department of Physics