Spurious alignment between large language models and brains can emerge from non-robust methods and overlooked confounds

N Nima Hadidi E Ebrahim Feghhi B Bryan H. Song I Idan A. Blank J Jonathan C. Kao

Abstract

Abstract Emerging research seeks to draw neuroscientific insights from the neural predictivity of large language models (LLMs). However, as results rapidly proliferate, there is a growing need for large-scale assessments of their robustness. Here, we analyze a wide range of models and methodological approaches across three widely used neural datasets. We find that the use of shuffled train-test splits has contributed to findings that are influential but spurious. Furthermore, how activations are extracted from LLMs can bias results in favor of specific model classes. Lastly, we find that confounding variables, particularly positional signals and word rate, perform competitively with trained LLMs and fully account for the neural predictivity of untrained LLMs on these neural datasets. Although many studies in the field avoid these pitfalls, our results indicate that some apparent alignment between LLMs and brains has emerged from non-robust methods and overlooked confounds.

Article Details

Volume / Issue Vol. 17, Issue 1
Published April 27, 2026
ISSN 2041-1723
Publisher Nature Portfolio

Journal Info

Nature Communications

Nature Portfolio

ISSN: 2041-1723 Open Access Life Sciences

Authors (5)

N

Nima Hadidi

E

Ebrahim Feghhi

B

Bryan H. Song

I

Idan A. Blank

J

Jonathan C. Kao