Evaluating the statistical realism of LLM-generated social science data

Y Yueqi Xie (Department of Computer Science & Engineering) L Lemeng Liang (Department of Sociology) S Shuzhen Li (Department of Electrical & Computer Engineering) Y Yifu Lu (Department of Electrical & Computer Engineering) Z Zhiwen Xiao M Mengdi Shi (Center for Social Research, Guanghua School of Management, Peking University) J Junming Huang (Paul and Marcia Center on Contemporary China) M Mengdi Wang (Department of Electrical and Computer Engineering, Princeton University, Princeton, NJ, USA.) Y Yu Xie

Abstract

Large language models (LLMs) hold great promise for generating social science data, potentially expanding the methodological toolkit of quantitative social research. Prior studies have primarily focused on individual-level predictability or behavioral plausibility of LLM-generated data. We propose a framework for assessing the validity of LLM-generated data by returning to the foundational principles of survey research in the social sciences. Just as surveys based on representative samples yield statistics that approximate the corresponding statistical moments of the target population, assessment should center on the ability of LLM-generated data to reproduce real-world, population-level statistical patterns. We introduce SSDataBench, a systematic benchmark designed to evaluate population-level statistical realism in LLM-generated social science data. The benchmark assesses five types of statistical patterns central to social research: univariate distributions, bivariate associations, multivariate outcome predictions, life event sequence distributions, and associations between life event sequences and covariates. We illustrate SSDataBench using four longitudinal datasets and three cross-sectional datasets spanning six major social domains: demographics, socioeconomic status, marriage, health, abilities, and attitudes. Our study reveals representational limitations in current LLMs under sparse conditioning settings, manifested in a pronounced tendency to compress real-world heterogeneity into simplified typological structures. Finally, we outline a roadmap toward improved statistical realism and report preliminary results indicating that domain-specific training can enhance population-level realism.

Article Details

Volume / Issue Vol. 123, Issue 19
Published May 12, 2026
ISSN 0027-8424
Publisher National Academy of Sciences

Authors (9)

Y

Yueqi Xie

Department of Computer Science & Engineering

L

Lemeng Liang

Department of Sociology

S

Shuzhen Li

Department of Electrical & Computer Engineering

Y

Yifu Lu

Department of Electrical & Computer Engineering

Z

Zhiwen Xiao

M

Mengdi Shi

Center for Social Research, Guanghua School of Management, Peking University

J

Junming Huang

Paul and Marcia Center on Contemporary China

M

Mengdi Wang

Department of Electrical and Computer Engineering, Princeton University, Princeton, NJ, USA.

Y

Yu Xie