Development and validation of large language model rating scales for automatically transcribed psychological therapy sessions

S Steffen T. Eberhardt A Antonia Vehlen J Jana Schaffrath B Brian Schwartz T Tobias Baur D Dominik Schiller T Tobias Hallmen E Elisabeth André (University of Augsburg) W Wolfgang Lutz (Asian Demographic Research Institute, Shanghai University)

Abstract

Abstract Rating scales have shaped psychological research, but are resource-intensive and can burden participants. Large Language Models (LLMs) offer a tool to assess latent constructs in text. This study introduces LLM rating scales, which use LLM responses instead of human ratings. We demonstrate this approach with an LLM rating scale measuring patient engagement in therapy transcripts. Automatically transcribed videos of 1,131 sessions from 155 patients were analyzed using DISCOVER, a software framework for local multimodal human behavior analysis. Llama 3.1 8B LLM rated 120 engagement items, averaging the top eight into a total score. Psychometric evaluation showed a normal distribution, strong reliability (ω = 0.953), and acceptable fit (CFI = 0.968, SRMR = 0.022), except RMSEA = 0.108. Validity was supported by significant correlations with engagement determinants (e.g., motivation, r = .413), processes (e.g., between-session efforts, r = .390), and outcomes (e.g., symptoms, r = − .304). Results remained robust across bootstrap resampling and cross-validation, accounting for nested data. The LLM rating scale exhibited strong psychometric properties, demonstrating the potential of the approach as an assessment tool. Importantly, this automated approach uses interpretable items, ensuring clear understanding of measured constructs, while supporting local implementation and protecting confidential data.

Article Details

Volume / Issue Vol. 15, Issue 1
Published August 12, 2025
ISSN 2045-2322
Publisher Nature Portfolio

Journal Info

Scientific Reports

Nature Portfolio

ISSN: 2045-2322 Open Access Life Sciences

Authors (9)

S

Steffen T. Eberhardt

A

Antonia Vehlen

J

Jana Schaffrath

B

Brian Schwartz

T

Tobias Baur

D

Dominik Schiller

T

Tobias Hallmen

E

Elisabeth André

University of Augsburg

W

Wolfgang Lutz

Asian Demographic Research Institute, Shanghai University