Data-driven prediction of ionic conductivity in solid-state electrolytes with machine learning and large language models

H Haewon Kim T Taekgi Lee (School of Chemical Engineering, Pusan National University 1 , Busan 46241,) S Seongeun Hong (School of Chemical Engineering, Pusan National University 1 , Busan 46241,) K Kyeong-Ho Kim (Department of Materials Science and Engineering, Pukyong National University 2 , Busan 48513,) Y Yongchul G. Chung

Abstract

Solid-state electrolytes (SSEs) are attractive for next-generation lithium-ion batteries due to improved safety and stability, but their low room-temperature ionic conductivity hinders practical application. Experimental synthesis and testing of new SSEs remain time-consuming and resource-intensive. Machine learning offers an accelerated route for SSE discovery; however, composition-only models neglect structural factors important for ion transport, while graph neural networks are challenged by the scarcity of structure-labeled conductivity data and the prevalence of crystallographic disorder in crystal structures (CIFs). Here, we train two complementary predictors on the same room-temperature, structure-labeled dataset (n = 499). A gradient-boosted tree regressor model using stoichiometric descriptors alone achieves a test MAE of 1.108 in log(S/cm); adding geometric descriptors (combined MAE = 1.172) does not lower the test error but reveals complementary structural information through Shapley Additive exPlanations, which shows that stoichiometric descriptors, particularly the oxygen ratio, dominate feature importance (seven of the top ten features), with three geometric descriptors (density, Lmax, and Lmin) also contributing meaningfully. In parallel, we fine-tune large language models (LLMs) using compact text prompts derived from CIF metadata (formula with optional symmetry and disorder tags), avoiding direct use of raw atomic coordinates. Notably, while Mistral-7B achieves the lowest absolute error [MAE = 0.798 in log(S/cm)], Qwen3-8B demonstrates the best overall ranking performance (SRCC = 0.849) using formula and disorder information, eliminating the need for numerical feature extraction from CIF files. Together, these results show that global geometric descriptors improve tree-based predictions and enable interpretable structure–property analysis, while LLMs provide a competitive low-preprocessing alternative for rapid SSE screening.

Article Details

Volume / Issue Vol. 164, Issue 11
Published March 21, 2026
ISSN 0021-9606
Publisher American Institute of Physics

Journal Info

The Journal of Chemical Physics

American Institute of Physics

ISSN: 0021-9606 Physical Sciences

Authors (5)

H

Haewon Kim

T

Taekgi Lee

School of Chemical Engineering, Pusan National University 1 , Busan 46241,

S

Seongeun Hong

School of Chemical Engineering, Pusan National University 1 , Busan 46241,

K

Kyeong-Ho Kim

Department of Materials Science and Engineering, Pukyong National University 2 , Busan 48513,

Y

Yongchul G. Chung