Predicting dominant terrestrial biomes at a global scale using machine learning algorithms, climate variable indices, and extreme event indices

H Hisashi Sato

Abstract

Understanding the global distribution of biomes is essential for biodiversity conservation, climate modeling, and land-use planning. Traditional approaches often summarize climate data into indices, and recent models sometimes include extreme events such as severe droughts or rare cold spells. This study evaluates how the choice of machine learning algorithm, climate data summarization, and extreme climate indices affect the accuracy and robustness of global biome modeling. Four algorithms were tested: random forest (RF), support vector machine (SVM), naive Bayes (NV), and LeNet convolutional neural network (CNN). RF and CNN achieved the highest accuracy, with CNN preferred due to RF’s stronger overfitting. Summarizing climate data into indices reduced accuracy by 1–2%, while adding extreme indices increased accuracy by <2% (except for NV, which performed poorly overall). However, extreme climate data caused large mismatches between observed and predicted climate values, reducing robustness as measured by prediction consistency. These results indicate that including extreme climate data in global biome prediction models offers limited accuracy gains but can significantly weaken robustness, so caution is advised.

Article Details

Journal PLoS ONE
Volume / Issue Vol. 21, Issue 2
Published February 26, 2026
Pages e0324107
ISSN 1932-6203
Publisher Public Library of Science

Journal Info

PLoS ONE

Public Library of Science

ISSN: 1932-6203 Open Access Health Sciences

Authors (1)

H

Hisashi Sato