CNN-transformer-based model explained by SHAP and multi-head attention weights for time series forecasting

S Stefano Frizzo Stefenon J João Pedro Matos-Carvalho V Valderi Reis Quietinho Leithardt K Kin-Choong Yow

Abstract

Abstract Convolutional Neural Networks (CNNs) and transformer architectures offer strengths for modeling temporal data: CNNs excel at capturing local patterns and translational invariances, while transformers effectively model long-range dependencies via self-attention. This paper proposes a hybrid architecture integrating convolutional feature extraction with a multi-head attention backbone to enhance multivariate time series forecasting. The CNN module first applies a hierarchy of one-dimensional convolutional layers to distill salient local patterns from raw input sequences, reducing noise and dimensionality. The resulting feature maps are then fed into the prediction model, which applies multi-head attention to capture both short- and long-term dependencies and to weigh relevant covariates adaptively. We evaluate the proposed CNN-transformer-based model on a hydroelectric natural flow time series dataset. Experimental results demonstrate that the proposed model outperforms well-established deep learning models, with a mean absolute percentage error of up to 2.2%. The explainability of the model is obtained by a proposed SHapley Additive exPlanations (SHAP) and Multi-Head Attention Weights (MHAW). Our novel architecture, named CNN-Transformer-SHAP-MHAW, is promising for applications requiring high-fidelity and multivariate time series forecasts.

Article Details

Volume / Issue Vol. 1, Issue 1
Published July 25, 2026
ISSN 2045-2322
Publisher Nature Portfolio

Journal Info

Scientific Reports

Nature Portfolio

ISSN: 2045-2322 Open Access Life Sciences

Authors (4)

S

Stefano Frizzo Stefenon

J

João Pedro Matos-Carvalho

V

Valderi Reis Quietinho Leithardt

K

Kin-Choong Yow