A two-stage deep learning approach for facial emotion recognition in RAVDESS videos

A Achraf Jallaglag M My Abdelouahed Sabri A Ali Yahyaouy A Abdellah Aarab

Abstract

Abstract Video-based emotion recognition is an important topic in affective computing, with applications in human–computer interaction, mental health, and multimedia systems. In this work, we propose a two-stage deep learning approach that extracts spatial features from individual frames and leverages temporal consistency across video sequences for facial expression recognition using the RAVDESS dataset. First, videos are split into frames, and a fine-tuned VGG16 CNN extracts discriminative spatial features from each frame. Second, these features are aggregated into sequences and the frame-level predictions are aggregated at the video level using a majority voting strategy to ensure temporal consistency. Our experiments show that the proposed method achieves 93.6% accuracy at the frame level and 98.1% at the video level, outperforming baseline models and remaining competitive with recent state-of-the-art approaches. Temporal aggregation helps reduce misclassifications of subtle emotions, while fine-tuning improves feature extraction. The approach is computationally efficient and provides a solid foundation for future research in multimodal emotion recognition and advanced video-level aggregation.

Article Details

Volume / Issue Vol. 1, Issue 1
Published June 29, 2026
ISSN 2045-2322
Publisher Nature Portfolio

Journal Info

Scientific Reports

Nature Portfolio

ISSN: 2045-2322 Open Access Life Sciences

Authors (4)

A

Achraf Jallaglag

M

My Abdelouahed Sabri

A

Ali Yahyaouy

A

Abdellah Aarab