MAF-Net: A multimodal data fusion approach for human action recognition

D Dongwei Xie X Xiaodan Zhang (Institute of Photoelectronic Thin Film Devices and Technology, Renewable Energy Conversion and Storage Center, State Key Laboratory of Photovoltaic Materials and Cells) X Xiang Gao H Hu Zhao D Dongyang Du

Abstract

3D skeleton-based human activity recognition has gained significant attention due to its robustness against variations in background, lighting, and viewpoints. However, challenges remain in effectively capturing spatiotemporal dynamics and integrating complementary information from multiple data modalities, such as RGB video and skeletal data. To address these challenges, we propose a multimodal fusion framework that leverages optical flow-based key frame extraction, data augmentation techniques, and an innovative fusion of skeletal and RGB streams using self-attention and skeletal attention modules. The model employs a late fusion strategy to combine skeletal and RGB features, allowing for more effective capture of spatial and temporal dependencies. Extensive experiments on benchmark datasets, including NTU RGB+D, SYSU, and UTD-MHAD, demonstrate that our method outperforms existing models. This work not only enhances action recognition accuracy but also provides a robust foundation for future multimodal integration and real-time applications in diverse fields such as surveillance and healthcare.

Article Details

Journal PLoS ONE
Volume / Issue Vol. 20, Issue 4
Published April 09, 2025
Pages e0319656
ISSN 1932-6203
Publisher Public Library of Science

Journal Info

PLoS ONE

Public Library of Science

ISSN: 1932-6203 Open Access Health Sciences

Authors (5)

D

Dongwei Xie

X

Xiaodan Zhang

Institute of Photoelectronic Thin Film Devices and Technology, Renewable Energy Conversion and Storage Center, State Key Laboratory of Photovoltaic Materials and Cells

X

Xiang Gao

H

Hu Zhao

D

Dongyang Du