SkelFormer: An adaptive hierarchical transformer-based approach on skeleton graphs for human action recognition in video sequences
Abstract
Human skeleton-based action recognition represents a pivotal field of study, capturing the intricate interplay between physical dynamics and intentional actions. Current research primarily focuses on extracting structural and temporal information from static skeleton-based graphs, but it grapples with a myriad of challenges. These include 1) an absence of hierarchical structure in encoding the skeleton-based graphs. 2) A requirement for substantial prior knowledge to interpret the diverse spatial dynamics within singular action labels. 3) An intricate task of representing the multifaceted temporal dynamics of individual actions. To address these challenges, we propose SkelFormer, a novel framework that captures spatiotemporal variations in skeleton-based graphs extracted from video sequences. The proposed SkelFormer incorporates the SKT Block as a central element, effectively facilitating information exchange through node concentration and diffusion across both structural and temporal dimensions. This design enables the extraction of hierarchical representations without relying on handcrafted rules, thereby improving the understanding of complex action patterns. Our rigorous experimental evaluations further substantiate SkelFormer’s supremacy, outperforming several state-of-the-art benchmarks in skeleton-based action recognition and achieving accuracy rates of 92.8% on the NTU RGB+D 60, 89.4% on the NTU RGB+D 120 (cross-subject split), and 96.1% on the NW-UCLA dataset.
Article Details
Authors (4)
Jiexing Yan
Xi Zhang
Caiyan Tan
Dawen Li