SkelFormer: An adaptive hierarchical transformer-based approach on skeleton graphs for human action recognition in video sequences

J Jiexing Yan X Xi Zhang C Caiyan Tan D Dawen Li

Abstract

Human skeleton-based action recognition represents a pivotal field of study, capturing the intricate interplay between physical dynamics and intentional actions. Current research primarily focuses on extracting structural and temporal information from static skeleton-based graphs, but it grapples with a myriad of challenges. These include 1) an absence of hierarchical structure in encoding the skeleton-based graphs. 2) A requirement for substantial prior knowledge to interpret the diverse spatial dynamics within singular action labels. 3) An intricate task of representing the multifaceted temporal dynamics of individual actions. To address these challenges, we propose SkelFormer, a novel framework that captures spatiotemporal variations in skeleton-based graphs extracted from video sequences. The proposed SkelFormer incorporates the SKT Block as a central element, effectively facilitating information exchange through node concentration and diffusion across both structural and temporal dimensions. This design enables the extraction of hierarchical representations without relying on handcrafted rules, thereby improving the understanding of complex action patterns. Our rigorous experimental evaluations further substantiate SkelFormer’s supremacy, outperforming several state-of-the-art benchmarks in skeleton-based action recognition and achieving accuracy rates of 92.8% on the NTU RGB+D 60, 89.4% on the NTU RGB+D 120 (cross-subject split), and 96.1% on the NW-UCLA dataset.

Article Details

Journal PLoS ONE
Volume / Issue Vol. 21, Issue 1
Published January 12, 2026
Pages e0340390
ISSN 1932-6203
Publisher Public Library of Science

Journal Info

PLoS ONE

Public Library of Science

ISSN: 1932-6203 Open Access Health Sciences

Authors (4)

J

Jiexing Yan

X

Xi Zhang

C

Caiyan Tan

D

Dawen Li