Hierarchical attentive transformer with label guided fusion for multimodal movie scene segmentation

W Weiwei Zhao J Jing Wang (Hunan Cancer Hospital Changsha China)

Abstract

Abstract Existing movie scene segmentation methods struggle with long-range temporal dependencies, attention degradation, and rigid multimodal fusion, often ignoring the inherent hierarchical structure of movies. To address these limitations, this paper proposes a Hierarchical Attentive Transformer (HATrans). First, to capture multi-granularity semantic consistency and discriminative scene boundaries, we adopt a shot-to-segment hierarchical encoding pipeline that explicitly models structural priors. Second, to mitigate attention degradation when processing extremely long shot sequences, we introduce a hierarchical masked attention mechanism combined with a temporal position-aware bias, which restricts irrelevant connections and enhances structural sensitivity. Third, to overcome the inflexibility of conventional multimodal integration, we propose a label-guided attention fusion module that leverages semantic category priors to dynamically weight visual, audio, and subtitle features based on varying semantic contexts. Experimental results on MovieNet-42 K suggest that HATrans achieves competitive performance, outperforming several baselines including CMTS.

Article Details

Volume / Issue Vol. 16, Issue 1
Published May 21, 2026
ISSN 2045-2322
Publisher Nature Portfolio

Journal Info

Scientific Reports

Nature Portfolio

ISSN: 2045-2322 Open Access Life Sciences

Authors (2)

W

Weiwei Zhao

J

Jing Wang

Hunan Cancer Hospital Changsha China