A multistream attention based neural network for visual speech recognition and sign language understanding

F Fatma M. Talaat B Basma M. Hassan

Abstract

Abstract This paper introduces SignKeyNet, a novel multi-stream keypoint-based neural network designed to enhance lip-reading recognition and support foundational sign language understanding. The architecture decouples signer movements into three primary streams hands, face, and body using 133 pose keypoints extracted via pose estimation techniques. Each stream is processed independently using specialized attention modules, followed by an attention-based fusion mechanism that models cross-modal spatiotemporal dependencies. SignKeyNet is evaluated on the MIRACL-VC1 lip-reading dataset, achieving superior performance over baseline models such as HMMs, DTW, CNNs, LSTMs, and Two-stream ConvNets, with results including an accuracy of 0.85, a Word Error Rate (WER) of 0.12, and a Character Error Rate (CER) of 0.06. These results highlight the effectiveness of attention-driven, multi-modal architectures for visual speech recognition tasks. While the current evaluation focuses on lip-reading due to dataset constraints, the proposed architecture is extendable to full Sign Language Translation (SLT) systems. SignKeyNet demonstrates strong potential for real-time deployment in accessibility technologies, particularly for the deaf and hard-of-hearing communities. 

Article Details

Volume / Issue Vol. 15, Issue 1
Published December 24, 2025
ISSN 2045-2322
Publisher Nature Portfolio

Journal Info

Scientific Reports

Nature Portfolio

ISSN: 2045-2322 Open Access Life Sciences

Authors (2)

F

Fatma M. Talaat

B

Basma M. Hassan