Viewport prediction with cross modal multiscale transformer for 360° video streaming

Y Yangsheng Tian Y Yi Zhong Y Yi Han F Fangyuan Chen (School of Materials Science and Engineering, National Institute of New Materials Research)

Abstract

Abstract In the realm of immersive video technologies, efficient 360° video streaming remains a challenge due to the high bandwidth requirements and the dynamic nature of user viewports. Most existing approaches neglect the dependencies between different modalities, and personal preferences are rarely considered. These limitations lead to inconsistent prediction performance. Here, we present a novel viewport prediction model leveraging a Cross Modal Multiscale Transformer (CMMST) that integrates user trajectory and video saliency features across different scales. Our approach outperforms baseline methods, maintaining high precision even with extended prediction intervals. By harnessing the Cross Modal attention mechanisms, CMMST captures intricate user preferences and viewing patterns, offering a promising solution for adaptive streaming in virtual reality and other immersive platforms. The code of this work is available at https://github.com/bbgua85776540/CMMST.

Article Details

Volume / Issue Vol. 15, Issue 1
Published August 19, 2025
ISSN 2045-2322
Publisher Nature Portfolio

Journal Info

Scientific Reports

Nature Portfolio

ISSN: 2045-2322 Open Access Life Sciences

Authors (4)

Y

Yangsheng Tian

Y

Yi Zhong

Y

Yi Han

F

Fangyuan Chen

School of Materials Science and Engineering, National Institute of New Materials Research