A lightweight dual branch masking network for environmental sound classification
Abstract
Abstract Environmental sound classification (ESC) is crucial for applications such as intelligent surveillance, urban acoustic monitoring, and human-computer interaction. Although deep neural networks (DNNs) have significantly improved ESC performance, these methods often rely on large models and extensive pretraining, making them difficult to deploy in resource-constrained environments. Some existing lightweight models, while having fewer parameters, still suffer from limited representational capacity, leading to suboptimal generalization, especially in low-data scenarios. To address these challenges, we propose SpectroMaskNet, a compact dual-branch architecture. This design integrates global-local attention mechanisms with block-masked spectrogram augmentation, allowing the model to capture both long-term temporal dependencies and fine-grained spectral features. This enhances robustness and generalization, particularly in data-scarce situations. Experimental results on four benchmark datasets–ESC-10, ESC-50, UrbanSound8K, and SpeechCommandV2–demonstrate that SpectroMaskNet achieves accuracies of 97.50%, 95.50%, 96.32%, and 96.52%, respectively, outperforming existing lightweight baselines without requiring large-scale pretraining. Furthermore, the model maintains low computational complexity, making it well-suited for real-world ESC applications that demand efficiency and scalability.
Article Details
Authors (10)
Guorong Chen
Bao Zhang
School of Chemical Engineering and Technology
Zhikang Ding
Ke Xiao
Pengyu Guan
Xianghan Xiao
Xiaoqiang Wang
Haixin Yi
Hong Hu
Weijie Zhang