A lightweight dual branch masking network for environmental sound classification
摘要
Environmental sound classification (ESC) is crucial for applications such as intelligent surveillance, urban acoustic monitoring, and human-computer interaction. Although deep neural networks (DNNs) have significantly improved ESC performance, these methods often rely on large models and extensive pretraining, making them difficult to deploy in resource-constrained environments. Some existing lightweight models, while having fewer parameters, still suffer from limited representational capacity, leading to suboptimal generalization, especially in low-data scenarios. To address these challenges, we propose SpectroMaskNet, a compact dual-branch architecture. This design integrates global-local attention mechanisms with block-masked spectrogram augmentation, allowing the model to capture both long-term temporal dependencies and fine-grained spectral features. This enhances robustness and generalization, particularly in data-scarce situations. Experimental results on four benchmark datasets–ESC-10, ESC-50, UrbanSound8K, and SpeechCommandV2–demonstrate that SpectroMaskNet achieves accuracies of 97.50%, 95.50%, 96.32%, and 96.52%, respectively, outperforming existing lightweight baselines without requiring large-scale pretraining. Furthermore, the model maintains low computational complexity, making it well-suited for real-world ESC applications that demand efficiency and scalability.