<p>Environmental sound classification (ESC) is crucial for applications such as intelligent surveillance, urban acoustic monitoring, and human-computer interaction. Although deep neural networks (DNNs) have significantly improved ESC performance, these methods often rely on large models and extensive pretraining, making them difficult to deploy in resource-constrained environments. Some existing lightweight models, while having fewer parameters, still suffer from limited representational capacity, leading to suboptimal generalization, especially in low-data scenarios. To address these challenges, we propose SpectroMaskNet, a compact dual-branch architecture. This design integrates global-local attention mechanisms with block-masked spectrogram augmentation, allowing the model to capture both long-term temporal dependencies and fine-grained spectral features. This enhances robustness and generalization, particularly in data-scarce situations. Experimental results on four benchmark datasets–ESC-10, ESC-50, UrbanSound8K, and SpeechCommandV2–demonstrate that SpectroMaskNet achieves accuracies of 97.50%, 95.50%, 96.32%, and 96.52%, respectively, outperforming existing lightweight baselines without requiring large-scale pretraining. Furthermore, the model maintains low computational complexity, making it well-suited for real-world ESC applications that demand efficiency and scalability.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A lightweight dual branch masking network for environmental sound classification

  • Guorong Chen,
  • Bao Zhang,
  • Zhikang Ding,
  • Ke Xiao,
  • Pengyu Guan,
  • Xianghan Xiao,
  • Xiaoqiang Wang,
  • Haixin Yi,
  • Hong Hu,
  • Weijie Zhang

摘要

Environmental sound classification (ESC) is crucial for applications such as intelligent surveillance, urban acoustic monitoring, and human-computer interaction. Although deep neural networks (DNNs) have significantly improved ESC performance, these methods often rely on large models and extensive pretraining, making them difficult to deploy in resource-constrained environments. Some existing lightweight models, while having fewer parameters, still suffer from limited representational capacity, leading to suboptimal generalization, especially in low-data scenarios. To address these challenges, we propose SpectroMaskNet, a compact dual-branch architecture. This design integrates global-local attention mechanisms with block-masked spectrogram augmentation, allowing the model to capture both long-term temporal dependencies and fine-grained spectral features. This enhances robustness and generalization, particularly in data-scarce situations. Experimental results on four benchmark datasets–ESC-10, ESC-50, UrbanSound8K, and SpeechCommandV2–demonstrate that SpectroMaskNet achieves accuracies of 97.50%, 95.50%, 96.32%, and 96.52%, respectively, outperforming existing lightweight baselines without requiring large-scale pretraining. Furthermore, the model maintains low computational complexity, making it well-suited for real-world ESC applications that demand efficiency and scalability.