<p>The mainstream approaches for environmental sound classification (ESC) typically employ audio spectrograms as input and leverage techniques from image processing. However, a key distinction exists between audio spectrograms and images: the former are generally rectangular, whereas the latter are commonly square or nearly square. This rectangular form stems from the inherent structure of audio signals, which span both time and frequency domains. Notably, the frequency dimension of spectrograms is often more compact and information-rich than the time dimension. In light of this, we propose to replace the conventional square convolution kernel–which treats both dimensions equally–with rectangular kernels combined with dilated convolutions. This design prioritises the more informative frequency axis. We further introduce a neural network model tailored for ESC, enhanced by self-distilled soft labels to enrich the input information, and a reconstructed loss function that boosts both accuracy and robustness. Experimental results on UrbanSound8K, ESC-10, and ESC-50 datasets achieve accuracies of 98.62%, 95.50%, and 89.3%, respectively, matching or surpassing state-of-the-art performance while maintaining a lower parameter count–demonstrating the efficiency of our approach.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Rectangular kernels for information-dense domains in environmental sound classification

  • Zhenghao Chang,
  • Ruhan He,
  • Yongsheng Yu

摘要

The mainstream approaches for environmental sound classification (ESC) typically employ audio spectrograms as input and leverage techniques from image processing. However, a key distinction exists between audio spectrograms and images: the former are generally rectangular, whereas the latter are commonly square or nearly square. This rectangular form stems from the inherent structure of audio signals, which span both time and frequency domains. Notably, the frequency dimension of spectrograms is often more compact and information-rich than the time dimension. In light of this, we propose to replace the conventional square convolution kernel–which treats both dimensions equally–with rectangular kernels combined with dilated convolutions. This design prioritises the more informative frequency axis. We further introduce a neural network model tailored for ESC, enhanced by self-distilled soft labels to enrich the input information, and a reconstructed loss function that boosts both accuracy and robustness. Experimental results on UrbanSound8K, ESC-10, and ESC-50 datasets achieve accuracies of 98.62%, 95.50%, and 89.3%, respectively, matching or surpassing state-of-the-art performance while maintaining a lower parameter count–demonstrating the efficiency of our approach.