<p>The temporal discontinuity and duration inconsistency of polyphonic sound events significantly complicate detection tasks. To address these limitations, we propose a novel detection framework that integrates cross-pooling depthwise separable convolution with multi-scale dilated convolution (CPDSC-MSDCN). The method first employs the Mel-Pseudo Constant Q-Transform (MPCQT) for feature extraction, which enhances spectral resolution and enables precise capture of subtle acoustic features in polyphonic signals. The resulting feature maps are subsequently fed into a CPDSC module. Within this module, the Depthwise Separable Convolution (DSC) component effectively reduces the number of model parameters, enhancing computational efficiency. Additionally, the introduction of cross-pooling within the CPDSC module serves to balance feature retention and compression. Finally, the MSDCN is used to model the extracted local features, capturing long short-term dependencies of sound events by fusing features with different lengths of temporal dependencies. Evaluations demonstrate a 22.3% F1-score improvement and a 32.9% reduction in error rate over the DCASE 2017 baseline, outperforming state-of-the-art models including MS-FCN and MSSF-Net.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimized Network Combining Depthwise Separable Convolutional with Multi-scale Dilated Convolution Used for Polyphonic Sound Event Detection

  • Liyan Luo,
  • Fukun Mi,
  • Mei Wang,
  • Zhenghong Liu,
  • Yun Li

摘要

The temporal discontinuity and duration inconsistency of polyphonic sound events significantly complicate detection tasks. To address these limitations, we propose a novel detection framework that integrates cross-pooling depthwise separable convolution with multi-scale dilated convolution (CPDSC-MSDCN). The method first employs the Mel-Pseudo Constant Q-Transform (MPCQT) for feature extraction, which enhances spectral resolution and enables precise capture of subtle acoustic features in polyphonic signals. The resulting feature maps are subsequently fed into a CPDSC module. Within this module, the Depthwise Separable Convolution (DSC) component effectively reduces the number of model parameters, enhancing computational efficiency. Additionally, the introduction of cross-pooling within the CPDSC module serves to balance feature retention and compression. Finally, the MSDCN is used to model the extracted local features, capturing long short-term dependencies of sound events by fusing features with different lengths of temporal dependencies. Evaluations demonstrate a 22.3% F1-score improvement and a 32.9% reduction in error rate over the DCASE 2017 baseline, outperforming state-of-the-art models including MS-FCN and MSSF-Net.