<p>Sound Event Detection (SED) is a pivotal task in audio signal processing with widespread applications, requiring the classification and temporal localization of sound events. However, there proves to be a challenge in balancing global features for event classification with local features for temporal localization. This paper introduces Group Feature Calibration (GFC), a plug-and-play solution designed to effectively representing diverse time-frequency characteristics of sound events. GFC consists of two sub-modules: Group Feature Learning (GFL) and Task-aware Activation (TA). GFL enhances the network’s ability to capture and integrate both global and local sound event features, while TA adaptively refines feature activation, facilitating more precise event classification and time localization. This approach effectively recalibrates the features, significantly improving the performance of various SED models on both two sub-tasks. Experimental results demonstrate that implementing GFC module in three mainstream SED networks improves the Polyphonic Sound Event Detection Score (PSDS) on the DCASE 2022 Task 4 dataset by up to 9.94%. Through class-wise analysis of the results, the consistent improvement of the GFC module in detecting sound events with diverse time-frequency characteristics is proven. Further ablation study also validates the effectiveness of the two sub-modules within the GFC module.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Group feature calibration for sound event detection

  • Yanzhen Ren,
  • Wuyang Liu,
  • Chenyu Liu,
  • Tingting Zhu

摘要

Sound Event Detection (SED) is a pivotal task in audio signal processing with widespread applications, requiring the classification and temporal localization of sound events. However, there proves to be a challenge in balancing global features for event classification with local features for temporal localization. This paper introduces Group Feature Calibration (GFC), a plug-and-play solution designed to effectively representing diverse time-frequency characteristics of sound events. GFC consists of two sub-modules: Group Feature Learning (GFL) and Task-aware Activation (TA). GFL enhances the network’s ability to capture and integrate both global and local sound event features, while TA adaptively refines feature activation, facilitating more precise event classification and time localization. This approach effectively recalibrates the features, significantly improving the performance of various SED models on both two sub-tasks. Experimental results demonstrate that implementing GFC module in three mainstream SED networks improves the Polyphonic Sound Event Detection Score (PSDS) on the DCASE 2022 Task 4 dataset by up to 9.94%. Through class-wise analysis of the results, the consistent improvement of the GFC module in detecting sound events with diverse time-frequency characteristics is proven. Further ablation study also validates the effectiveness of the two sub-modules within the GFC module.