<p>Self-supervised learning (SSL) models offer powerful representations for sound event detection (SED), yet their synergistic potential remains underexplored. This study systematically evaluates state-of-the-art SSL models to guide optimal model selection and integration for SED. We propose a framework combining heterogeneous SSL model (BEATs, HuBERT, WavLM, etc.) representations through three fusion strategies: individual SSL embedding integration, dual-modal fusion, and full aggregation. Our experiments on the DCASE 2023 Task 4 Challenge reveal that dual-modal fusion (e.g., CRNN+BEATs+WavLM) achieves complementary performance gains, while CRNN+BEATs alone attains superior individual SSL embedding integration results. We further introduce a normalized sound event bounding boxes (nSEBBs), an adaptive post-processing method that dynamically adjusts target events detection boundaries, improving PSDS<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="44443_2025_181_Article_IEq1.gif" Format="GIF" Height="10" Rendition="HTML" Resolution="72" Type="Linedraw" Width="8" /> </InlineMediaObject> <EquationSource Format="TEX">\(_1\)</EquationSource> <EquationSource Format="MATHML"><math> <mmultiscripts> <mrow /> <mn>1</mn> <mrow /> </mmultiscripts> </math></EquationSource> </InlineEquation> by up to 4% for standalone SSL models. These findings provide important insights into SSL model compatibility, demonstrating that task-specific fusion and dynamic post-processing enhance robustness. Our work establishes a reference for selecting and integrating SSL architectures in SED systems, balancing efficiency and accuracy for real-world deployment.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Revisiting SSL for sound event detection: complementary fusion and adaptive post-processing

  • Hanfang Cui,
  • Longfei Song,
  • Li Li,
  • Dongxing Xu,
  • Yanhua Long

摘要

Self-supervised learning (SSL) models offer powerful representations for sound event detection (SED), yet their synergistic potential remains underexplored. This study systematically evaluates state-of-the-art SSL models to guide optimal model selection and integration for SED. We propose a framework combining heterogeneous SSL model (BEATs, HuBERT, WavLM, etc.) representations through three fusion strategies: individual SSL embedding integration, dual-modal fusion, and full aggregation. Our experiments on the DCASE 2023 Task 4 Challenge reveal that dual-modal fusion (e.g., CRNN+BEATs+WavLM) achieves complementary performance gains, while CRNN+BEATs alone attains superior individual SSL embedding integration results. We further introduce a normalized sound event bounding boxes (nSEBBs), an adaptive post-processing method that dynamically adjusts target events detection boundaries, improving PSDS \(_1\) 1 by up to 4% for standalone SSL models. These findings provide important insights into SSL model compatibility, demonstrating that task-specific fusion and dynamic post-processing enhance robustness. Our work establishes a reference for selecting and integrating SSL architectures in SED systems, balancing efficiency and accuracy for real-world deployment.