Revisiting SSL for sound event detection: complementary fusion and adaptive post-processing
摘要
Self-supervised learning (SSL) models offer powerful representations for sound event detection (SED), yet their synergistic potential remains underexplored. This study systematically evaluates state-of-the-art SSL models to guide optimal model selection and integration for SED. We propose a framework combining heterogeneous SSL model (BEATs, HuBERT, WavLM, etc.) representations through three fusion strategies: individual SSL embedding integration, dual-modal fusion, and full aggregation. Our experiments on the DCASE 2023 Task 4 Challenge reveal that dual-modal fusion (e.g., CRNN+BEATs+WavLM) achieves complementary performance gains, while CRNN+BEATs alone attains superior individual SSL embedding integration results. We further introduce a normalized sound event bounding boxes (nSEBBs), an adaptive post-processing method that dynamically adjusts target events detection boundaries, improving PSDS