Audio-visual segmentation is a recently proposed task, whose main goal is to locate the target of the sound in the image at the pixel level. In practical scenarios, multiple types of audio can coexist, with different objects emitting sounds at different frequencies. However, existing methods only use single-frequency audio information when fusing audio and visual modalities. Moreover, the process of combining images and audio can be quite rough. Therefore, we propose a multi-frequency fine-grained matching method for multiple sound sources scenario. Firstly, we use short-time Fourier transform (STFT) to extract different frequency spectrograms and input them into the audio encoder to extract multi-frequency audio features. Secondly, multi-frequency audio information serves as a prompt in the pixel decoder stage to guide model segmentation. To obtain high-quality prompts, we use an attention method in the Audio-Visual Matching Module (AVMM) to match visual and audio information. The experiments show that our method has a significant improvement over the baseline and achieves state-of-the-art results on the MS3 benchmark (64.1 mIoU on MS3).

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-frequency Fine-Grained Matching for Audio-Visual Segmentation

  • Yinhao Zhang,
  • Tianyang Xu,
  • Xiao-Jun Wu,
  • Shaochuan Zhao,
  • Josef Kittler

摘要

Audio-visual segmentation is a recently proposed task, whose main goal is to locate the target of the sound in the image at the pixel level. In practical scenarios, multiple types of audio can coexist, with different objects emitting sounds at different frequencies. However, existing methods only use single-frequency audio information when fusing audio and visual modalities. Moreover, the process of combining images and audio can be quite rough. Therefore, we propose a multi-frequency fine-grained matching method for multiple sound sources scenario. Firstly, we use short-time Fourier transform (STFT) to extract different frequency spectrograms and input them into the audio encoder to extract multi-frequency audio features. Secondly, multi-frequency audio information serves as a prompt in the pixel decoder stage to guide model segmentation. To obtain high-quality prompts, we use an attention method in the Audio-Visual Matching Module (AVMM) to match visual and audio information. The experiments show that our method has a significant improvement over the baseline and achieves state-of-the-art results on the MS3 benchmark (64.1 mIoU on MS3).