Recent efforts have focused on training audio-visual pairs through self-supervised contrastive learning, which relies on the assumption of audio-visual correspondence (AVC). This assumption posits that positive pairs consist of audio and visual from the same video, while negative pairs are formed from different videos. However, this assumption is too strict and may be unreliable in practice. This unreliable assumption inevitably introduces two types of noisy correspondence. False positive pairs arise from weak AVC caused by invisible-sounding objects or background noise. Conversely, false negative pairs arise from strong AVC caused by random pairing. In this paper, we focus on the visual sound localization task, aiming to localize the visual regions that emit sound. To address the issue of noisy correspondence in visual sound localization, an optimized soft contrastive loss is proposed to alleviate the impact of false positives. Additionally, the hard contrastive set mixing strategy is utilized to suppress the effect of false negatives. Experimental results demonstrate that our methods significantly reduce the impact of noisy correspondences and achieve competitive results on standard benchmarks. Furthermore, the proposed method shows potential for generalization to other multi-modal tasks based on contrastive learning.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Robust Contrastive Learning Against Audio-Visual Noisy Correspondence

  • Yihan Zhao,
  • Wei Xi,
  • Gairui Bai,
  • Xinhui Liu,
  • Jizhong Zhao

摘要

Recent efforts have focused on training audio-visual pairs through self-supervised contrastive learning, which relies on the assumption of audio-visual correspondence (AVC). This assumption posits that positive pairs consist of audio and visual from the same video, while negative pairs are formed from different videos. However, this assumption is too strict and may be unreliable in practice. This unreliable assumption inevitably introduces two types of noisy correspondence. False positive pairs arise from weak AVC caused by invisible-sounding objects or background noise. Conversely, false negative pairs arise from strong AVC caused by random pairing. In this paper, we focus on the visual sound localization task, aiming to localize the visual regions that emit sound. To address the issue of noisy correspondence in visual sound localization, an optimized soft contrastive loss is proposed to alleviate the impact of false positives. Additionally, the hard contrastive set mixing strategy is utilized to suppress the effect of false negatives. Experimental results demonstrate that our methods significantly reduce the impact of noisy correspondences and achieve competitive results on standard benchmarks. Furthermore, the proposed method shows potential for generalization to other multi-modal tasks based on contrastive learning.