Recent years have witnessed remarkable progress in self-supervised pretraining of Vision Transformers. However, all state-of-the-art methods utilize solely image data for pretraining, missing out on the additional information of motion and sound that videos can offer. In this paper, we present a novel approach to enhancing the dense performance of Vision Transformers (ViT) by leveraging joint audio-visual information from web-scraped video datasets. We begin with image encoders pretrained on image data and employ a dense cross-modal contrastive loss function, which helps establish meaningful associations between visual components and their corresponding auditory cues. Furthermore, we explore a novel unsupervised and few-shot segmentation approach by creating audio prototypes from an audio database, and then mapping each image patch to the most similar audio prototype, yielding segmentation masks. We validate our approach on three diverse and challenging datasets: Pascal VOC 2012, COCO-Things, and COCO-Stuff, and demonstrate that audio-visual pre-training can significantly boost performance over models solely pretrained on image data, establishing new state-of-the-art records in certain scenarios, while also enabling the use of audio prototypes to perform semantic segmentation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Self-supervised Semantic Segmentation from Audio-Visual Data

  • Apostolos Panagiotopoulos,
  • Mandy Toh,
  • Triantafyllos Afouras,
  • Yuki M. Asano

摘要

Recent years have witnessed remarkable progress in self-supervised pretraining of Vision Transformers. However, all state-of-the-art methods utilize solely image data for pretraining, missing out on the additional information of motion and sound that videos can offer. In this paper, we present a novel approach to enhancing the dense performance of Vision Transformers (ViT) by leveraging joint audio-visual information from web-scraped video datasets. We begin with image encoders pretrained on image data and employ a dense cross-modal contrastive loss function, which helps establish meaningful associations between visual components and their corresponding auditory cues. Furthermore, we explore a novel unsupervised and few-shot segmentation approach by creating audio prototypes from an audio database, and then mapping each image patch to the most similar audio prototype, yielding segmentation masks. We validate our approach on three diverse and challenging datasets: Pascal VOC 2012, COCO-Things, and COCO-Stuff, and demonstrate that audio-visual pre-training can significantly boost performance over models solely pretrained on image data, establishing new state-of-the-art records in certain scenarios, while also enabling the use of audio prototypes to perform semantic segmentation.