Learning Supportive Two-Stream Network for Audio-Visual Segmentation
摘要
Audio Visual Segmentation is an emerging problem in multi-modality video analysis. Sound signals in the task enable to help the network segment the vocal object in the video. However, how to effectively utilize sound signals for segmentation remains an open challenge. We introduce an innovative segmentation network in this manuscript, referred to as Two-stream Network for Audio Visual Segmentation. The network comprises two distinct streams, namely Multi-Modality Stream and Video-Feature Stream, which respectively provide the model with fusion features and deep visual features. In particular, Multi-Modality Stream is responsible for fusing features from two modalities to generate fusion features, while Video-Feature Stream employs self-attention to extract deeper visual features. Then, Two output features are combined by Cross-Modality Fusion Module, which enables to adaptively keep the balance between two different streams based on the influence extent of noise in the audio. In our experiments on AVSBench, our method surpasses several current methods, showcasing advanced performance while utilizing the same backbone.