Visual saliency prediction is critical for understanding human attention in video content and supports various applications. In this paper, we introduce AVSal, an advanced audio-visual saliency prediction model designed to enhance the accuracy of video saliency prediction. AVSal leverages foundation models, specifically CLIP and ImageBind, for robust and high-quality feature extraction from both visual and auditory inputs. Then, a novel cross-attention-based fusion mechanism is employed to effectively integrate audio and visual features at multiple levels, capturing the intricate relationships between these modalities. Additionally, a spatio-temporal GRU architecture is implemented to preserve critical temporal dynamics, improving the model’s accuracy in predicting saliency in dynamic scenes. Extensive experimental results demonstrate that the proposed AVSal model performs excellently in the ECCV AIM Video Saliency Prediction Challenge 2024 and significantly outperforms other state-of-the-art models in six other mainstream audio-visual saliency datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AVSal: Enhancing Video Saliency Prediction Through Audio-Visual Fusion and Temporal Aggregation

  • Yuxin Zhu,
  • Yinan Sun,
  • Huiyu Duan,
  • Yuqin Cao,
  • Ziheng Jia,
  • Qiang Hu,
  • Xiongkuo Min,
  • Guangtao Zhai

摘要

Visual saliency prediction is critical for understanding human attention in video content and supports various applications. In this paper, we introduce AVSal, an advanced audio-visual saliency prediction model designed to enhance the accuracy of video saliency prediction. AVSal leverages foundation models, specifically CLIP and ImageBind, for robust and high-quality feature extraction from both visual and auditory inputs. Then, a novel cross-attention-based fusion mechanism is employed to effectively integrate audio and visual features at multiple levels, capturing the intricate relationships between these modalities. Additionally, a spatio-temporal GRU architecture is implemented to preserve critical temporal dynamics, improving the model’s accuracy in predicting saliency in dynamic scenes. Extensive experimental results demonstrate that the proposed AVSal model performs excellently in the ECCV AIM Video Saliency Prediction Challenge 2024 and significantly outperforms other state-of-the-art models in six other mainstream audio-visual saliency datasets.