This review discusses the development and challenges of computational saliency, focusing on the integration of visual and auditory features. Over the past four decades, numerous visual saliency models have been developed, first relying on handcrafted features and later adopting deep learning techniques. While deep saliency models have significantly improved visual attention predictions, they still fall short in fully replicating human behavior, particularly when accounting for audiovisual interactions. The review emphasizes the underexplored role of sound in dynamic visual scenes, highlighting that auditory cues, such as speech or environmental sounds, can influence where people direct their gaze. Early audiovisual saliency models focused on simple handcrafted features like visual motion and sound onsets. However, more recent advancements have incorporated deep features using various fusion schemes, such as consistency-aware predictive coding. While these models are efficient, especially in specific contexts like conversation scenes, they struggle with generalizing to natural settings, particularly when the auditory and visual components are weakly correlated. The review suggests that future models should go beyond simple audiovisual consistency by considering more complex and subtle auditory effects, such as emotional tone in speech or music, which can significantly alter visual attention.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Audiovisual Saliency Models: A Short Review

  • Antoine Coutrot

摘要

This review discusses the development and challenges of computational saliency, focusing on the integration of visual and auditory features. Over the past four decades, numerous visual saliency models have been developed, first relying on handcrafted features and later adopting deep learning techniques. While deep saliency models have significantly improved visual attention predictions, they still fall short in fully replicating human behavior, particularly when accounting for audiovisual interactions. The review emphasizes the underexplored role of sound in dynamic visual scenes, highlighting that auditory cues, such as speech or environmental sounds, can influence where people direct their gaze. Early audiovisual saliency models focused on simple handcrafted features like visual motion and sound onsets. However, more recent advancements have incorporated deep features using various fusion schemes, such as consistency-aware predictive coding. While these models are efficient, especially in specific contexts like conversation scenes, they struggle with generalizing to natural settings, particularly when the auditory and visual components are weakly correlated. The review suggests that future models should go beyond simple audiovisual consistency by considering more complex and subtle auditory effects, such as emotional tone in speech or music, which can significantly alter visual attention.