Audiovisual Saliency Models: A Short Review
摘要
This review discusses the development and challenges of computational saliency, focusing on the integration of visual and auditory features. Over the past four decades, numerous visual saliency models have been developed, first relying on handcrafted features and later adopting deep learning techniques. While deep saliency models have significantly improved visual attention predictions, they still fall short in fully replicating human behavior, particularly when accounting for audiovisual interactions. The review emphasizes the underexplored role of sound in dynamic visual scenes, highlighting that auditory cues, such as speech or environmental sounds, can influence where people direct their gaze. Early audiovisual saliency models focused on simple handcrafted features like visual motion and sound onsets. However, more recent advancements have incorporated deep features using various fusion schemes, such as consistency-aware predictive coding. While these models are efficient, especially in specific contexts like conversation scenes, they struggle with generalizing to natural settings, particularly when the auditory and visual components are weakly correlated. The review suggests that future models should go beyond simple audiovisual consistency by considering more complex and subtle auditory effects, such as emotional tone in speech or music, which can significantly alter visual attention.