In this chapter we review the main modalities where saliency is computed from the point of view of both classic models and deep learning-based models along with the main datasets used to train and test the different models. This review covers models for image, audio and video, \(360^{\circ}\) images and videos, and RGB-D and 3D data. Overall, we have foreseen two main future research tracks in attention modeling: (1) dynamic fixation models including also fixation duration and (2) multimodality where images, video, sound, depth map, and maybe even new modalities (coming from other sensors, touch, etc.) are taken into account. The arrival of Visual Language Models (VLMs) and other generic large multimodal models might boost the multimodality results.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Attention, Multimodality, and Datasets for Validation

  • Matei Mancas,
  • Erwan David,
  • Nicolas Riche,
  • Julien Leroy

摘要

In this chapter we review the main modalities where saliency is computed from the point of view of both classic models and deep learning-based models along with the main datasets used to train and test the different models. This review covers models for image, audio and video, \(360^{\circ}\) images and videos, and RGB-D and 3D data. Overall, we have foreseen two main future research tracks in attention modeling: (1) dynamic fixation models including also fixation duration and (2) multimodality where images, video, sound, depth map, and maybe even new modalities (coming from other sensors, touch, etc.) are taken into account. The arrival of Visual Language Models (VLMs) and other generic large multimodal models might boost the multimodality results.