<p>Video memorability refers to the extent to which videos are recalled after viewing, playing a crucial role in obtaining memorable video content. Existing models often rely on extracting multimodal features to predict video memorability scores, yet they tend to overlook the effective utilization of motion cues. The representation capability of motion features is adversely affected during the fine-tuning stage of the motion feature extractor due to insufficient labeled data. In this paper, we introduce the Text-Motion Cross-modal Contrastive Loss (TMCCL) multimodal video memorability prediction model to bolster the representation of motion features. We address the challenge of enhancing motion feature representation by considering text description similarities among videos as a criterion for establishing positive and negative motion sample sets for a target sample. The improved motion features are anticipated to learn similar feature representations for semantically related motion content, leading to more accurate predictions of video memorability scores. Our proposed model demonstrates state-of-the-art performance on two video memorability prediction datasets. Additionally, the application prospects related to video memorability prediction have received limited attention. To address this, we introduce Memorability Weighted Correction for Video Summarization (MWCVS), applying video memorability prediction to mitigate the issue of human subjectivity in video summarization labels. Experimental results on two video summarization datasets underscore the effectiveness of MWCVS, highlighting the potential of video memorability prediction applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Text-motion cross-modal contrastive loss for video memorability prediction and memorability weighted correction

  • Zhiyi Zhu,
  • Xiaoyu Wu,
  • Youwei Lu

摘要

Video memorability refers to the extent to which videos are recalled after viewing, playing a crucial role in obtaining memorable video content. Existing models often rely on extracting multimodal features to predict video memorability scores, yet they tend to overlook the effective utilization of motion cues. The representation capability of motion features is adversely affected during the fine-tuning stage of the motion feature extractor due to insufficient labeled data. In this paper, we introduce the Text-Motion Cross-modal Contrastive Loss (TMCCL) multimodal video memorability prediction model to bolster the representation of motion features. We address the challenge of enhancing motion feature representation by considering text description similarities among videos as a criterion for establishing positive and negative motion sample sets for a target sample. The improved motion features are anticipated to learn similar feature representations for semantically related motion content, leading to more accurate predictions of video memorability scores. Our proposed model demonstrates state-of-the-art performance on two video memorability prediction datasets. Additionally, the application prospects related to video memorability prediction have received limited attention. To address this, we introduce Memorability Weighted Correction for Video Summarization (MWCVS), applying video memorability prediction to mitigate the issue of human subjectivity in video summarization labels. Experimental results on two video summarization datasets underscore the effectiveness of MWCVS, highlighting the potential of video memorability prediction applications.