Video data has become indispensable to people’s daily lives. Retrieving video clips relevant to user queries from massive datasets has become a significant research focus. However, current methods still have two major limitations: (1) independent moment encoding, which fails to consider the contextual information fully, (2) lack of utilization of fine-grained information, ignoring boundary information. This paper proposes a video clip retrieval method based on multimodal information fusion to address the above issues and improve retrieval accuracy, namely VRMF. This method derives video encodings that simultaneously integrate contextual, fine-grained, and user query information by exploring the relationships between candidate clips, fusing fine-grained cross-modal information of videos and user queries, and incorporating boundary information of different video clips. Finally, results are obtained through a scoring and ranking algorithm. We conduct experiments on three public datasets, showing that VRMF outperforms baseline models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Video Moment Retrieval Based on Multimodal Information Fusion

  • Lu Zhang,
  • Xunyuan Liu,
  • Yingxuan Guan,
  • Ying Xing,
  • Yaru Zhao

摘要

Video data has become indispensable to people’s daily lives. Retrieving video clips relevant to user queries from massive datasets has become a significant research focus. However, current methods still have two major limitations: (1) independent moment encoding, which fails to consider the contextual information fully, (2) lack of utilization of fine-grained information, ignoring boundary information. This paper proposes a video clip retrieval method based on multimodal information fusion to address the above issues and improve retrieval accuracy, namely VRMF. This method derives video encodings that simultaneously integrate contextual, fine-grained, and user query information by exploring the relationships between candidate clips, fusing fine-grained cross-modal information of videos and user queries, and incorporating boundary information of different video clips. Finally, results are obtained through a scoring and ranking algorithm. We conduct experiments on three public datasets, showing that VRMF outperforms baseline models.