Video Moment Retrieval Based on Multimodal Information Fusion
摘要
Video data has become indispensable to people’s daily lives. Retrieving video clips relevant to user queries from massive datasets has become a significant research focus. However, current methods still have two major limitations: (1) independent moment encoding, which fails to consider the contextual information fully, (2) lack of utilization of fine-grained information, ignoring boundary information. This paper proposes a video clip retrieval method based on multimodal information fusion to address the above issues and improve retrieval accuracy, namely VRMF. This method derives video encodings that simultaneously integrate contextual, fine-grained, and user query information by exploring the relationships between candidate clips, fusing fine-grained cross-modal information of videos and user queries, and incorporating boundary information of different video clips. Finally, results are obtained through a scoring and ranking algorithm. We conduct experiments on three public datasets, showing that VRMF outperforms baseline models.