Learning Fine-Grained Cross-Modal Features for Weakly Supervised Video Grounding
摘要
The video grounding task aims to locate clips that are semantically consistent with a given query from untrimmed video. Since the annotation process of this task is labor-intensive and suffers from severe annotation bias, researchers have shown increased interest in the weakly supervised paradigm in recent years. However, due to the differences between video and text features, how to better align video and text is an urgent problem to be solved. Inspired by recent applications of large language models in the field of video understanding, in this paper, we propose an interaction network to learn Fine-grained Cross-modal Feature (FCF) between video and query, which enables it to focus on key information under sparse video frames and keywords interference. Additionally, to address the issue of the one-sidedness in reconstruction-based methods, We meticulously design a similarity loss to measure the association between the query and proposals, and combined this similarity loss with reconstruction loss to rank the proposals. We conduct experiments on the public dataset Charades-STA to verify the effectiveness of the proposed method. Compared with the baseline, the performance of the proposed method has been comprehensively improved.