Multi-hierarchical semantic graph learning for video moment retrieval
摘要
Retrieving video clips that match textual descriptions from untrimmed videos poses a significant challenge, often encountering issues of modality misalignment when integrating full-text information with clip or moment visual embeddings. To address this, we introduce the Multi-Hierarchical Semantic Graph Learning Network (MHSG), which aligns holistic visual data with comprehensive text information and segments visual details with keyword features, ensuring a more accurate cross-modal correspondence. MHSG leverages multi-hierarchical semantic structures to capture visual information at various levels of content granularity, aligning these extracted features with hierarchical text graphs based on inherent semantic relationships. Furthermore, we propose a transition to dense sampling and a local-global feature extraction strategy to improve clip prediction accuracy. Experiments demonstrate that MHSG achieves state-of-the-art performance on multiple benchmarks. Specifically, when using C3D features in the Charades-STA, ActivityNet Caption, and TACoS benchmarks, the model improves the R@1, IoU = 0.5 scores by 1.32%, 1.22%, and 0.12%, respectively. By achieving clear and explicit alignment through multi-hierarchical semantic structures, MHSG effectively enhances the retrieval of video clips matching textual descriptions.