Semantic Pruning Video Moment Localization
摘要
Video moment localization, a crucial task in video content analysis, has garnered significant attention in recent years. However, challenges such as cross-modal semantic alignment and localization efficiency still hinder progress in this field. To address these issues, we propose a cross-modal semantic alignment network. Specifically, we design a video encoder that generates moment candidates, learns their representations, and models their semantic relevance. Simultaneously, we develop a query encoder to effectively capture diverse query intentions. To further enhance localization accuracy, we introduce a multi-granularity interaction module that explores semantic correlations across different modalities, enabling precise localization through comprehensive cross-modal understanding. Additionally, we propose a semantic pruning strategy that reduces the computational overhead of cross-modal retrieval, thereby improving localization efficiency. Experimental results on two benchmark datasets demonstrate the superior performance of our model compared to several state-of-the-art approaches.