Weakly-Supervised Video Moment Localization
摘要
Video moment localization, which aims to localize specific moments within a video based on a descriptive query, has gained significant attention in recent years. Despite substantial progress, most existing methods are supervised and rely on moment-level temporal annotations. In contrast, weakly-supervised approaches, which require only video-level annotations, remain relatively underexplored. In this chapter, we propose a novel end-to-end Siamese alignment network for weakly-supervised video moment retrieval. Specifically, we introduce a multi-scale Siamese module that progressively reduces the semantic gap between the visual and textual modalities. Furthermore, we propose a context-aware multiple instance learning module, which incorporates the influence of adjacent contexts, enhancing both moment-query and video-query alignment simultaneously. By optimizing both moment-level and video-level matching, our model effectively improves retrieval performance even with only weak video-level annotations. Extensive experiments on two benchmark datasets, ActivityNet Captions and Charades-STA, demonstrate the superior performance of our model over several state-of-the-art baselines.