Semantic Enhanced Video Moment Localization
摘要
Video moment localization involves accurately identifying specific segments within a video that correspond to a given query, which includes determining the precise start and end points of the segment. This task becomes challenging due to the varying durations and diverse spatial-temporal characteristics of video moments. Effective localization requires a deeper understanding not only of the features of the target moment but also of its surrounding context. For instance, temporal constraint words such as “first” in the query necessitate a comprehensive understanding of the broader temporal relationships within the video for accurate localization. To address these challenges, we propose an Attentive Cross-Modal Retrieval Network (ACRN) in this chapter. Our method introduces a memory attention mechanism that prioritizes the visual features referenced in the query while incorporating contextual information to generate enriched moment representations. Additionally, we design a cross-modal fusion sub-network that captures both intra-modality and inter-modality interactions, further enhancing the representation of the moment-query pair. We evaluate our approach on two benchmark datasets: DiDeMo and TACoS. Extensive experimental results demonstrate the effectiveness of our method, showing significant improvements over state-of-the-art approaches.