Introduction
摘要
Video moment localization, also known as video moment retrieval, aims to identify specific segments within a video that correspond to a given query expressed in natural language. Unlike traditional temporal action localization, where the target actions are predefined and fixed, video moment localization enables the querying of a wide range of complex and dynamic activities that may not be limited to a specific set of action categories. This chapter provides a comprehensive background on the task, highlighting its significance in video understanding and its distinction from related tasks such as action recognition and temporal action localization. We outline the main challenges associated with video moment localization, including the complexity of natural language queries, the varying durations and characteristics of video moments, and the multimodal nature of the task. Additionally, the chapter discusses the structure of the book.