Multimodal representation reconstruction for video localization with natural language
摘要
We focus on localizing a specific video moment given a description query within a raw and untrimmed video. Accurate localization is challenging due to the requirement for distinction among similar candidates. Generally, though existing approaches consider the verbs and nouns in the query, they fail to pay attention to the inferential and key words that play crucial roles in a query and contribute a lot to video localization with natural language. To improve localization, we proposed a Multimodal Representation Reconstruction for Video Localization (MRRVL) model to jointly model visual and textual representations. To be specific, the model transforms the query into a semantic graph according to semantic roles. The detailed textual representations are updated by graph reasoning, which can guide visual representations to be reconstructed. As the specific and context information is mined from both the video and query sufficiently, it is very discriminative for similar segments. Comprehensive experimental results demonstrate that our MRRVL outperforms related approaches on two benchmark datasets significantly.