<p>We focus on localizing a specific video moment given a description query within a raw and untrimmed video. Accurate localization is challenging due to the requirement for distinction among similar candidates. Generally, though existing approaches consider the verbs and nouns in the query, they fail to pay attention to the inferential and key words that play crucial roles in a query and contribute a lot to video localization with natural language. To improve localization, we proposed a Multimodal Representation Reconstruction for Video Localization (MRRVL) model to jointly model visual and textual representations. To be specific, the model transforms the query into a semantic graph according to semantic roles. The detailed textual representations are updated by graph reasoning, which can guide visual representations to be reconstructed. As the specific and context information is mined from both the video and query sufficiently, it is very discriminative for similar segments. Comprehensive experimental results demonstrate that our MRRVL outperforms related approaches on two benchmark datasets significantly.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal representation reconstruction for video localization with natural language

  • Qin Wang,
  • Libiao Jiang,
  • Le Ma,
  • Hailan Jiang,
  • Xizhi Hu,
  • Siyu Jiang

摘要

We focus on localizing a specific video moment given a description query within a raw and untrimmed video. Accurate localization is challenging due to the requirement for distinction among similar candidates. Generally, though existing approaches consider the verbs and nouns in the query, they fail to pay attention to the inferential and key words that play crucial roles in a query and contribute a lot to video localization with natural language. To improve localization, we proposed a Multimodal Representation Reconstruction for Video Localization (MRRVL) model to jointly model visual and textual representations. To be specific, the model transforms the query into a semantic graph according to semantic roles. The detailed textual representations are updated by graph reasoning, which can guide visual representations to be reconstructed. As the specific and context information is mined from both the video and query sufficiently, it is very discriminative for similar segments. Comprehensive experimental results demonstrate that our MRRVL outperforms related approaches on two benchmark datasets significantly.