The video moment retrieval with the natural language area aims to locate the segment (moment) of the video most relevant to a textual description (natural language). However, existing methods are based only on the image sequence analysis and neglect the information derived from the audio. Thus, the main objective of this study is to combine both features (from image and audio) to make the retrieval more comprehensive and robust. For this, a model is built on audio and image sequence extractors aligned that relate to the textual description to retrieve the desired moment of the video. We proposed a weakly supervised model that uses attention mechanisms and the audio component for video moment retrieval by natural language. Results demonstrate that the proposed model outperforms the current state-of-the-art in the metric mIoU by more than 27%, in addition to decreasing the response time of the video moment retrieval (reducing the computational complexity from polynomial to linear).

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Combining Audio and Image Sequence for Video Moment Retrieval by Natural Language

  • Luís G. de Souza,
  • Sílvio R. R. Sanches,
  • Pedro H. Bugatti,
  • Priscila T. M. Saito

摘要

The video moment retrieval with the natural language area aims to locate the segment (moment) of the video most relevant to a textual description (natural language). However, existing methods are based only on the image sequence analysis and neglect the information derived from the audio. Thus, the main objective of this study is to combine both features (from image and audio) to make the retrieval more comprehensive and robust. For this, a model is built on audio and image sequence extractors aligned that relate to the textual description to retrieve the desired moment of the video. We proposed a weakly supervised model that uses attention mechanisms and the audio component for video moment retrieval by natural language. Results demonstrate that the proposed model outperforms the current state-of-the-art in the metric mIoU by more than 27%, in addition to decreasing the response time of the video moment retrieval (reducing the computational complexity from polynomial to linear).