Most real-life videos encompass a diverse array of events, which may be sequential or overlapping. Dense video captioning involves the identification and localization of events within a video, as well as the generation of descriptive captions for each event. To accurately describe an event at a specific time step, it is crucial to integrate contextual information from both past and future events. Our proposed approach utilizes two uni-directional LSTM-based captioning modules that synthesize contextual information from both visual and textual data in forward and backward directions to generate dense video caption. This model demonstrates a significant advancement over the leading dense video captioning methods, achieving a relative improvement of approximately 9% on the ActivityNet dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Two Uni-directional LSTMs-Based Captioning Module for Dense Video Captioning

  • Dvijesh Bhatt,
  • Priyank Thakkar

摘要

Most real-life videos encompass a diverse array of events, which may be sequential or overlapping. Dense video captioning involves the identification and localization of events within a video, as well as the generation of descriptive captions for each event. To accurately describe an event at a specific time step, it is crucial to integrate contextual information from both past and future events. Our proposed approach utilizes two uni-directional LSTM-based captioning modules that synthesize contextual information from both visual and textual data in forward and backward directions to generate dense video caption. This model demonstrates a significant advancement over the leading dense video captioning methods, achieving a relative improvement of approximately 9% on the ActivityNet dataset.