Real-world video recordings frequently encompass a diverse range of events that can either unfold sequentially or happen simultaneously. Dense video captioning encompasses the tasks of event localization within a video, and the generation of detailed captions for each identified event. To provide a precise depiction of an occurrence at a certain moment, it is essential to include relevant contextual information from both preceding and subsequent events. The transformer, an LLM-based captioning module, and a Bi-LSTM-based event proposal module are used in the proposed dense video captioning model. The proposed model demonstrates a noteworthy improvement of approximately 8% on the ActivityNet dataset. This represents a notable improvement over the current leading techniques in dense video captioning.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Transformer and LLM-Based Captioning Module for Dense Video Captioning

  • Dvijesh Bhatt,
  • Priyank Thakkar

摘要

Real-world video recordings frequently encompass a diverse range of events that can either unfold sequentially or happen simultaneously. Dense video captioning encompasses the tasks of event localization within a video, and the generation of detailed captions for each identified event. To provide a precise depiction of an occurrence at a certain moment, it is essential to include relevant contextual information from both preceding and subsequent events. The transformer, an LLM-based captioning module, and a Bi-LSTM-based event proposal module are used in the proposed dense video captioning model. The proposed model demonstrates a noteworthy improvement of approximately 8% on the ActivityNet dataset. This represents a notable improvement over the current leading techniques in dense video captioning.