Transformer and LLM-Based Captioning Module for Dense Video Captioning
摘要
Real-world video recordings frequently encompass a diverse range of events that can either unfold sequentially or happen simultaneously. Dense video captioning encompasses the tasks of event localization within a video, and the generation of detailed captions for each identified event. To provide a precise depiction of an occurrence at a certain moment, it is essential to include relevant contextual information from both preceding and subsequent events. The transformer, an LLM-based captioning module, and a Bi-LSTM-based event proposal module are used in the proposed dense video captioning model. The proposed model demonstrates a noteworthy improvement of approximately 8% on the ActivityNet dataset. This represents a notable improvement over the current leading techniques in dense video captioning.