Scene Knowledge Enhanced Multimodal Retrieval Model for Dense Video Captioning
摘要
Dense video captioning aims to automatically localize and generate captions for all events in untrimmed videos. Motivated by cognitive informatics, which shows how humans recall relevant scenes from observed cues, we integrate a scene knowledge bank to emulate this process using both visual and speech data. While prior work predominantly focused on visual inputs, speech as a critical multimodal cue has been underutilized, despite offering contextual and descriptive insights that enhance visual understanding. To fully leverage speech alongside scene knowledge, we propose the Scene Knowledge Enhanced Multimodal Retrieval model. This model constructs a scene knowledge bank from ground truth captions in the training videos, then enhances visual representations by merging them with pertinent text features retrieved through multimodal retrieval involving speech and visual inputs. A specialized Knowledge Enhancement Transformer facilitates enhancement of visual features. Experimental results on the ViTT and YouCook2 datasets demonstrate that our proposed method achieves competitive performance compared to existing approaches.