Dense video captioning aims to automatically localize and generate captions for all events in untrimmed videos. Motivated by cognitive informatics, which shows how humans recall relevant scenes from observed cues, we integrate a scene knowledge bank to emulate this process using both visual and speech data. While prior work predominantly focused on visual inputs, speech as a critical multimodal cue has been underutilized, despite offering contextual and descriptive insights that enhance visual understanding. To fully leverage speech alongside scene knowledge, we propose the Scene Knowledge Enhanced Multimodal Retrieval model. This model constructs a scene knowledge bank from ground truth captions in the training videos, then enhances visual representations by merging them with pertinent text features retrieved through multimodal retrieval involving speech and visual inputs. A specialized Knowledge Enhancement Transformer facilitates enhancement of visual features. Experimental results on the ViTT and YouCook2 datasets demonstrate that our proposed method achieves competitive performance compared to existing approaches.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Scene Knowledge Enhanced Multimodal Retrieval Model for Dense Video Captioning

  • Mingru Huang,
  • Pengfei Duan,
  • Yifang Zhang,
  • Huimin Chen,
  • Jiawang Peng,
  • Shengwu Xiong

摘要

Dense video captioning aims to automatically localize and generate captions for all events in untrimmed videos. Motivated by cognitive informatics, which shows how humans recall relevant scenes from observed cues, we integrate a scene knowledge bank to emulate this process using both visual and speech data. While prior work predominantly focused on visual inputs, speech as a critical multimodal cue has been underutilized, despite offering contextual and descriptive insights that enhance visual understanding. To fully leverage speech alongside scene knowledge, we propose the Scene Knowledge Enhanced Multimodal Retrieval model. This model constructs a scene knowledge bank from ground truth captions in the training videos, then enhances visual representations by merging them with pertinent text features retrieved through multimodal retrieval involving speech and visual inputs. A specialized Knowledge Enhancement Transformer facilitates enhancement of visual features. Experimental results on the ViTT and YouCook2 datasets demonstrate that our proposed method achieves competitive performance compared to existing approaches.