An end-to-end dense video captioning network with multimodal data fusion and multi-level visual feature enhancement via the ViFi-CLIP pretrained model
摘要
Dense video captioning is a multimodal task that involves visual feature extraction and natural language generation. Previous approaches face several challenges, such as reliance on complex "localize-then-describe" frameworks, limited contextual information, and the potential for inaccurate detection. In this paper, we propose an end-to-end dense video captioning model that leverages multimodal data and multi-level visual features, based on the pretrained ViFi-CLIP model. Our approach constructs an integrated network that unifies video segmentation and caption generation into a single end-to-end framework, effectively utilizing visual, audio, and textual modalities. In selecting visual features, we comprehensively consider the impact of grid and region-based features and employ ViFi-CLIP to map these features into a joint vision–language embedding space, thereby enhancing the semantic relevance of the visual representations. Extensive experiments conducted on the ActivityNet dataset demonstrate the effectiveness of our model. Results show that our approach outperforms PDVC by 8–10% across various metrics, and exceeds MDVC by 20–40% in METEOR and ROUGE-L scores, with over 100% improvements in BLEU-4 and CIDEr scores. Furthermore, our model consistently surpasses existing state-of-the-art dense video captioning methods, highlighting its superior performance.