Video Summarization Algorithm Based on Multimodal Multiscale Temporal Conjugate Position Coding
摘要
To address the limitations of the existing DSNet model in video summarization, including insufficient temporal information utilization, poor multi-modal fusion, and weak multi-scale feature capture capabilities leading to incoherent contexts and low-quality summaries, this paper proposes a video summarization algorithm based on multi-modal multi-scale temporal conjugate positional encoding. The framework integrates temporal conjugate positional encoding into the multi-head attention mechanism to fuse video, audio, and text modalities, while introducing a multi-scale feature pyramid structure for video segmentation. The algorithm uses temporal encoding and multi-modal fusion for temporal modeling, boosts cross-hierarchical features via the pyramid, and hierarchically predicts important frame and video segment. Combined with segment classification-regression methods, it simultaneously outputs shot confidence scores and positional offsets to select key shots. Experimental results on standard and augmented versions of SumMe and TVSum datasets demonstrate F-Score values of 54.5%, 54.9% and 62.6%, 63.0%, outperforming the VJMHT model by 3.2%, 2.7% and 1.7%, 1.1% respectively. This demonstrates the algorithm’s enhanced summarization accuracy and theoretical contributions to video summarization.