Abstract <p>Storing similar videos will seriously waste the storage resources. However, many methods have low detection accuracy and efficiency, which cannot extract and analyze the dynamic features of specific local targets effectively. To solve the above problems, a multimodal long video detection algorithm based on multiscenario learning is proposed, which effectively integrates the multimode features of four modes, i.e., text flow, keyframes, optical flow, and audio flow. Firstly, a coevolutionary neural network model is constructed to extract the four pattern features based on the scene as a unit. Our innovation is that keyframe features are obtained by selecting the frames with the lowest similarity to aggregate the features, and optical flow features are obtained by extracting the dynamic running mode of local specific targets. Then, a position cross attention module and a cross modal fusion model is designed to capture the correlation between different modal data effectively. Finally, the chamfer similarity is applied to various video scenes to calculate the similarity matrix. Experiments conducted on CC_WEB_VIDEO and VCDB datasets reveal that the mAP values obtained by the present algorithm are 98.5 and 94.9% respectively, demonstrating the best performance in terms of accuracy in comparison to the baseline methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Scene-Based Multimodal Video Similarity Detection Algorithm

  • Xue Li

摘要

Abstract

Storing similar videos will seriously waste the storage resources. However, many methods have low detection accuracy and efficiency, which cannot extract and analyze the dynamic features of specific local targets effectively. To solve the above problems, a multimodal long video detection algorithm based on multiscenario learning is proposed, which effectively integrates the multimode features of four modes, i.e., text flow, keyframes, optical flow, and audio flow. Firstly, a coevolutionary neural network model is constructed to extract the four pattern features based on the scene as a unit. Our innovation is that keyframe features are obtained by selecting the frames with the lowest similarity to aggregate the features, and optical flow features are obtained by extracting the dynamic running mode of local specific targets. Then, a position cross attention module and a cross modal fusion model is designed to capture the correlation between different modal data effectively. Finally, the chamfer similarity is applied to various video scenes to calculate the similarity matrix. Experiments conducted on CC_WEB_VIDEO and VCDB datasets reveal that the mAP values obtained by the present algorithm are 98.5 and 94.9% respectively, demonstrating the best performance in terms of accuracy in comparison to the baseline methods.