MMI-FM: Multimodal Interaction and Fusion Mechanism for Video Anomaly Detection
摘要
In video anomaly detection tasks, multi-modal input can encode video content across different feature spaces, offering richer and more complementary semantic representations than single-modal data. However, most existing multi-modal VAD methods employ independent modeling strategies and fuse features at the decision stage typically through simple operations such as concatenation or gating. To address this limitation, we propose a novel Multimodal Interaction and Fusion Mechanism (MMI-FM) based on RGB features and skeleton features. Specifically, we design the Dual Contrastive Loss Alignment Module (DCL-AM), which aligns features by mining both intra-modal and cross-modal semantic correlations. Furthermore, we introduce a Cross-Modal Bidirectional Knowledge Distillation Module (CM-BKDM) to mitigate potential representational biases in single modality representation learning by performing bidirectional knowledge transfer between modalities. Finally, an Adaptive Multimodal Feature Fusion Module (AMFFM) is proposed, which dynamically fuses RGB regions and skeleton points with strong semantic associations. Extensive experiments conducted on two public VAD datasets, CUHK Avenue and ShanghaiTech, demonstrate that MMI-FM achieves AUC scores of 91.8% and 78.4%, outperforming most state-of-the-art methods and validating the effectiveness of our proposed framework.