<p>The application of visual language pretraining (VLP) models has gradually become a research hotspot in multi-modal sentiment analysis due to the explosive growth of multi-modal data on social media. However, existing methods still have limitations in information interaction between modalities, feature fusion, and handling the negative impact of weak correlations between modalities. This paper proposes the CLIP-driven attention network (CDAN), a framework based on the contrastive language-image pretraining (CLIP) pre-trained model. CDAN further explores the potential of VLP models. Specifically, CDAN uses two single-modality encoders and two CLIP encoders that separately extract text and image features to capture deep semantic information from both modalities. For the image–text data projected into a shared feature space, we design a novel attention mechanism for fine-grained feature extraction and modality fusion. At the same time, we fully leverage the deep semantic relational information captured by the CLIP pre-trained model as the basis for dynamically adjusting the image–text features, thereby weighting the important features. Furthermore, we use a self-supervised decoder to reconstruct the CLIP-fused features and obtain additional CLIP-optimized features, which are incorporated as part of the self-adaptive weighted aggregation features in the final fusion module’s modality attention mechanism. This method successfully mitigates the negative impact of weakly correlated modalities on the task. Experimental results show that CDAN achieves an accuracy of 78.5% on the MVSA-Single (MVSA-S) dataset and 73.2% on the MVSA-Multiple (MVSA-M) dataset. Our model addresses the challenges of modality interaction and cross-modal integration, demonstrating CDAN’s potential advantages in multi-modal sentiment analysis tasks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CLIP-driven attention network for multimodal sentiment analysis

  • Jialun Lv,
  • Qimeng Yang,
  • Shengwei Tian,
  • Bo Liu,
  • Long Yu

摘要

The application of visual language pretraining (VLP) models has gradually become a research hotspot in multi-modal sentiment analysis due to the explosive growth of multi-modal data on social media. However, existing methods still have limitations in information interaction between modalities, feature fusion, and handling the negative impact of weak correlations between modalities. This paper proposes the CLIP-driven attention network (CDAN), a framework based on the contrastive language-image pretraining (CLIP) pre-trained model. CDAN further explores the potential of VLP models. Specifically, CDAN uses two single-modality encoders and two CLIP encoders that separately extract text and image features to capture deep semantic information from both modalities. For the image–text data projected into a shared feature space, we design a novel attention mechanism for fine-grained feature extraction and modality fusion. At the same time, we fully leverage the deep semantic relational information captured by the CLIP pre-trained model as the basis for dynamically adjusting the image–text features, thereby weighting the important features. Furthermore, we use a self-supervised decoder to reconstruct the CLIP-fused features and obtain additional CLIP-optimized features, which are incorporated as part of the self-adaptive weighted aggregation features in the final fusion module’s modality attention mechanism. This method successfully mitigates the negative impact of weakly correlated modalities on the task. Experimental results show that CDAN achieves an accuracy of 78.5% on the MVSA-Single (MVSA-S) dataset and 73.2% on the MVSA-Multiple (MVSA-M) dataset. Our model addresses the challenges of modality interaction and cross-modal integration, demonstrating CDAN’s potential advantages in multi-modal sentiment analysis tasks.