Multimodal Sentiment Analysis (MSA) leverages multimodal signals to detect the sentiment of a video clip. Despite the impressive performance of previous MSA approaches, the issue of inherent multimodal representation redundancy still persists, and the contribution of different modalities varies significantly. In this work, we propose a novel contrastive decoupling distillation (ConD2) framework to reduce information redundancy between multimodal representations and mitigate distribution gaps between modalities through graph distillation. Specifically, we design a contrastive learning strategy to learn fully decoupled two types of multimodal representations: General Information Representations (GIR) and Sentiment Information Representations (SIR). Additionally, ConD2 utilizes a graph distillation unit to balance the differences between distributions for better multimodal fusion. Experiments on two popular MSA benchmarks, CMU-MOSI and CMU-MOSEI, show that ConD2 outperforms all prior methods on a variety of performance metrics.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ConD2: Contrastive Decomposition Distilling for Multimodal Sentiment Analysis

  • Xi Yu,
  • Wenti Huang,
  • Jun Long

摘要

Multimodal Sentiment Analysis (MSA) leverages multimodal signals to detect the sentiment of a video clip. Despite the impressive performance of previous MSA approaches, the issue of inherent multimodal representation redundancy still persists, and the contribution of different modalities varies significantly. In this work, we propose a novel contrastive decoupling distillation (ConD2) framework to reduce information redundancy between multimodal representations and mitigate distribution gaps between modalities through graph distillation. Specifically, we design a contrastive learning strategy to learn fully decoupled two types of multimodal representations: General Information Representations (GIR) and Sentiment Information Representations (SIR). Additionally, ConD2 utilizes a graph distillation unit to balance the differences between distributions for better multimodal fusion. Experiments on two popular MSA benchmarks, CMU-MOSI and CMU-MOSEI, show that ConD2 outperforms all prior methods on a variety of performance metrics.