ConD2: Contrastive Decomposition Distilling for Multimodal Sentiment Analysis
摘要
Multimodal Sentiment Analysis (MSA) leverages multimodal signals to detect the sentiment of a video clip. Despite the impressive performance of previous MSA approaches, the issue of inherent multimodal representation redundancy still persists, and the contribution of different modalities varies significantly. In this work, we propose a novel contrastive decoupling distillation (ConD2) framework to reduce information redundancy between multimodal representations and mitigate distribution gaps between modalities through graph distillation. Specifically, we design a contrastive learning strategy to learn fully decoupled two types of multimodal representations: General Information Representations (GIR) and Sentiment Information Representations (SIR). Additionally, ConD2 utilizes a graph distillation unit to balance the differences between distributions for better multimodal fusion. Experiments on two popular MSA benchmarks, CMU-MOSI and CMU-MOSEI, show that ConD2 outperforms all prior methods on a variety of performance metrics.