WConF: Weighted Contrastive Fusion for Multimodal Sentiment Analysis
摘要
With the continuous development of multimedia, Multimodal Sentiment Analysis has become a highly regarded field. Recent research proposes learning effective unimodal representations to facilitate multimodal fusion, which mainly contain two parts of information: modality-common and modality-specific information. However, previous work does not consider the sentiment span information between different samples during the fusion process. In this paper, we propose a novel framework Weighted Contrastive Fusion (WConF) to extract sentiment span information for multimodal fusion. First, we apply modality contrastive learning to capture modality-common sentiment information and separate modality-specific sentiment information. Then, considering sentiment spans as weights, we perform weighted contrastive learning on the multimodal representations. Moreover, we design a loss function to assist in weighted contrastive fusion. Finally, we conduct extensive experiments on four benchmark datasets: MOSI, MOSEI, CH-SIMS, and CH-SIMSV2. Experimental results demonstrate that our model achieves superior performance compared to previous models.