A dissimilarity feature-driven decomposition network for multimodal sentiment analysis
摘要
Multimodal sentiment analysis (MSA) aims to predict human emotions via language, visual and acoustic modalities. Based on the idea of extracting modalities into modality-invariant and modality-specific components, feature decomposition methods decompose features into similarity and dissimilarity features. However, dissimilarity features contribute less than similarity features due to emotional divergence. And they have not been effectively utilized due to insufficient methods for feature extraction and processing. To address these issues, we propose a dissimilarity feature-driven decomposition network (DFDDN) for MSA. We deploy a feature extraction module to extract and enhance features. This not only increases the differences between the features, but also enables us to focus more on the emotional information contained in the features. We utilize different encoders to decompose features, and design loss functions to increase the differences of modality features. Compared to state-of-the-art (SOTA) method on the CMU-MOSI dataset, there are improvements of 0.35%/1.57% and 0.28%/2.09% in Acc2 and macro F1, a 3.31% improvement in Acc7, and the MAE decreases by 0.1. Compared to SOTA method on the CMU-MOSEI dataset, there is an improvement of 1.34% in Acc7, the Corr increases by 0.016, and the MAE decreases by 0.005.