Fine-Grained Feature Fusion for Self-supervised Multimodal Sentiment Analysis
摘要
Accurate emotion detection relies on the fusion of information from multiple modalities, including text, audio, and video. However, existing fusion methods face challenges such as information redundancy and significant differences in representation between modalities, making it difficult to detect subtle emotional cues, particularly in audio and visual data, and when unimodal annotations are unavailable. In response to the aforementioned issues, we propose the Fine-Grained Multimodal Fusion Network (FG-MFN), which effectively captures subtle emotional variations by adopting multi-directional and multi-scale feature extraction techniques. Specifically, we introduce a fine-grained fusion mechanism that enhances feature representations by promoting cross-dimensional interaction through rotation operations and residual transformations. Additionally, a self-supervised mechanism is employed to generate modality-wise labels, reducing annotation costs and enhancing modality-specific learning. Numerous experiments on the CMU-MOSI, CMU-MOSEI, and CH-SIMS datasets show that FG-MFN consistently surpasses the leading baselines. Quantitative and qualitative results validate the robustness of the proposed approach across multilingual settings.