Text-Guided Enhanced Transformer Fusion for Multimodal Sentiment Analysis
摘要
Multimodal sentiment analysis benefits from leveraging text, video, and audio data to better infer emotional states, but the heterogeneous nature of these modalities presents challenges. This paper proposes the adaptive text-guided multimodal gated-fusion transformer (ATMGT), a novel model designed to address the challenges. The ATMGT leverages Transformer-based self-attention and cross-attention mechanisms to enable the deep interaction and fusion of the text, audio, and visual modalities. Then, a gated fusion mechanism effectively integrates these audio-visual features, alleviating information redundancy. Additionally, text is treated as a local feature to guide the scaling of global features formed by audio-visual data, emphasizing emotionally significant regions. Furthermore, a self-supervised label generation module enhances modality-specific learning and robust sentiment classification. Experimental comparisons demonstrate that the proposed model achieves competitive performance across multiple metrics on the CMU-MOSI, CMU-MOSEI, and CH-SIMS datasets, outperforming state-of-the-art models. Finally, ablation studies validate the contributions of the core modules to the overall model performance.