Dual-stream multi-level interaction network for aspect-based multimodal sentiment analysis
摘要
Aspect-based multimodal sentiment analysis (ABMSA) has made significant progress, but existing methods often overlook the differences between text and image modalities, and have a limited understanding of sentiment, resulting in suboptimal performance. To address these issues, we propose the Dual-Stream Multi-level Interaction Network (DMIN), which leverages comprehensive sentiment information in text-image pairs. Specifically, to prevent information loss caused by reconstructing spatial encodings, we designed a fully convolutional encoding stream for the image modality, forming a dual-stream architecture alongside the text encoding stream. The image stream incorporates an innovative transposed-convolutional attention mechanism to highlight image features relevant to the text. Furthermore, we designed a multi-level fusion architecture that integrates feature-level and decision-level fusion to enhance cross-modal interactions at different depths. Finally, inspired by psychological understanding of sentiment, we propose the Ordinal Valence Loss (OVL) function specifically for sentiment analysis. Experimental results on three benchmark datasets proves that our model has better performance and better robustness to noise compared to previous models, with fewer parameters. Our study validates the effectiveness of DMIN and the ordinal valence loss function, advancing the development of ABMSA and sentiment understanding.