<p>Multimodal sentiment analysis has made significant progress in the field of natural language processing by fusing various modalities of data. The main challenges in this field lie in the complexity of modality feature extraction and the issue of multimodal sentiment semantic inconsistency. To address these challenges, we propose a Self-attention Mechanism Prior to Modality Fusion for multimodal sentiment analysis (SAM-PMF). Specifically, this paper first constructs a modality feature extraction module to encode the extracted multimodal features. Then, we design a self-attention mechanism module to enhance the quality of each modality feature, facilitating subsequent multimodal fusion. Finally, we propose a modality feature fusion module that combines CMD and a Cross-Attention to fuse information across modalities into a unified multimodal feature representation, effectively addressing the semantic inconsistency between modalities. The proposed model is evaluated on two publicly available multimodal sentiment analysis datasets, MOSI and MOSEI. Experimental results show that compared to the baseline models, SAM-PMF demonstrated significant improvements across most evaluation metrics, confirming its effectiveness in multimodal sentiment analysis tasks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Self-attention mechanism prior to modality fusion for multimodal sentiment analysis

  • Zhongyuan Chen,
  • Chong Lu,
  • Yihan Wang

摘要

Multimodal sentiment analysis has made significant progress in the field of natural language processing by fusing various modalities of data. The main challenges in this field lie in the complexity of modality feature extraction and the issue of multimodal sentiment semantic inconsistency. To address these challenges, we propose a Self-attention Mechanism Prior to Modality Fusion for multimodal sentiment analysis (SAM-PMF). Specifically, this paper first constructs a modality feature extraction module to encode the extracted multimodal features. Then, we design a self-attention mechanism module to enhance the quality of each modality feature, facilitating subsequent multimodal fusion. Finally, we propose a modality feature fusion module that combines CMD and a Cross-Attention to fuse information across modalities into a unified multimodal feature representation, effectively addressing the semantic inconsistency between modalities. The proposed model is evaluated on two publicly available multimodal sentiment analysis datasets, MOSI and MOSEI. Experimental results show that compared to the baseline models, SAM-PMF demonstrated significant improvements across most evaluation metrics, confirming its effectiveness in multimodal sentiment analysis tasks.