In today’s era, Cross-modal Sentiment Analysis (CSA) finds widespread application, where the combination of image and text data allows for a richer expression of emotional content. Effectively capturing the semantic information contained in associated image and text data plays a crucial role. This paper proposes a Text-Dominated Interactive Attention cross-modal sentiment analysis network (TDiA), which utilizes cross-modal interactive encoding and semantic correlation of uni-modal features for sentiment analysis. In TDiA, the text-dominant Cross-Modal Attention encoding (CMA) process is employed to adequately interact the encoded uni-modal features into cross-modal features. Simultaneously, it is designed to utilize the Semantic Correlation Space (SCS) to fully associate semantically similar parts of heterogeneous uni-modal features in a low-dimensional space. The encoded cross-modal features undergo modal fusion modules to obtain the final classification. TDiA validates the improvement of the model on the publicly available cross-fmodal sentiment analysis dataset MVSA.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Text-Dominant Interactive Attention for Cross-Modal Sentiment Analysis

  • Zebao Zhang,
  • Shuang Yang,
  • Haiwei Pan

摘要

In today’s era, Cross-modal Sentiment Analysis (CSA) finds widespread application, where the combination of image and text data allows for a richer expression of emotional content. Effectively capturing the semantic information contained in associated image and text data plays a crucial role. This paper proposes a Text-Dominated Interactive Attention cross-modal sentiment analysis network (TDiA), which utilizes cross-modal interactive encoding and semantic correlation of uni-modal features for sentiment analysis. In TDiA, the text-dominant Cross-Modal Attention encoding (CMA) process is employed to adequately interact the encoded uni-modal features into cross-modal features. Simultaneously, it is designed to utilize the Semantic Correlation Space (SCS) to fully associate semantically similar parts of heterogeneous uni-modal features in a low-dimensional space. The encoded cross-modal features undergo modal fusion modules to obtain the final classification. TDiA validates the improvement of the model on the publicly available cross-fmodal sentiment analysis dataset MVSA.