Text-Dominant Interactive Attention for Cross-Modal Sentiment Analysis
摘要
In today’s era, Cross-modal Sentiment Analysis (CSA) finds widespread application, where the combination of image and text data allows for a richer expression of emotional content. Effectively capturing the semantic information contained in associated image and text data plays a crucial role. This paper proposes a Text-Dominated Interactive Attention cross-modal sentiment analysis network (TDiA), which utilizes cross-modal interactive encoding and semantic correlation of uni-modal features for sentiment analysis. In TDiA, the text-dominant Cross-Modal Attention encoding (CMA) process is employed to adequately interact the encoded uni-modal features into cross-modal features. Simultaneously, it is designed to utilize the Semantic Correlation Space (SCS) to fully associate semantically similar parts of heterogeneous uni-modal features in a low-dimensional space. The encoded cross-modal features undergo modal fusion modules to obtain the final classification. TDiA validates the improvement of the model on the publicly available cross-fmodal sentiment analysis dataset MVSA.