Multimodal aspect-based sentiment analysis based on a dual syntactic graph network and joint contrastive learning
摘要
Multimodal aspect-based sentiment analysis aims to integrate various modalities of information, such as images and text, through comprehensive analysis to more accurately capture user sentiments across different aspects. Existing multimodal research methods fail to fully utilize rich syntactic information in the text and do not comprehensively align intermodal and intramodal features before downstream tasks, thereby limiting the overall model performance. Therefore, we propose a novel method, multimodal aspect-based sentiment analysis based on a dual syntactic graph network and joint contrastive learning, to enhance the model’s sentiment classification performance. Specifically, we first construct a sentiment knowledge-enhanced graph by combining sentiment knowledge and syntactic features. Then, we utilize the graph convolutional networks and the graph attention networks to capture local syntactic information and global attention weights, respectively. Subsequently, we devise a bidirectional fusion mechanism to integrate dual-channel features. Furthermore, a multisource feature alignment method is designed through contrastive learning to achieve consistency between and within modalities in the feature representation space. Finally, the integrated features are used for multimodal aspect-based sentiment analysis, and contrastive learning is used to improve the overall model’s sentiment classification accuracy. Extensive experiments are conducted on two benchmark datasets, TWITTER-2015 and TWITTER-2017, and the results validate the effectiveness of our proposed model.