<p>Multimodal sentiment analysis aims to accurately assess the sentiment expressed in a given data source by integrating and analyzing multiple modalities, such as text and images. Extracting discriminative features for sentiment prediction is a powerful approach to address the challenges in multimodal sentiment analysis. Most methods in this domain leverage pre-trained unimodal models to extract features from individual modalities. Subsequently, these features undergo integration via sophisticated fusion mechanisms. However, these models often need to be improved in their ability to proficiently process multimodal data, potentially risking the loss of semantic associations between the different modalities. This study aims to address this problem in multimodal sentiment analysis by developing a simple end-to-end model that avoids the need for sophisticated ensemble techniques for feature extraction. In contrast, the proposed methodology capitalizes on the benefits of transfer learning through the deployment of a vision-language pre-trained model. This model efficiently extracts both visual and textual features within a cohesive framework. Extracted features are subsequently integrated via the proposed feature interaction module, which facilitates capturing potential semantic information in an image-text pair through explicit and implicit feature interaction. Finally, the derived representations undergo transmission to the classification module, thereby augmenting performance in sentiment analysis tasks. The effectiveness of the proposed approach is substantiated through a rigorous experimental evaluation. Assessments conducted on two publicly available real-world datasets reveal significant enhancements in sentiment analysis performance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving multimodal sentiment prediction through vision-language feature interaction

  • Jieyu An,
  • Binfen Ding,
  • Wan Mohd Nazmee Wan Zainon

摘要

Multimodal sentiment analysis aims to accurately assess the sentiment expressed in a given data source by integrating and analyzing multiple modalities, such as text and images. Extracting discriminative features for sentiment prediction is a powerful approach to address the challenges in multimodal sentiment analysis. Most methods in this domain leverage pre-trained unimodal models to extract features from individual modalities. Subsequently, these features undergo integration via sophisticated fusion mechanisms. However, these models often need to be improved in their ability to proficiently process multimodal data, potentially risking the loss of semantic associations between the different modalities. This study aims to address this problem in multimodal sentiment analysis by developing a simple end-to-end model that avoids the need for sophisticated ensemble techniques for feature extraction. In contrast, the proposed methodology capitalizes on the benefits of transfer learning through the deployment of a vision-language pre-trained model. This model efficiently extracts both visual and textual features within a cohesive framework. Extracted features are subsequently integrated via the proposed feature interaction module, which facilitates capturing potential semantic information in an image-text pair through explicit and implicit feature interaction. Finally, the derived representations undergo transmission to the classification module, thereby augmenting performance in sentiment analysis tasks. The effectiveness of the proposed approach is substantiated through a rigorous experimental evaluation. Assessments conducted on two publicly available real-world datasets reveal significant enhancements in sentiment analysis performance.