<p>Traditionally, textual data has been used to perform sentiment analysis for applications such as customer feedback analysis, and social media monitoring. There is a shift toward multimodal sentiment analysis (MSA) due to the growing presence of multimedia content such as images, videos, and text to understand the sentiment more comprehensively. Despite these advancements, MSA's current models lack the diversity to assign appropriate weights to each modality based on their ability to distinguish between sentiments, resulting in decreased reliability and potential biases when dealing with datasets with imbalanced distributions. To address these constraints, we proposed a novel framework named the multihead attention for memes (MHAM) model which allocates equal importance to each modality. The MHAM framework has a multimodal feature encoder (MMFE) for extracting text and image features by using BERT with BiLSTM and ResNeXt-SE to represent the high-feature dimensionality, and for enhancing the sentiment information, an additional module named text–image interaction (TII) is introduced which facilitates cross-modal interactions. Our model provides superior performance across various state-of-the-art methods and achieves AUC scores ranging from 0.66 to 0.91 and an average accuracy of up to 85.21% on three widely used datasets, and MHAM sets a new standard in the area of multimodal sentiment analysis.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MHAM: a novel framework for multimodal sentiment analysis in memes

  • Bhavana Verma,
  • Priyanka Meel,
  • Dinesh Kumar Vishwakarma

摘要

Traditionally, textual data has been used to perform sentiment analysis for applications such as customer feedback analysis, and social media monitoring. There is a shift toward multimodal sentiment analysis (MSA) due to the growing presence of multimedia content such as images, videos, and text to understand the sentiment more comprehensively. Despite these advancements, MSA's current models lack the diversity to assign appropriate weights to each modality based on their ability to distinguish between sentiments, resulting in decreased reliability and potential biases when dealing with datasets with imbalanced distributions. To address these constraints, we proposed a novel framework named the multihead attention for memes (MHAM) model which allocates equal importance to each modality. The MHAM framework has a multimodal feature encoder (MMFE) for extracting text and image features by using BERT with BiLSTM and ResNeXt-SE to represent the high-feature dimensionality, and for enhancing the sentiment information, an additional module named text–image interaction (TII) is introduced which facilitates cross-modal interactions. Our model provides superior performance across various state-of-the-art methods and achieves AUC scores ranging from 0.66 to 0.91 and an average accuracy of up to 85.21% on three widely used datasets, and MHAM sets a new standard in the area of multimodal sentiment analysis.