<p>Sarcasm is now widespread across social media and customer feedback platforms with seamless intertwining of multimodal elements, including text, images, and audio. Sarcasm detection improves sentiment analysis for ensuring accurate interpretation of user opinions and supports mental health assessment by identifying hidden emotions in online conversations. However the combination of vision and language elements makes it harder to detect the true sentiment in sarcasm, as the meaning often depends on context and contradictions between them. Despite the growing interest from the sentiment analysis community in detecting the sentiments in sarcastic cues, their inherent ambiguity poses a significant challenge for both humans and machines to detect the real sentiments in them. Motivated by these challenges, this paper proposes the ALBEF-Sarc model for the image-text based sarcasm detection. The model leverages the Align Before Fuse (ALBEF) framework to effectively align and integrate textual and visual inputs, enhancing the detection of sarcasm across multiple modalities. The performance of the proposed model has been evaluated using a popular benchmarking dataset on multimodal sarcasm detection and achieved an overall accuracy of 87.95% and an F-score of 86.52% that outperforms other state-of-the art methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Applying cross-modal feature alignment and fusion for effective sarcasm detection

  • Sarang P. Karun,
  • V. Adithya

摘要

Sarcasm is now widespread across social media and customer feedback platforms with seamless intertwining of multimodal elements, including text, images, and audio. Sarcasm detection improves sentiment analysis for ensuring accurate interpretation of user opinions and supports mental health assessment by identifying hidden emotions in online conversations. However the combination of vision and language elements makes it harder to detect the true sentiment in sarcasm, as the meaning often depends on context and contradictions between them. Despite the growing interest from the sentiment analysis community in detecting the sentiments in sarcastic cues, their inherent ambiguity poses a significant challenge for both humans and machines to detect the real sentiments in them. Motivated by these challenges, this paper proposes the ALBEF-Sarc model for the image-text based sarcasm detection. The model leverages the Align Before Fuse (ALBEF) framework to effectively align and integrate textual and visual inputs, enhancing the detection of sarcasm across multiple modalities. The performance of the proposed model has been evaluated using a popular benchmarking dataset on multimodal sarcasm detection and achieved an overall accuracy of 87.95% and an F-score of 86.52% that outperforms other state-of-the art methods.