<p>Sarcasm is a complex linguistic phenomenon in which the literal meaning of words contrasts with the intended message. Detecting sarcasm in multimodal data, especially combining text and images, is a significant challenge due to the inconsistency between the modalities. Current multimodal sarcasm detection systems often rely on fixed interaction methods and simple feature fusion, which may neglect crucial relationships between image regions and textual words, and the models lack flexibility. To address these challenges, we propose the Multi-Granularity Dynamic Interaction Network (MDIN), which introduces two novel components: the Dynamic Routing Interaction Network (DRIN) and the Multimodal Synergistic Similarity Modeling (MSSM). DRIN utilizes dynamic routing to explore multimodal interactions at multiple granularities, while MSSM enhances feature alignment by modeling semantic and structural relationships across modalities. Experimental results show that MDIN outperforms state-of-the-art models in terms of both accuracy and adaptability, achieving a 0.46% improvement in the F1-score. This demonstrates the model's effectiveness in handling complex multimodal sarcasm detection tasks and suggests its potential applications in social media sentiment analysis and misinformation detection.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-granular dynamic interaction network for multimodal sarcasm detection

  • Jiexia Lin,
  • Xiaodong Zhu

摘要

Sarcasm is a complex linguistic phenomenon in which the literal meaning of words contrasts with the intended message. Detecting sarcasm in multimodal data, especially combining text and images, is a significant challenge due to the inconsistency between the modalities. Current multimodal sarcasm detection systems often rely on fixed interaction methods and simple feature fusion, which may neglect crucial relationships between image regions and textual words, and the models lack flexibility. To address these challenges, we propose the Multi-Granularity Dynamic Interaction Network (MDIN), which introduces two novel components: the Dynamic Routing Interaction Network (DRIN) and the Multimodal Synergistic Similarity Modeling (MSSM). DRIN utilizes dynamic routing to explore multimodal interactions at multiple granularities, while MSSM enhances feature alignment by modeling semantic and structural relationships across modalities. Experimental results show that MDIN outperforms state-of-the-art models in terms of both accuracy and adaptability, achieving a 0.46% improvement in the F1-score. This demonstrates the model's effectiveness in handling complex multimodal sarcasm detection tasks and suggests its potential applications in social media sentiment analysis and misinformation detection.