<p>Sarcasm is a complex linguistic act in which the literal meaning is opposite to the true attitude. Multimodal sarcasm detection (MSD) aims to identify whether a given multimodal data sample is sarcastic. Although existing methods have achieved impressive success, they ignored sarcastic information in text and potential clues in images and failed to fully explore cross-modal feature representation and inconsistency between modalities. To solve the above problems, we propose a knowledge enhanced and incongruity perceiving network (KEIPN) for MSD. We design a text knowledge pretraining module that includes RoBERTa models enhanced with biased sentiment knowledge and sarcasm knowledge. Next, we develop an image knowledge amalgamation module to integrate different types of knowledge into visual model, enhancing its visual understanding capability. Pretrained language and visual models can extract better feature representations. Then, we propose a cross-modal knowledge distillation module to model the relationship between text and images and achieve adaptive weighted fusion. Finally, we design an incongruity perception module to capture the inconsistency between images and text and weight the loss using keyless attention mechanisms. Extensive experiments on widely used datasets MMSD and MMSD2.0 demonstrate the superiority of our model over state-of-the-art (SOTA) methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Knowledge Enhanced and Incongruity Perceiving Network for Multimodal Sarcasm Detection

  • Mingqi Liu,
  • Zhixin Li

摘要

Sarcasm is a complex linguistic act in which the literal meaning is opposite to the true attitude. Multimodal sarcasm detection (MSD) aims to identify whether a given multimodal data sample is sarcastic. Although existing methods have achieved impressive success, they ignored sarcastic information in text and potential clues in images and failed to fully explore cross-modal feature representation and inconsistency between modalities. To solve the above problems, we propose a knowledge enhanced and incongruity perceiving network (KEIPN) for MSD. We design a text knowledge pretraining module that includes RoBERTa models enhanced with biased sentiment knowledge and sarcasm knowledge. Next, we develop an image knowledge amalgamation module to integrate different types of knowledge into visual model, enhancing its visual understanding capability. Pretrained language and visual models can extract better feature representations. Then, we propose a cross-modal knowledge distillation module to model the relationship between text and images and achieve adaptive weighted fusion. Finally, we design an incongruity perception module to capture the inconsistency between images and text and weight the loss using keyless attention mechanisms. Extensive experiments on widely used datasets MMSD and MMSD2.0 demonstrate the superiority of our model over state-of-the-art (SOTA) methods.