Multimodal sarcasm detection is of great significance for understanding users’ intentions and emotions. Existing methods suffer from weaknesses in visual modality representation, insufficient granularity in cross-modal fusion, and a lack of semantic reasoning knowledge. This paper proposes a multimodal sarcasm detection model based on cross-modal fine-grained fusion and reasoning chain enhancement, conducting research from three dimensions: modal modeling, knowledge-enhanced reasoning, and semantic fusion. (1) We utilize multimodal large language models to perform deep semantic parsing of visual modality images; (2) based on the long-chain reasoning technology of large language models, we achieve knowledge enhancement under the constraints of semantic and logical consistency; (3) we design an image region-text token-level association mapping to achieve more detailed and accurate cross-modal information fusion. Experimental results on the HFM and SarcNet datasets show that the method proposed in this paper has significant advantages over baseline methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal Sarcasm Detection Based on Cross-Modal Fine-Grained Fusion and Reasoning Chain Enhancement

  • Jian Liao,
  • Yujin Zheng,
  • Suge Wang,
  • Jianxing Zheng,
  • Xin Guo,
  • Jianjun Li

摘要

Multimodal sarcasm detection is of great significance for understanding users’ intentions and emotions. Existing methods suffer from weaknesses in visual modality representation, insufficient granularity in cross-modal fusion, and a lack of semantic reasoning knowledge. This paper proposes a multimodal sarcasm detection model based on cross-modal fine-grained fusion and reasoning chain enhancement, conducting research from three dimensions: modal modeling, knowledge-enhanced reasoning, and semantic fusion. (1) We utilize multimodal large language models to perform deep semantic parsing of visual modality images; (2) based on the long-chain reasoning technology of large language models, we achieve knowledge enhancement under the constraints of semantic and logical consistency; (3) we design an image region-text token-level association mapping to achieve more detailed and accurate cross-modal information fusion. Experimental results on the HFM and SarcNet datasets show that the method proposed in this paper has significant advantages over baseline methods.