<p>Metaphors are ubiquitous in natural language, and metaphor detection, as an important prerequisite for metaphor understanding, is widely used in natural language processing tasks such as sentiment analysis, sarcasm interpretation, and text comprehension. Current metaphor detection methods rely mainly on text and identify metaphorical language through language analysis. However, these methods usually focus too much on text content, ignore the importance of visual metaphors, and lack effective multimodal metaphor feature integration methods. This paper proposes a metaphor detection model with visual information enhancement based on multimodal split fusion. Specifically, we first use a multidimensional attention enhancement module to process image information. This module optimizes the recognition and processing of key features by sequentially integrating channel and spatial attention mechanisms, thereby improving the performance of the model in visual tasks. To achieve two-way interaction of multimodal metaphor features, we design a multimodal split-fusion module. This module enhances the model’s metaphor detection ability by dividing each modal data into feature blocks of equal size and aggregating and weighting these blocks. Extensive experimental results on the public multimodal metaphor dataset METMeme and the sarcasm dataset Sarcasm verify the effectiveness of our model.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SFVE: visual information enhancement metaphor detection with multimodal splitting fusion

  • Qimeng Yang,
  • Hao Meng,
  • Yuanbo Yan,
  • Shisong Guo,
  • Qixing Wei

摘要

Metaphors are ubiquitous in natural language, and metaphor detection, as an important prerequisite for metaphor understanding, is widely used in natural language processing tasks such as sentiment analysis, sarcasm interpretation, and text comprehension. Current metaphor detection methods rely mainly on text and identify metaphorical language through language analysis. However, these methods usually focus too much on text content, ignore the importance of visual metaphors, and lack effective multimodal metaphor feature integration methods. This paper proposes a metaphor detection model with visual information enhancement based on multimodal split fusion. Specifically, we first use a multidimensional attention enhancement module to process image information. This module optimizes the recognition and processing of key features by sequentially integrating channel and spatial attention mechanisms, thereby improving the performance of the model in visual tasks. To achieve two-way interaction of multimodal metaphor features, we design a multimodal split-fusion module. This module enhances the model’s metaphor detection ability by dividing each modal data into feature blocks of equal size and aggregating and weighting these blocks. Extensive experimental results on the public multimodal metaphor dataset METMeme and the sarcasm dataset Sarcasm verify the effectiveness of our model.