Multimodal Stance Detection aims to classify public opinions on specific targets in social media, incorporating both text and image data. However, prior studies have overemphasized the significance of images, neglecting the presence of irrelevant images in the dataset. Moreover, previous research has shown that employing the Chain of Thought approach with large language models can introduce noise from the generated text as well as noise from the image modality into the text input modality. These noises  can degrade the performance of multimodal models. Additionally, both image and text modalities exhibit complex data patterns, resulting in significant disparities in training difficulty across the dataset. To address these issues, We proposed Target-Oriented Dynamic Denoising Curriculum Learning(TODDCL), which effectively measures different types of noise and automatically tackles the noise in both image and text modalities based on their relevance to the target, in an escalating complexity order. Experimental results on five benchmark datasets demonstrate that our proposed TODDCL method achieves state-of-the-art performance in multimodal stance detection.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Target-Oriented Dynamic Denosing Curriculum Learning for Multimodel Stance Detection

  • Zihao Suo,
  • Shanliang Pan

摘要

Multimodal Stance Detection aims to classify public opinions on specific targets in social media, incorporating both text and image data. However, prior studies have overemphasized the significance of images, neglecting the presence of irrelevant images in the dataset. Moreover, previous research has shown that employing the Chain of Thought approach with large language models can introduce noise from the generated text as well as noise from the image modality into the text input modality. These noises  can degrade the performance of multimodal models. Additionally, both image and text modalities exhibit complex data patterns, resulting in significant disparities in training difficulty across the dataset. To address these issues, We proposed Target-Oriented Dynamic Denoising Curriculum Learning(TODDCL), which effectively measures different types of noise and automatically tackles the noise in both image and text modalities based on their relevance to the target, in an escalating complexity order. Experimental results on five benchmark datasets demonstrate that our proposed TODDCL method achieves state-of-the-art performance in multimodal stance detection.