The swift proliferation of multimodal rumors on social media, particularly those with manipulated images and complex intermodal interactions, significantly challenges current detection methods. In response, we utilize statistical image features, including mean and variance, to capture spatial attributes effectively and improve the detection of image-tampered tweets. To tackle complex intermodal correlations, we introduce a contrastive learning approach that aligns features across modalities efficiently. Additionally, we introduce a cross-attention fusion module (CAFM) that enhances the integration of image and text modalities, thereby improving multimodal rumor detection performance. In conclusion, we propose the cross-attention fusion network (ConCAFN), leveraging contrastive learning for robust multimodal rumor detection. Extensive experiments on two real-world datasets confirm the model’s enhanced capability to detect multimodal rumors accurately, demonstrating our methods’ effectiveness.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Multimodal Rumor Detection with Statistical Image Features and Modal Alignment via Contrastive Learning

  • Chenyu Zhou,
  • Xiuhong Li,
  • Zhe Li,
  • Fan Chen,
  • Jiabao Sheng,
  • Bin Chen,
  • Haoyu Wang

摘要

The swift proliferation of multimodal rumors on social media, particularly those with manipulated images and complex intermodal interactions, significantly challenges current detection methods. In response, we utilize statistical image features, including mean and variance, to capture spatial attributes effectively and improve the detection of image-tampered tweets. To tackle complex intermodal correlations, we introduce a contrastive learning approach that aligns features across modalities efficiently. Additionally, we introduce a cross-attention fusion module (CAFM) that enhances the integration of image and text modalities, thereby improving multimodal rumor detection performance. In conclusion, we propose the cross-attention fusion network (ConCAFN), leveraging contrastive learning for robust multimodal rumor detection. Extensive experiments on two real-world datasets confirm the model’s enhanced capability to detect multimodal rumors accurately, demonstrating our methods’ effectiveness.