Mitigating the dissemination of misinformation requires effective fake news detection. Existing models rely on simple global features and shallow language models for visual and textual feature extraction, making it difficult to capture deep semantic relationships. Furthermore, traditional models lack detailed fusion mechanisms, leading to inefficient filtering of redundant information and affecting overall performance. This paper proposes a multimodal fake news detection model, MSFND-Net, which combines advanced feature extraction and semantic relationship modeling techniques. MSFND-Net includes four components: feature extraction, internal semantic relationship modeling, cross-modal semantic relationship modeling, and classification. Visual features are obtained through Faster-RCNN combined with a pre-trained ResNet-101 model, whereas textual features are derived from the BERT model. The internal semantic relationship modeling module constructs semantic relationship graphs for visual regions and text words, using a graph attention network to capture these relationships. The cross-modal semantic relationship modeling module uses a two-layer graph attention network to process image and text features, forming a combined feature graph, which is further processed by a cross-modal graph neural network to capture complex cross-modal semantic relationships. The feature fusion attention module calculates the cosine similarity between textual and visual features, normalizes it to obtain cross-modal similarity scores, and adaptively reweights these features. Experimental results show that MSFND-Net performs excellently on several well-known fake news detection datasets, with overall accuracy improvements of 4.1%, 1.1%, and 3.8% on the Weibo, Politifact, and Gossipcop datasets, respectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal Semantic Fusion Network for Fake News Detection

  • Jiaqian Liu,
  • Xiaolong Deng

摘要

Mitigating the dissemination of misinformation requires effective fake news detection. Existing models rely on simple global features and shallow language models for visual and textual feature extraction, making it difficult to capture deep semantic relationships. Furthermore, traditional models lack detailed fusion mechanisms, leading to inefficient filtering of redundant information and affecting overall performance. This paper proposes a multimodal fake news detection model, MSFND-Net, which combines advanced feature extraction and semantic relationship modeling techniques. MSFND-Net includes four components: feature extraction, internal semantic relationship modeling, cross-modal semantic relationship modeling, and classification. Visual features are obtained through Faster-RCNN combined with a pre-trained ResNet-101 model, whereas textual features are derived from the BERT model. The internal semantic relationship modeling module constructs semantic relationship graphs for visual regions and text words, using a graph attention network to capture these relationships. The cross-modal semantic relationship modeling module uses a two-layer graph attention network to process image and text features, forming a combined feature graph, which is further processed by a cross-modal graph neural network to capture complex cross-modal semantic relationships. The feature fusion attention module calculates the cosine similarity between textual and visual features, normalizes it to obtain cross-modal similarity scores, and adaptively reweights these features. Experimental results show that MSFND-Net performs excellently on several well-known fake news detection datasets, with overall accuracy improvements of 4.1%, 1.1%, and 3.8% on the Weibo, Politifact, and Gossipcop datasets, respectively.