Scene Graph-Based Semantic Enhancement for Multimodal Fake News Detection
摘要
Multimodal fake news has caused significant harm to economic and political systems, thereby emerging as a serious societal issue. Current methodologies predominantly concentrate on modeling superficial image features (e.g., texture and color characteristics). Although certain investigations have attempted to enhance detection performance by extracting deep semantic information through identification of entity objects within images, existing approaches still demonstrate insufficient representation learning of spatial positional relationships between entities. This limitation consequently restricts their capacity to effectively characterize complex semantic associations. Therefore, we propose a scene graph-based semantic enhancement framework (SGSE), which extracts relationships between entities to help the model deeply understand semantic information in images. Utilizing fine-grained information obtained through scene graphs, we employ a fusion mechanism based on text filtering to achieve relation-level alignment between images and text. Furthermore, we employ a cross-granularity integration module, which contains coarse-grained extraction and multi-granularity fusion, to ensure a comprehensive representation of fake news. Experiments conducted on three benchmark datasets consistently demonstrate the superiority of SGSE, achieving state-of-the-art performance.