Detect Text Forgery with Non-forged Image Features: A Framework for Detection and Grounding of Image-Text Manipulation
摘要
With the rapid development of generative models, multimodal fake media has proliferated across the Internet. Detecting and Grounding forgery images and text is crucial in advancing cybersecurity. Most existing approaches utilize image-text inconsistency to detect and ground the multi-modal forgery. However, simultaneous manipulations in visual and textual modalities may still maintain the consistency between forged images and forged text, making detecting forgery challenging. To address this problem, we divide the task of detecting multi-modal forgery into two sub-tasks: detecting image forgery and detecting text forgery with non-forged image areas. Specifically, we propose a novel progressive reasoning framework that only fuses the feature between text and authentic image area controlled by the Forgery-Aware Feature Gate (FAFG). Additionally, we introduce Multi-Scale Feature Aggregation (MSFA) to enhance image forgery detection by aggregating multi-scale image features. Experimental results demonstrate that our method outperforms previous state-of-the-art methods with even fewer training epochs.