Multi-scale query-based transformer for image forgery localization
摘要
Deep learning has achieved remarkable results in the field of Image Forgery Localization (IFL), but the existing methods pay insufficient attention to the training efficiency and the spatial structure and size of the tampered region. To address this limitation, we propose the Multi-scale Query-based Transformer (MQFormer) model, that employs Ground-truth Mask Token (GMT) to facilitate the identification of forged regions using Image Feature Token (IFT) and Learnable Query Token (LQT). In particular, IFT are initially extracted through the utilization of a dual-stream encoder (comprising RGB and Noise branches), while the feature embedding of the ground-truth mask is then employed as GMT. The key contribution of MQFormer is the design of a multi-scale query transformer module, which leverage attention mechanism to update IFT and LQT through GMT. The module continuously optimizes the localization accuracy from coarse to fine grains through gradual refinement, thus enhancing the model’s ability to identify forged regions of different sizes. Furthermore, we propose a novel contrastive loss function, which aims to reduce the distance between GMT and LQT in the feature space, thereby enhancing the guiding role of mask information. Extensive experimentation on multiple benchmarks has demonstrated that MQFormer markedly outperforms existing state-of-the-art methodologies and displays adaptability to diverse tampering types and robust resilience to attacks.