Semantic-agnostic and edge-guided for image manipulation detection and localization
摘要
Image Manipulation Detection and Localization (IMDL) is a crucial task in the field of intelligent image security. The challenges lie in making the model focus on tampering traces while ignoring semantic information and accurately localizing the boundaries of manipulated regions. This paper proposes an end-to-end, generalized network, SAEG-Net, based on semantic agnosticism and edge guidance for IMDL. SAEG-Net utilizes MviT2-B as the backbone network to capture correlations between image patches via the self-attention mechanism. A Semantic Agnostic Module (SAM) is designed to mitigate the impact of rich image semantics on tampering trace extraction through subtraction operations. An Edge Guidance Module (EGM) guides the model in precisely localizing tampered regions using multiplication operations to learn boundary artifacts. Hybrid Attention Feature Fusion (HAFF) integrates these two features, and a Refinement Module refines multi-scale features in a top-down iterative manner to enhance representation. The model is co-supervised using edge loss, detection loss, and segmentation loss. Extensive experiments on five benchmark datasets demonstrate that SAEG-Net outperforms state-of-the-art methods in robustness to attacks and adaptability to different manipulation types, achieving superior performance at both the image-level and pixel-level.