Image Manipulation Localization via Enhanced Cross-Modal Fusion with Edge Supervision
摘要
Image manipulation localization aims to verify image authenticity and precisely segment tampered areas. Recent advances in deep learning have greatly improved detection performance, primarily relying on RGB features or a combination of RGB features with a single type of noise feature for detection. However, this approach often lacks adaptability to complex manipulation types. Additionally, most existing feature fusion methods depend on attention mechanisms for channel-wise weighting, lacking direct cross-modal interaction, which limits the deep fusion capability of multimodal information. To resolve these constraints, our study develops a novel noise extractor that enhances adaptability to various manipulation types by integrating Bayar and NP++ noise features in addition to conventional high-pass filtering. Furthermore, a cross-modal feature correction and fusion module enables more direct and tighter interactions between RGB features and multiple noise features, thereby improving the accuracy and robustness of manipulation detection. Experimental results demonstrate that the proposed method outperforms existing mainstream approaches in manipulation region localization tasks, effectively enhancing detection precision and generalization ability.