<p>With the rapid development of Deepfake and generative AI technologies, accurately distinguishing real from synthetic images has become a key challenge in digital security and media forensics. Existing CNN-based methods are limited by their local receptive fields, making it difficult to capture subtle forgery traces distributed across the global spatial domain. Moreover, existing methods cannot model the continuous dynamic feature evolution intrinsic to image generation processes. To address these issues, this paper proposes ODE-SAT Net, a synthetic image detection method that integrates Neural Ordinary Differential Equations (Neural ODEs) with a Spatial Attention Transformer Module (SATM). First, to highlight subtle artifacts masked by low-frequency content, we design a high-frequency residual preprocessing module that amplifies high-frequency artifacts introduced by generative models via cascaded downsampling and upsampling operations. Second, to overcome the limited receptive field of CNNs, we introduce a spatial self-attention Transformer that models long-range semantic dependencies among pixels to precisely localize forgery regions dispersed throughout the image. Finally, we introduce a Neural ODE module to model the continuous-depth evolution of intermediate feature representations by transforming discrete features into a continuous dynamical system. However, the iterative numerical solving of Neural ODEs and the O(N<sup>2</sup>) complexity of self-attention introduce significant computational bottlenecks. Experimental results on a dataset covering 16 generative models demonstrate that our method achieves state-of-the-art performance in both detection accuracy and generalization capability. Nevertheless, in large-scale high-resolution image detection scenarios, the substantial computational and memory requirements far exceed the capacity of single-machine architectures, making high-performance computing or parallel distributed architectures essential for real-time detection. Importantly, the modular structure of our method provides a solid foundation for such parallel implementation on supercomputing platforms.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Synthetic image detection method integrating neural ordinary differential equations and spatial attention transformer

  • Tao Yang,
  • Qian Zhang,
  • Bin Hu,
  • Haiyan Jin,
  • Lian Zhu

摘要

With the rapid development of Deepfake and generative AI technologies, accurately distinguishing real from synthetic images has become a key challenge in digital security and media forensics. Existing CNN-based methods are limited by their local receptive fields, making it difficult to capture subtle forgery traces distributed across the global spatial domain. Moreover, existing methods cannot model the continuous dynamic feature evolution intrinsic to image generation processes. To address these issues, this paper proposes ODE-SAT Net, a synthetic image detection method that integrates Neural Ordinary Differential Equations (Neural ODEs) with a Spatial Attention Transformer Module (SATM). First, to highlight subtle artifacts masked by low-frequency content, we design a high-frequency residual preprocessing module that amplifies high-frequency artifacts introduced by generative models via cascaded downsampling and upsampling operations. Second, to overcome the limited receptive field of CNNs, we introduce a spatial self-attention Transformer that models long-range semantic dependencies among pixels to precisely localize forgery regions dispersed throughout the image. Finally, we introduce a Neural ODE module to model the continuous-depth evolution of intermediate feature representations by transforming discrete features into a continuous dynamical system. However, the iterative numerical solving of Neural ODEs and the O(N2) complexity of self-attention introduce significant computational bottlenecks. Experimental results on a dataset covering 16 generative models demonstrate that our method achieves state-of-the-art performance in both detection accuracy and generalization capability. Nevertheless, in large-scale high-resolution image detection scenarios, the substantial computational and memory requirements far exceed the capacity of single-machine architectures, making high-performance computing or parallel distributed architectures essential for real-time detection. Importantly, the modular structure of our method provides a solid foundation for such parallel implementation on supercomputing platforms.