Joint spatial-frequency deepfake detection network based on dual-domain attention-enhanced deformable convolution
摘要
This paper proposes a novel Spatio-Frequency Deepfake Detection Network based on Dual-Domain Attention and Deformable Convolution (HFDCDNet). The core of the network is the HFDCD module, which simultaneously extracts spatial domain features and high-frequency information from the input feature maps. These two types of features are then fused using a specially designed bidirectional cross-attention mechanism. Specifically, the spatial features are extracted using a Deformable Convolution and Dual Attention-based module (DCD), which leverages deformable convolutions (DCN) and a spatial-channel attention mechanism to capture more comprehensive spatial representations. Meanwhile, high-frequency features are extracted using the High-Frequency Extraction (HFE) module, which learns frequency domain representations through high-pass filtering and frequency-aware convolutional learning. The bidirectional cross-attention mechanism facilitates complementary fusion and mutual enhancement of spatial and frequency features, enabling more fine-grained and holistic feature learning and improving detection performance. To evaluate the effectiveness of the proposed HFDCDNet, extensive experiments were conducted on two widely used public datasets: FaceForensics++ and Celeb-DF (V2). The results demonstrate that HFDCDNet achieves an accuracy (ACC) of 98.31% and an area under the curve (AUC) of 99.51% on the FaceForensics++ dataset, and 98.29% ACC and 99.13% AUC on the Celeb-DF (V2) dataset, outperforming many state-of-the-art methods. These results confirm that the proposed DCD module significantly enhances the model’s ability to detect manipulated facial content.