DFADNet: A Diverse-Feature Adaptive Network for Web3.0-Oriented Deep Forgery Detection
摘要
The pervasive spread of deepfakes has underscored an urgent requirement for advanced forgery detection systems, especially in the context of Web 3.0, where user-generated content and decentralized applications are becoming increasingly prevalent. While significant progress has been made in this field through deep learning, the constantly evolving nature of forgery techniques necessitates continual innovation. Current detection methods, constrained by inflexible parameters, struggle to capture the nuances of noise residuals and often overlook the interaction between texture and semantic clues, as well as the temporal subtleties that can significantly enhance detection capabilities. This paper introduces DFADnet, the Diverse Features Adaptive Detection Network, a pioneering approach to deep forgery detection that integrates a diverse range of features for effectively identifying manipulated videos. DFADnet comprises three essential components: TextureNoiseAdapt (TNA) for adaptive extraction of texture noise; SemanticGuidance (SG) for directing focus towards semantic inconsistencies indicative of tampering; and TemporalMultiscale (TMS) for analyzing temporal dynamics across video frames. Our comprehensive experiments on the FF++ and Celeb-DF datasets demonstrate DFADnet’s superior performance, achieving an accuracy of 97.65% in the HQ mode and an AUC of 76.95% on Celeb-DF, thus highlighting its robust generalization capabilities.