LEDNet: a multimodal foundation model for robust deepfake detection
摘要
The proliferation of artificial intelligence (AI)-generated content on social media has escalated concerns regarding misinformation, privacy, and security. Current deepfake detection methods struggle to generalize effectively, failing to keep pace with evolving generative models like generative adversarial networks (GANs) and diffusion models. Inspired by the robustness of “real” and “fake” descriptions in the language world, in this paper, we introduce the language-enhanced deepfake detection network (LEDNet), a multimodal foundation model that leverages both linguistic and visual information to improve detection accuracy and adaptability. We are the first to exploit multi-modal technologies for generalized deepfake detection through an un-trained-backbone mechanism, which utilizes pre-trained vision-language models without the need for additional retraining. The model is powered by two key components: language-guided knowledge aggregation (LKA), which assembles a linguistic knowledge base of real and fake indicators, and attention-based vision-language mutualism (AVM), which aligns multimodal features to enhance detection precision. Extensive experiments across 25 popular deepfake datasets demonstrate that LEDNet outperforms state-of-the-art methods, achieving superior accuracy and generalization across diverse deepfake generation techniques. This work sets a new benchmark in multimodal deepfake detection, advancing the application of foundation models in this domain.