<p>The proliferation of artificial intelligence (AI)-generated content on social media has escalated concerns regarding misinformation, privacy, and security. Current deepfake detection methods struggle to generalize effectively, failing to keep pace with evolving generative models like generative adversarial networks (GANs) and diffusion models. Inspired by the robustness of “real” and “fake” descriptions in the language world, in this paper, we introduce the language-enhanced deepfake detection network (LEDNet), a multimodal foundation model that leverages both linguistic and visual information to improve detection accuracy and adaptability. We are the first to exploit multi-modal technologies for generalized deepfake detection through an un-trained-backbone mechanism, which utilizes pre-trained vision-language models without the need for additional retraining. The model is powered by two key components: language-guided knowledge aggregation (LKA), which assembles a linguistic knowledge base of real and fake indicators, and attention-based vision-language mutualism (AVM), which aligns multimodal features to enhance detection precision. Extensive experiments across 25 popular deepfake datasets demonstrate that LEDNet outperforms state-of-the-art methods, achieving superior accuracy and generalization across diverse deepfake generation techniques. This work sets a new benchmark in multimodal deepfake detection, advancing the application of foundation models in this domain.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LEDNet: a multimodal foundation model for robust deepfake detection

  • Renshuai Tao,
  • Shijie Tang,
  • Haotong Qin,
  • Wei Wang,
  • Yunchao Wei,
  • Yao Zhao

摘要

The proliferation of artificial intelligence (AI)-generated content on social media has escalated concerns regarding misinformation, privacy, and security. Current deepfake detection methods struggle to generalize effectively, failing to keep pace with evolving generative models like generative adversarial networks (GANs) and diffusion models. Inspired by the robustness of “real” and “fake” descriptions in the language world, in this paper, we introduce the language-enhanced deepfake detection network (LEDNet), a multimodal foundation model that leverages both linguistic and visual information to improve detection accuracy and adaptability. We are the first to exploit multi-modal technologies for generalized deepfake detection through an un-trained-backbone mechanism, which utilizes pre-trained vision-language models without the need for additional retraining. The model is powered by two key components: language-guided knowledge aggregation (LKA), which assembles a linguistic knowledge base of real and fake indicators, and attention-based vision-language mutualism (AVM), which aligns multimodal features to enhance detection precision. Extensive experiments across 25 popular deepfake datasets demonstrate that LEDNet outperforms state-of-the-art methods, achieving superior accuracy and generalization across diverse deepfake generation techniques. This work sets a new benchmark in multimodal deepfake detection, advancing the application of foundation models in this domain.