Recent advancements in producing visually convincing Deepfake images have precipitated concerns regarding their potential misuse. To address these concerns, developing robust detection mechanisms has become a priority in recent times. However, the susceptibility of these Deepfake detectors to adversarial attacks continues to pose a substantial challenge, necessitating in-depth exploration. We present a novel framework called the Latent Diffusion-based Momentum Iterative Fast Gradient Sign Method (LDMI-FGSM) that generates authentic and effective adversarial samples. The LDMI-FGSM method uses a latent diffusion model to steer the momentum gradient method, keeping the generated attacks within the proximity of the initial data distribution. This method ensures that the generated attacks remain covert while guaranteeing their effectiveness. Our experiments demonstrate that this novel adversarial attack significantly reduces the accuracy of five state-of-the-art (SOTA) Deepfake detectors and demonstrates excellent transferability and resistance to purification. Even detection models enhanced for improved defense are vulnerable to our transfer black-box attacks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing the Transferability and Stealth of Deepfake Detection Attacks Through Latent Diffusion Models

  • Yu Zhang,
  • Shoukun Xu,
  • Huajun Zhang

摘要

Recent advancements in producing visually convincing Deepfake images have precipitated concerns regarding their potential misuse. To address these concerns, developing robust detection mechanisms has become a priority in recent times. However, the susceptibility of these Deepfake detectors to adversarial attacks continues to pose a substantial challenge, necessitating in-depth exploration. We present a novel framework called the Latent Diffusion-based Momentum Iterative Fast Gradient Sign Method (LDMI-FGSM) that generates authentic and effective adversarial samples. The LDMI-FGSM method uses a latent diffusion model to steer the momentum gradient method, keeping the generated attacks within the proximity of the initial data distribution. This method ensures that the generated attacks remain covert while guaranteeing their effectiveness. Our experiments demonstrate that this novel adversarial attack significantly reduces the accuracy of five state-of-the-art (SOTA) Deepfake detectors and demonstrates excellent transferability and resistance to purification. Even detection models enhanced for improved defense are vulnerable to our transfer black-box attacks.