With the advent of Deep Learning (DL), speech synthesis technologies have made remarkable progress, enabling the creation of highly realistic human voices. Although this technology offers numerous benefits, it also introduces substantial security risks, particularly through “Deepfake” speech attacks. These attacks pose severe threats to personal security and societal trust. Existing defense mechanisms, primarily focused on post-attack detection, fall short when confronted with the advanced sophisticate speech synthesis techniques. Worse still, these defenses are often insufficient because significant harm may have already been inflicted by the time deepfake audio is identified. In this paper, we present a novel proactive defense method against unauthorized speech synthesis named “Defend from Scratch” (DFS). By leveraging the idea of adversarial examples, our method can proactively hinder the creation of deepfake speeches. To do that, we propose to incorporate a pretrained decoupled denoising diffusion model (DDDM) to introduce robust and imperceptible adversarial perturbations into the source audio, which enables effectively defense against adaptive attacks without significant audio quality downgrade. Furthermore, Enhanced defenses have been implemented through the application of ensemble learning, which extend its capability to counter a diverse range of threats, thereby ensuring robust voice privacy protection without compromising the integrity or usability of the original audio. Experiments show that our approach stands out for its efficacy in maintaining a high standard of voice privacy in the face of emerging Deepfake technologies.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Defend from Scratch: A Diffusion-Based Proactive Defense Method for Unauthorized Speech Synthesis

  • Ying Wang,
  • Yuchuan Luo,
  • Zhenyu Qiu,
  • Lin Liu,
  • Shaojing Fu

摘要

With the advent of Deep Learning (DL), speech synthesis technologies have made remarkable progress, enabling the creation of highly realistic human voices. Although this technology offers numerous benefits, it also introduces substantial security risks, particularly through “Deepfake” speech attacks. These attacks pose severe threats to personal security and societal trust. Existing defense mechanisms, primarily focused on post-attack detection, fall short when confronted with the advanced sophisticate speech synthesis techniques. Worse still, these defenses are often insufficient because significant harm may have already been inflicted by the time deepfake audio is identified. In this paper, we present a novel proactive defense method against unauthorized speech synthesis named “Defend from Scratch” (DFS). By leveraging the idea of adversarial examples, our method can proactively hinder the creation of deepfake speeches. To do that, we propose to incorporate a pretrained decoupled denoising diffusion model (DDDM) to introduce robust and imperceptible adversarial perturbations into the source audio, which enables effectively defense against adaptive attacks without significant audio quality downgrade. Furthermore, Enhanced defenses have been implemented through the application of ensemble learning, which extend its capability to counter a diverse range of threats, thereby ensuring robust voice privacy protection without compromising the integrity or usability of the original audio. Experiments show that our approach stands out for its efficacy in maintaining a high standard of voice privacy in the face of emerging Deepfake technologies.