The vulnerability of deep learning models to adversarial attacks poses a significant challenge in artificial intelligence, particularly in deepfake detection. Adversarial training has emerged as a crucial defense mechanism, highlighting the importance of highly transferable adversarial samples. However, current attack methods often struggle with transferability in black-box scenarios. This research identifies a key limitation: white-box methods tend to overfit surrogate models, reducing their effectiveness against unknown attacks. To address this vulnerability, we propose the iterative “Random Masking Perturbation” method. This method reduces dependence on gradient or loss function behavior by not using full gradient information to create perturbations. Furthermore, by iteratively masking perturbations while still generating successful adversarial samples, our method automatically and selectively perturbs sensitive features that easily mislead the surrogate model, particularly those with the highest density in the low-frequency band, which is critical for achieving high transferability rates. Experimental results demonstrate significant enhancements in transferability rates, with Random Masking yielding improvements of up to 60%, offering promising strides in fortifying the robustness of adversarial training for deepfake detection models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Boosting Black-Box Transferability of Weak Audio Adversarial Attacks with Random Masking

  • Mai Bui,
  • Thien-Phuc Doan,
  • Kihun Hong,
  • Souhwan Jung

摘要

The vulnerability of deep learning models to adversarial attacks poses a significant challenge in artificial intelligence, particularly in deepfake detection. Adversarial training has emerged as a crucial defense mechanism, highlighting the importance of highly transferable adversarial samples. However, current attack methods often struggle with transferability in black-box scenarios. This research identifies a key limitation: white-box methods tend to overfit surrogate models, reducing their effectiveness against unknown attacks. To address this vulnerability, we propose the iterative “Random Masking Perturbation” method. This method reduces dependence on gradient or loss function behavior by not using full gradient information to create perturbations. Furthermore, by iteratively masking perturbations while still generating successful adversarial samples, our method automatically and selectively perturbs sensitive features that easily mislead the surrogate model, particularly those with the highest density in the low-frequency band, which is critical for achieving high transferability rates. Experimental results demonstrate significant enhancements in transferability rates, with Random Masking yielding improvements of up to 60%, offering promising strides in fortifying the robustness of adversarial training for deepfake detection models.