One of the most prominent security challenges to neural networks are adversarial examples - inputs with often barely perceptible perturbations causing misclassification. In this study, we propose a defense mechanism that uses an autoencoder to restore adversarial examples before classification. That is, the autoencoder purifies input data points from potential adversarial perturbations. The method is titled Autoencoder-based Adversarial Purification (AAP). We demonstrate the effectiveness of AAP on multiple datasets, attack methods, and perturbation levels. While certain limitations exist, this research offers valuable insights and a promising direction for robust defense mechanisms in adversarial deep learning.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Purifying Adversarial Examples Using an Autoencoder

  • Thijs van Weezel,
  • Famke van Ree,
  • Tychon Bos,
  • Patrick Bastiaanssen,
  • Sibylle Hess

摘要

One of the most prominent security challenges to neural networks are adversarial examples - inputs with often barely perceptible perturbations causing misclassification. In this study, we propose a defense mechanism that uses an autoencoder to restore adversarial examples before classification. That is, the autoencoder purifies input data points from potential adversarial perturbations. The method is titled Autoencoder-based Adversarial Purification (AAP). We demonstrate the effectiveness of AAP on multiple datasets, attack methods, and perturbation levels. While certain limitations exist, this research offers valuable insights and a promising direction for robust defense mechanisms in adversarial deep learning.