<p>The concept of adversarial machine learning has become a central research area within the domain of artificial intelligence, focusing on how machine learning systems can be exploited for malicious purposes. This is particularly relevant in computer vision, where adversarial attacks pose significant security risks in areas such as autonomous vehicle and weapon detection. These attacks aim to modify correctly classified images with minor modifications to convert them into adversarial instances to deceive machine learning models. This paper introduces a novel white-box attack based on an optimization process with the objective to minimize the number of altered pixels during the creation of adversarial images. To this end, we design a novel algorithm called the Minimal <InlineEquation ID="IEq4"> <EquationSource Format="TEX">\(L_0\)</EquationSource> </InlineEquation> attack (MLA) efficient in 100% of the tested cases, significantly reducing the quantity of altered pixels and computation time during adversarial image creation. Our methodology is benchmarked against existing JSMA and CW-<InlineEquation ID="IEq5"> <EquationSource Format="TEX">\(L_0\)</EquationSource> </InlineEquation> attacks on MNIST, CIFAR10 and GTSRB datasets demonstrating superior performance in both efficiency and computational speed. To successfully deceive models trained on the CIFAR-10 dataset, only 3.26 pixels are necessary, whereas for models trained on the GTSRB dataset, a slightly higher number of 5.03 pixels are required.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Minimal adversarial attack on the \(L_0\) norm

  • Pierre-Francois Maillard,
  • Bimal Kumar Roy

摘要

The concept of adversarial machine learning has become a central research area within the domain of artificial intelligence, focusing on how machine learning systems can be exploited for malicious purposes. This is particularly relevant in computer vision, where adversarial attacks pose significant security risks in areas such as autonomous vehicle and weapon detection. These attacks aim to modify correctly classified images with minor modifications to convert them into adversarial instances to deceive machine learning models. This paper introduces a novel white-box attack based on an optimization process with the objective to minimize the number of altered pixels during the creation of adversarial images. To this end, we design a novel algorithm called the Minimal \(L_0\) attack (MLA) efficient in 100% of the tested cases, significantly reducing the quantity of altered pixels and computation time during adversarial image creation. Our methodology is benchmarked against existing JSMA and CW- \(L_0\) attacks on MNIST, CIFAR10 and GTSRB datasets demonstrating superior performance in both efficiency and computational speed. To successfully deceive models trained on the CIFAR-10 dataset, only 3.26 pixels are necessary, whereas for models trained on the GTSRB dataset, a slightly higher number of 5.03 pixels are required.