Towards a Metric to Assess Neural Network Resilience Against Adversarial Samples
摘要
Neural networks are vulnerable to adversarial attacks. Existing robustness evaluation methods have notable limitations, which makes robustness assessment challenging. This work explores robustness evaluation techniques and identifies key factors, including distance metrics, loss functions, attack generation algorithms, attacker models, specificity, and computational resources. Building on those factors, a novel robustness metric for classification tasks is proposed. Our metric accounts for both, targeted and untargeted attacks across three attacker models, while incorporating accuracy and loss into a weighted aggregation. The scoring includes robustness-versus-perturbation and loss-versus-perturbation curves. Our robustness metric offers a more reliable evaluation and deeper insights into model vulnerability compared to previous approaches.