Raising the numbers: multi-generation adversarial attack and frequency-based defense for heightened NLP security
摘要
The integration of Artificial Intelligence (AI) into applications of Natural Language Processing (NLP), such as spam detection and sentiment analysis, necessitates robust and explainable defenses against adversarial attacks—subtle input perturbations that can compromise model integrity. In this paper, we propose two novel methods, driven by Explainable AI (XAI), for enhancing the robustness of deep learning models in NLP. Traditional methods for generating adversarial examples are often resource-intensive relative to the number of samples produced, limiting their effectiveness in large-scale adversarial training. To address this problem, we propose, as our first contribution, a Multi-Generation Attack Strategy (