<p>Natural language processing models are widely acknowledged for their strong data fitting capabilities, diverse application scenarios, and adaptable learning methodologies. However, these models, including the large language models, exhibit sensitivity to adversarial example attacks. These examples are slightly perturbed from the pristine text but mislead the model classification. Nevertheless, the existing attack methods primarily focus on the attack effectiveness without semantics-preservation considered. Moreover, the trade-off between evasion effectiveness and concealment of perturbed texts is less investigated. In this study, we propose a multi-objective adversarial text generation framework (MOATG) that simultaneously optimizes attack success rate, imperceptibility, and semantic similarity. Tailored objective functions and dominance relations are designed for character-, word-, and sentence-level perturbations. MOATG is evaluated against five baselines across five benchmark datasets. Experimental results show that MOATG achieves a 5.38% average improvement in attack success rate and reduces word error rate by 1.48%, demonstrating its effectiveness in balancing attack strength and stealth.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Towards optimal adversarial texts: character, word, and sentence

  • Pengchuan Wang,
  • Deqiang Li,
  • Qianmu Li

摘要

Natural language processing models are widely acknowledged for their strong data fitting capabilities, diverse application scenarios, and adaptable learning methodologies. However, these models, including the large language models, exhibit sensitivity to adversarial example attacks. These examples are slightly perturbed from the pristine text but mislead the model classification. Nevertheless, the existing attack methods primarily focus on the attack effectiveness without semantics-preservation considered. Moreover, the trade-off between evasion effectiveness and concealment of perturbed texts is less investigated. In this study, we propose a multi-objective adversarial text generation framework (MOATG) that simultaneously optimizes attack success rate, imperceptibility, and semantic similarity. Tailored objective functions and dominance relations are designed for character-, word-, and sentence-level perturbations. MOATG is evaluated against five baselines across five benchmark datasets. Experimental results show that MOATG achieves a 5.38% average improvement in attack success rate and reduces word error rate by 1.48%, demonstrating its effectiveness in balancing attack strength and stealth.