Backdoor attacks pose significant threats to Natural Language Processing (NLP) models. Various backdoor defense methods for NLP models primarily function by identifying and subsequently manipulating backdoor triggers within provided samples. However, such methods predominantly operate at the level of data filtering, essentially failing to cleanse the affected model. To solve this problem, we present ELUDE—a groundbreaking method designed to excise the backdoor triggers embedded within the corrupted model. ELUDE’s architecture comprises two core components: the backdoor trigger identifier and the backdoor trigger remover, operating synergistically in a pipeline procedure. While the former employs a perplexity-based approach to locate the backdoor trigger, the latter eradicates the inserted backdoor’s influence on the tainted model using machine unlearning. To counteract the issue of catastrophic forgetting engendered by machine unlearning, we incorporate Elastic Weight Consolidation (EWC) within the backdoor trigger remover. Our experiments on SST-2, OLID, and AG News text classification datasets exemplify the efficacy of ELUDE, as comparative results indicate that ELUDE effectively reduces the success rate of three cutting-edge backdoor attack methods by an average of 60%—simultaneously maintaining comparable performance on the original task.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Defense Against Textual Backdoors via Elastic Weighted Consolidation-Based Machine Unlearning

  • Haojun Xuan,
  • Yajie Wang,
  • Huishu Wu,
  • Tao Liu,
  • Chuan Zhang,
  • Liehuang Zhu

摘要

Backdoor attacks pose significant threats to Natural Language Processing (NLP) models. Various backdoor defense methods for NLP models primarily function by identifying and subsequently manipulating backdoor triggers within provided samples. However, such methods predominantly operate at the level of data filtering, essentially failing to cleanse the affected model. To solve this problem, we present ELUDE—a groundbreaking method designed to excise the backdoor triggers embedded within the corrupted model. ELUDE’s architecture comprises two core components: the backdoor trigger identifier and the backdoor trigger remover, operating synergistically in a pipeline procedure. While the former employs a perplexity-based approach to locate the backdoor trigger, the latter eradicates the inserted backdoor’s influence on the tainted model using machine unlearning. To counteract the issue of catastrophic forgetting engendered by machine unlearning, we incorporate Elastic Weight Consolidation (EWC) within the backdoor trigger remover. Our experiments on SST-2, OLID, and AG News text classification datasets exemplify the efficacy of ELUDE, as comparative results indicate that ELUDE effectively reduces the success rate of three cutting-edge backdoor attack methods by an average of 60%—simultaneously maintaining comparable performance on the original task.