<p>Neural Code Models (NCMs) have shown strong performance across a wide range of code understanding tasks. However, recent studies reveal that such models are vulnerable to backdoor attacks. Backdoored NCMs behave normally on clean inputs, but produce attacker-expected outputs on samples injected with backdoor triggers, posing a covert, yet critical security threat. Existing backdoor defense techniques employ trigger inversion and trigger unlearning to mitigate such attacks. However, they often struggle to recover accurate triggers and to sufficiently eliminate backdoor behaviors due to reliance on coarse-grained heuristics and restricted unlearning strategies, leaving models still sensitive to inputs containing triggers. To address these issues, we propose BADERASER, a novel backdoor defense technique for backdoor elimination in NCMs. BADERASER introduces code naturalness as an auxiliary constraint and incorporates statistical indicators in trigger inversion to improve the quality of recovered triggers. Furthermore, we propose a distillation-based unlearning method to purify backdoor features while preserving clean knowledge. We evaluate BADERASER on multiple code understanding tasks, model architectures, and advanced backdoor attacks. Experimental results demonstrate that BADERASER outperforms state-of-the-art defenses in reducing the attack success rate across all evaluated attack scenarios, achieving an average attack success rate decrease of 54.37% over ONION, 62.50% over DBS, and 20.75% over ELIBADCODE.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Defending neural code understanding models by eliminating backdoors

  • Wei Cheng,
  • Yu Zhou,
  • Guang Yang,
  • Xiangyu Zhang,
  • Wenhua Yang,
  • Taolue Chen

摘要

Neural Code Models (NCMs) have shown strong performance across a wide range of code understanding tasks. However, recent studies reveal that such models are vulnerable to backdoor attacks. Backdoored NCMs behave normally on clean inputs, but produce attacker-expected outputs on samples injected with backdoor triggers, posing a covert, yet critical security threat. Existing backdoor defense techniques employ trigger inversion and trigger unlearning to mitigate such attacks. However, they often struggle to recover accurate triggers and to sufficiently eliminate backdoor behaviors due to reliance on coarse-grained heuristics and restricted unlearning strategies, leaving models still sensitive to inputs containing triggers. To address these issues, we propose BADERASER, a novel backdoor defense technique for backdoor elimination in NCMs. BADERASER introduces code naturalness as an auxiliary constraint and incorporates statistical indicators in trigger inversion to improve the quality of recovered triggers. Furthermore, we propose a distillation-based unlearning method to purify backdoor features while preserving clean knowledge. We evaluate BADERASER on multiple code understanding tasks, model architectures, and advanced backdoor attacks. Experimental results demonstrate that BADERASER outperforms state-of-the-art defenses in reducing the attack success rate across all evaluated attack scenarios, achieving an average attack success rate decrease of 54.37% over ONION, 62.50% over DBS, and 20.75% over ELIBADCODE.