Backdoor threats in large language models—a survey
摘要
Large language models (LLMs), with their advanced language comprehension and text generation capabilities, have demonstrated remarkable performance across diverse application scenarios involving code processing, search engines, and translation, among others. However, these models have become increasingly vulnerable to security threats, particularly to backdoor attacks. Therefore, a timely and comprehensive review of the existing backdoor threats is urgently required. In this paper, we present a systematic and timely review of the research on backdoor attacks on LLMs, categorising existing attack and defence methods according to the LLM. Additionally, we draw comparisons with backdoor attacks in traditional deep learning to provide a more intuitive understanding of backdoor threats in LLMs. Through this effective analysis and an evaluation of the reviewed studies, we identify the current research challenges and propose potential future research directions to address these issues.