Personalized Prompt Attack Strategies Against Large Language Models
摘要
With the widespread application of large language models (LLMs) in open ended dialogue and content generation, the interplay between security mechanisms and adversarial attack strategies has become a central issue in AI safety. This study investigates the security risks of LLMs and proposes seven personalized prompt attack paradigms (PPAP), including three-level jailbreak, prompt induction, feedback misdirection, rule substitution, text continuation, directed disguise, and iterative many-shot jailbreaking (IMSJ) attacks, along with corresponding countermeasures tailored to each attack type. Experimental results demonstrate that PPAP achieves a significantly higher average attack success rate of 16.08%, compared to conventional attack methods. These findings reveal that LLMs remain vulnerable in specific scenarios, as customized attacks are more effective than general ones, thereby exposing the fragility of model security boundaries and vulnerabilities arising from model dependency, and offering critical insights into the development of robust and trustworthy AI systems.