<p>Large Language Models (LLMs) are becoming vulnerable to adversarial attacks due to the creation crafted inputs that may deceive even the most advanced protections. These models can produce the wrong prediction with slight changes in the input data. This flaw can be noticed in many areas, such as computer vision, speech recognition, and natural language processing, and there are serious doubts about the robustness of AraBERT LLMs in sensitive use. Consequently, this paper assesses the robustness of AraBERT LLMs to character-level and word-level black-box attacks that cause spelling errors in the input data. The adversarial samples were generated using the Chain-of-Thought prompting method, to evaluate the robustness of a fine-tuned AraBERT model and the original AarBERT v2.0 LLM. A considerable reduction in accuracy was found, especially with regard to word-level delete attacks, and the largest reduction of 44.93% of the original model was obtained. The deletions at the word and character levels were also significant. This highlights how important it is to strengthen the model’s robustness to deletions, in order to obtain reliable performance for a variety of text lengths and challenging crafted input.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Robustness of AraBERT models to black-box adversarial deletions and perturbations

  • Malik Qasaimeh,
  • Saja Nakhleh,
  • Raad Al-Qassas,
  • Ammar Abdallah,
  • Muneer Bani Yasin

摘要

Large Language Models (LLMs) are becoming vulnerable to adversarial attacks due to the creation crafted inputs that may deceive even the most advanced protections. These models can produce the wrong prediction with slight changes in the input data. This flaw can be noticed in many areas, such as computer vision, speech recognition, and natural language processing, and there are serious doubts about the robustness of AraBERT LLMs in sensitive use. Consequently, this paper assesses the robustness of AraBERT LLMs to character-level and word-level black-box attacks that cause spelling errors in the input data. The adversarial samples were generated using the Chain-of-Thought prompting method, to evaluate the robustness of a fine-tuned AraBERT model and the original AarBERT v2.0 LLM. A considerable reduction in accuracy was found, especially with regard to word-level delete attacks, and the largest reduction of 44.93% of the original model was obtained. The deletions at the word and character levels were also significant. This highlights how important it is to strengthen the model’s robustness to deletions, in order to obtain reliable performance for a variety of text lengths and challenging crafted input.