Large Language Models (LLMs) play a crucial role in natural language processing, yet their expanding applications highlight increasing security risks. This paper introduces CS-Eval, a novel benchmark designed to assess LLMs’ capability in addressing safety concerns. CS-Eval focuses on seven key security risks: ethical dilemmas, marginal topics, error discovery, detailed events, consciousness bias, logical reasoning, and privacy identification, and establishes the Multi-Security Hazard Dataset (MSHD). The study evaluated mainstream models including GPT-4o, Llama-3-70B, and DeepSeek-V3. Performance analysis reveals notable differences among models, and improvement strategies are proposed. Comparative analysis with existing benchmarks demonstrates CS-Eval superior efficiency (18.86%), surpassing SafetyBench, SafetyPrompts, S-Eval, and SALAD-Bench by significant margins. Additionally, findings indicate a nonlinear relationship between enhanced security measures and overall model performance, underscoring the necessity of multifaceted safety strategies in future LLM development and deployment.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Towards a Concise Benchmark for Security Risk Evaluation in Large Language Models

  • Yu Zhang,
  • Yongbing Gao,
  • Weihao Li,
  • Xinguang Wang

摘要

Large Language Models (LLMs) play a crucial role in natural language processing, yet their expanding applications highlight increasing security risks. This paper introduces CS-Eval, a novel benchmark designed to assess LLMs’ capability in addressing safety concerns. CS-Eval focuses on seven key security risks: ethical dilemmas, marginal topics, error discovery, detailed events, consciousness bias, logical reasoning, and privacy identification, and establishes the Multi-Security Hazard Dataset (MSHD). The study evaluated mainstream models including GPT-4o, Llama-3-70B, and DeepSeek-V3. Performance analysis reveals notable differences among models, and improvement strategies are proposed. Comparative analysis with existing benchmarks demonstrates CS-Eval superior efficiency (18.86%), surpassing SafetyBench, SafetyPrompts, S-Eval, and SALAD-Bench by significant margins. Additionally, findings indicate a nonlinear relationship between enhanced security measures and overall model performance, underscoring the necessity of multifaceted safety strategies in future LLM development and deployment.