As generative artificial intelligence evolves, understanding the capabilities in the cybersecurity domain becomes crucial. This paper examines the capability of Large Language Models (LLMs) models in solving cybersecurity certification Multiple Choice Question Answering (MCQA) exams, comparing proprietary and open-weights models. Challenges related to test-set leakage, notably on the widely used MMLU benchmark, emphasize the need for continuous validation of benchmarking results. Open-weights models, namely Mistral Large 2, Qwen 2, and Phi 3, seem to overfit the MMLU Computer Security and indicate less usability for cybersecurity knowledge tasks. The study also introduces the first visual cybersecurity MCQA benchmark, assessing the capability of Large Multimodal Models (LMMs) in interpreting and responding to visual questions. Among the tested models, the proprietary Anthropic Claude 3.5 Sonnet and OpenAI GPT-4o outperformed others in the language and vision-language setting. However, Llama 3.1 model series demonstrated significant advancement in the open-weights domain, signaling potential parity in cybersecurity knowledge with proprietary models in the near future. Code and datasets are available at: https://github.com/GKeppler/GenAICyberSecMCQA .

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating Large Language Models in Cybersecurity Knowledge with Cisco Certificates

  • Gustav Keppler,
  • Jeremy Kunz,
  • Veit Hagenmeyer,
  • Ghada Elbez

摘要

As generative artificial intelligence evolves, understanding the capabilities in the cybersecurity domain becomes crucial. This paper examines the capability of Large Language Models (LLMs) models in solving cybersecurity certification Multiple Choice Question Answering (MCQA) exams, comparing proprietary and open-weights models. Challenges related to test-set leakage, notably on the widely used MMLU benchmark, emphasize the need for continuous validation of benchmarking results. Open-weights models, namely Mistral Large 2, Qwen 2, and Phi 3, seem to overfit the MMLU Computer Security and indicate less usability for cybersecurity knowledge tasks. The study also introduces the first visual cybersecurity MCQA benchmark, assessing the capability of Large Multimodal Models (LMMs) in interpreting and responding to visual questions. Among the tested models, the proprietary Anthropic Claude 3.5 Sonnet and OpenAI GPT-4o outperformed others in the language and vision-language setting. However, Llama 3.1 model series demonstrated significant advancement in the open-weights domain, signaling potential parity in cybersecurity knowledge with proprietary models in the near future. Code and datasets are available at: https://github.com/GKeppler/GenAICyberSecMCQA .