<p>While there are frameworks that allow for some autonomy in penetration testing, they still do not permit full autonomy, effective management of context, have no formal safety validation, or reproducible evaluation pipelines. In this paper, we introduce a fully autonomous framework for scalable penetration testing called AutoSec-Agent that uses large language models (LLMs) in a formalized Planner–Summarizer–Validator (PSV) iterative reasoning loop while maintaining safety constraints. In contrast to previous semi-automated solutions, AutoSec-Agent introduces relevance-aware dynamic memory compression, dual-layer pre-execution safety validation, and adaptive hyperparameter selection, which assure reproducible and ethically controlled penetration-testing outcomes in diverse cybersecurity areas. All agent operations are done in a containerized, hardened sandbox, with network scoping, kernel-level isolation and in-depth audit logging. We present SecureCTF-AgentBench, a large-scale dynamic benchmark of 300 challenges, covering web exploits, binary exploits, cryptography, digital forensics, reverse engineering, and multi-host Active Directory enterprise scenarios, whose flags are dynamically generated and whose correctness is verified by solving the challenges. AutoSec-Agent’s macro-average task success rate of 61.3% and best-model success rate of 81.3% (GPT-4o) are 15.5 percentage points higher than the strongest baseline, PentestGPT. The proposed architecture decreases unsafe command generation by 87% and decreases hallucination rates by 62% from baseline systems and with an average of 6,820 token for reasoning per challenge it is still token efficient. The results of ablation studies validate that the three modules which make up the PSV loop, Summarizer, Validator, and incremental planning constraint, all add their own unique value and importantly, work in synergy, in terms of overall performances, safety, and reasoning stability. AutoSec-Agent is freely released as an open source research platform, setting the standard for reproducible, ethically responsible and scalable autonomous cybersecurity research through AI.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AutoSec-Agent: a fully autonomous and ethical multi-agent framework for scalable penetration testing using large language models

  • Rashid Amin,
  • Sajid Mehmood,
  • Syafiq Fauzi Kamarulzaman,
  • Antonio Costanzo,
  • Faisal S. Alsubaei,
  • Asma Hassan Alshehri

摘要

While there are frameworks that allow for some autonomy in penetration testing, they still do not permit full autonomy, effective management of context, have no formal safety validation, or reproducible evaluation pipelines. In this paper, we introduce a fully autonomous framework for scalable penetration testing called AutoSec-Agent that uses large language models (LLMs) in a formalized Planner–Summarizer–Validator (PSV) iterative reasoning loop while maintaining safety constraints. In contrast to previous semi-automated solutions, AutoSec-Agent introduces relevance-aware dynamic memory compression, dual-layer pre-execution safety validation, and adaptive hyperparameter selection, which assure reproducible and ethically controlled penetration-testing outcomes in diverse cybersecurity areas. All agent operations are done in a containerized, hardened sandbox, with network scoping, kernel-level isolation and in-depth audit logging. We present SecureCTF-AgentBench, a large-scale dynamic benchmark of 300 challenges, covering web exploits, binary exploits, cryptography, digital forensics, reverse engineering, and multi-host Active Directory enterprise scenarios, whose flags are dynamically generated and whose correctness is verified by solving the challenges. AutoSec-Agent’s macro-average task success rate of 61.3% and best-model success rate of 81.3% (GPT-4o) are 15.5 percentage points higher than the strongest baseline, PentestGPT. The proposed architecture decreases unsafe command generation by 87% and decreases hallucination rates by 62% from baseline systems and with an average of 6,820 token for reasoning per challenge it is still token efficient. The results of ablation studies validate that the three modules which make up the PSV loop, Summarizer, Validator, and incremental planning constraint, all add their own unique value and importantly, work in synergy, in terms of overall performances, safety, and reasoning stability. AutoSec-Agent is freely released as an open source research platform, setting the standard for reproducible, ethically responsible and scalable autonomous cybersecurity research through AI.