E2E-AutoPT: An End-to-End Automated Penetration Testing with LSTM-PPO Approach
摘要
With the rapid growth of the Internet, malicious cyber activities have caused significant user losses, underscoring the need for enhanced cybersecurity. Automated Penetration Testing (AutoPT) has emerged as an effective method for identifying vulnerabilities. However, effective AutoPT requires utilizing contextual information across temporal dimensions and comprehensive text data. To address these challenges, we propose an advanced end-to-end (E2E) model for AutoPT that integrates scanning information embedding with a novel Long Short-Term Memory-Proximal Policy Optimization (LSTM-PPO) approach. The LSTM-PPO architecture enhances action selection by incorporating scanning information in an E2E manner, improving the use of spatial and historical data. Experimental results show that the proposed approach outperforms traditional PPO algorithms, achieving higher cumulative rewards in fewer steps. Specifically, the LSTM-PPO approach achieves 18.86%, 19.32%, 10.29%, and 29.47% higher convergence rewards than PPO algorithms in networks with 20, 30, 40 hosts, and large-scale networks, respectively. These findings highlight significant improvements in addressing decision-making challenges in AutoPT, potentially enhancing penetration testing efficiency and strengthening cybersecurity.