As cybersecurity faces increasingly severe challenges, phishing attacks remain one of the most threatening malicious activities in the digital landscape. To effectively counter this threat, this paper proposes a multi-layer security level phishing URL detection model called ILSPP, which is based on an ensemble incremental learning algorithm and the sum of predicted probabilities. As a model integrating multiple incremental learning algorithms, it not only outperforms single incremental learning models in robustness and generalization ability but also overcomes the limitation of traditional machine learning methods, which cannot perform continuous training. Building upon prior research, we have also implemented a two-stage selection strategy for the incremental learning algorithms: the first stage operates independently of the model, while the second stage is embedded within the model. Based on the three incremental learning algorithms selected in the second stage, a combined prediction is made using a soft-voting mechanism. By setting multiple security levels, the model determines the legitimacy of websites according to different thresholds, thereby adapting to various scenario requirements. Experimental results demonstrate that the proposed model achieves a high accuracy of 99.35% on a real-world phishing dataset, significantly outperforming existing detection methods. Furthermore, by constructing an extensible dataset based on URL features, this paper eliminates the reliance on HTML features, reducing the complexity of the feature extraction process and the risk of network attacks. A balanced dataset containing 42 features and 50,693 latest URLs are provided, contributing to the advancement of research in this field.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ILSPP: A Multi-Layer Security Level Phishing URL Detection Model

  • Yuqing Liu,
  • Yong Wang,
  • Lin Zhou,
  • Tingting Wang

摘要

As cybersecurity faces increasingly severe challenges, phishing attacks remain one of the most threatening malicious activities in the digital landscape. To effectively counter this threat, this paper proposes a multi-layer security level phishing URL detection model called ILSPP, which is based on an ensemble incremental learning algorithm and the sum of predicted probabilities. As a model integrating multiple incremental learning algorithms, it not only outperforms single incremental learning models in robustness and generalization ability but also overcomes the limitation of traditional machine learning methods, which cannot perform continuous training. Building upon prior research, we have also implemented a two-stage selection strategy for the incremental learning algorithms: the first stage operates independently of the model, while the second stage is embedded within the model. Based on the three incremental learning algorithms selected in the second stage, a combined prediction is made using a soft-voting mechanism. By setting multiple security levels, the model determines the legitimacy of websites according to different thresholds, thereby adapting to various scenario requirements. Experimental results demonstrate that the proposed model achieves a high accuracy of 99.35% on a real-world phishing dataset, significantly outperforming existing detection methods. Furthermore, by constructing an extensible dataset based on URL features, this paper eliminates the reliance on HTML features, reducing the complexity of the feature extraction process and the risk of network attacks. A balanced dataset containing 42 features and 50,693 latest URLs are provided, contributing to the advancement of research in this field.