Phishing attacks remain a significant cybersecurity threat, often exploiting unsuspecting users by disguising malicious URLs as legitimate ones. In this study, we address the challenge of identifying phishing URLs through a comprehensive analysis of machine learning and deep learning algorithms. A curated dataset of 11,430 URLs, featuring 87 extracted features from different aspects, serves as the basis for our investigation. The dataset is thoughtfully balanced, consisting of an equal distribution of phishing and legitimate URLs. Multiple machine learning algorithms were meticulously evaluated and a custom deep learning algorithm was developed from scratch to complement the ensemble of machine learning techniques. Performance metrics including accuracy, precision, recall, and F1 score were employed to assess the efficacy of the models. The results indicate the deep learning algorithm outperformed the machine learning counterparts, achieving an accuracy, precision, recall, and F1 score of 98.03%, while maintaining an efficient time-in-seconds metric of 0.34.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comprehensive Analysis of Machine Learning and Deep Learning Algorithms for Phishing URL Detection

  • Azmi H. Alsaqqa,
  • Samy S. Abu-Naser

摘要

Phishing attacks remain a significant cybersecurity threat, often exploiting unsuspecting users by disguising malicious URLs as legitimate ones. In this study, we address the challenge of identifying phishing URLs through a comprehensive analysis of machine learning and deep learning algorithms. A curated dataset of 11,430 URLs, featuring 87 extracted features from different aspects, serves as the basis for our investigation. The dataset is thoughtfully balanced, consisting of an equal distribution of phishing and legitimate URLs. Multiple machine learning algorithms were meticulously evaluated and a custom deep learning algorithm was developed from scratch to complement the ensemble of machine learning techniques. Performance metrics including accuracy, precision, recall, and F1 score were employed to assess the efficacy of the models. The results indicate the deep learning algorithm outperformed the machine learning counterparts, achieving an accuracy, precision, recall, and F1 score of 98.03%, while maintaining an efficient time-in-seconds metric of 0.34.