Phishing attack has always been a latent threat to the privacy and security of Internet users. With the thriving development in the past decade, machine learning techniques have great potential to reduce this risk via efficiently classifying URLs as either legitimate or malicious. In this paper, extensive experiments are carried out to detect phishing websites based on variable machine learning models. After collecting and preprocessing data of both legitimate and malicious URLs, the mutual information method is adopted to select the most relevant features to achieve higher prediction accuracy. Then the model training and testing are implemented via variable machine learning and deep learning models, including Logistic Regression, Decision Tree, Random Forest, Gradient Boosting, SVM, MLP, Adaboost, GaussianNB, XGB, Light GBM and CatBoost. The results have revealed that Decision Tree Classifier can achieve the highest accuracy of 98.80% and along with good precision, recall, and F1 score. These results show that decision tree models are efficient in detecting phishing URLs and proves its potential for enhancing users security and privacy.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative Study of Artificial Intelligent Approaches for Phishing Website Detection

  • Bingbing Li,
  • Ogbebisi Chukwuebuka Amandi,
  • Mingwu Zhang

摘要

Phishing attack has always been a latent threat to the privacy and security of Internet users. With the thriving development in the past decade, machine learning techniques have great potential to reduce this risk via efficiently classifying URLs as either legitimate or malicious. In this paper, extensive experiments are carried out to detect phishing websites based on variable machine learning models. After collecting and preprocessing data of both legitimate and malicious URLs, the mutual information method is adopted to select the most relevant features to achieve higher prediction accuracy. Then the model training and testing are implemented via variable machine learning and deep learning models, including Logistic Regression, Decision Tree, Random Forest, Gradient Boosting, SVM, MLP, Adaboost, GaussianNB, XGB, Light GBM and CatBoost. The results have revealed that Decision Tree Classifier can achieve the highest accuracy of 98.80% and along with good precision, recall, and F1 score. These results show that decision tree models are efficient in detecting phishing URLs and proves its potential for enhancing users security and privacy.