Phishing attacks are one of the most critical attacks in the field of cybersecurity, characterized by deceptive websites that trick users into disclosing sensitive information. Traditional rule-based phishing detection methods struggle to keep up with the evolving attack landscape, while machine learning models often operate as “black boxes,” limiting interpretability. This study explores the application of Explainable Artificial Intelligence (XAI) techniques, specifically SHAP and LIME, to improve the transparency and trustworthiness of phishing website detection models. We evaluate multiple machines learning algorithms, including Gradient Boosting, XGBoost, CatBoost, and Autoencoders, on a real-world dataset of over 11,000 URLs. Our findings show that the ensemble models outperform traditional classifiers in terms of high accuracy and interpretable results. Moreover, our analysis highlights the significant phishing features, including HTTPS, URL length, and domain registration information. This research also significantly advances phishing detection by supplementing it with XAI capability that may be leveraged by security professionals as a means to render actionable insights, disseminating a functional balance between cybersecurity and explainability.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Explainable Machine Learning Models for Phishing Website Detection: Enhancing Transparency and Accuracy in Cybersecurity

  • Sonkarlay J. Y. Weamie,
  • Vinothkumar Kolluru,
  • Kahsay Birhanu Tsadik,
  • Charan Sundar Telaganeni

摘要

Phishing attacks are one of the most critical attacks in the field of cybersecurity, characterized by deceptive websites that trick users into disclosing sensitive information. Traditional rule-based phishing detection methods struggle to keep up with the evolving attack landscape, while machine learning models often operate as “black boxes,” limiting interpretability. This study explores the application of Explainable Artificial Intelligence (XAI) techniques, specifically SHAP and LIME, to improve the transparency and trustworthiness of phishing website detection models. We evaluate multiple machines learning algorithms, including Gradient Boosting, XGBoost, CatBoost, and Autoencoders, on a real-world dataset of over 11,000 URLs. Our findings show that the ensemble models outperform traditional classifiers in terms of high accuracy and interpretable results. Moreover, our analysis highlights the significant phishing features, including HTTPS, URL length, and domain registration information. This research also significantly advances phishing detection by supplementing it with XAI capability that may be leveraged by security professionals as a means to render actionable insights, disseminating a functional balance between cybersecurity and explainability.