The widespread use of the Internet has increased the number of malicious domains used for cybercrimes. However, traditional DNS firewalls struggle to detect malicious domains as they rely on blacklists. To overcome this limitation, machine learning-based approaches have been explored, but they often suffer from limited accuracy. In this paper, we propose two novel methodologies for malicious domain detection: (i) an ensemble-based traditional machine learning framework leveraging ensemble feature selection and (ii) a deep learning model with an optimized architecture designed using a tree-structured Parzen estimator (TPE), a type of Bayesian optimization, for hyperparameter tuning. The ensemble feature selection method integrates four feature selection techniques—univariate feature selection, feature importance, recursive feature elimination, and correlation matrix analysis—reducing the number of features by 62.5%. Our results demonstrate that the proposed ensemble classifier achieves an accuracy of 99.58%, while the optimized deep learning model attains an accuracy of 99.94%. Notably, the deep learning model surpasses traditional ensemble classifiers, including voting, stacking, and random forest, and outperforms state-of-the-art methods by a margin of 1.99%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Malicious Domain Detection with Deep Learning and Bayesian Optimization Approaches

  • Tanu Sree Roy Tama,
  • Md. Sakir Hossain,
  • B. M. Rahful Hasan Shawon,
  • Md. Abeer Hossain,
  • S. M. Farhan Tanvir

摘要

The widespread use of the Internet has increased the number of malicious domains used for cybercrimes. However, traditional DNS firewalls struggle to detect malicious domains as they rely on blacklists. To overcome this limitation, machine learning-based approaches have been explored, but they often suffer from limited accuracy. In this paper, we propose two novel methodologies for malicious domain detection: (i) an ensemble-based traditional machine learning framework leveraging ensemble feature selection and (ii) a deep learning model with an optimized architecture designed using a tree-structured Parzen estimator (TPE), a type of Bayesian optimization, for hyperparameter tuning. The ensemble feature selection method integrates four feature selection techniques—univariate feature selection, feature importance, recursive feature elimination, and correlation matrix analysis—reducing the number of features by 62.5%. Our results demonstrate that the proposed ensemble classifier achieves an accuracy of 99.58%, while the optimized deep learning model attains an accuracy of 99.94%. Notably, the deep learning model surpasses traditional ensemble classifiers, including voting, stacking, and random forest, and outperforms state-of-the-art methods by a margin of 1.99%.