An intrusion detection system, commonly referred to as IDS, is a network security tool that continuously monitors network communications. Upon detecting potentially malicious transmissions, it triggers alarms or initiates responsive actions. Numerous researchers have sought to enhance intrusion detection performance by integrating traditional machine learning models with alternative optimization methods. Despite the capability of current intrusion detection models to significantly improve performance, persistent challenges such as inaccurate detection and data preparation activities continue to impact accuracy adversely. In our study, utilizing the CICIDS2017 dataset and the Boruta feature selection algorithm, this paper presents an analytical model featuring various classifiers, including Naive Bayes (NB), Random Forest (RF), K-Nearest Neighbors (KNN), Decision Tree (DT), Stacking (utilizing NB, RF, KNN), Bagging (utilizing DT), and Extreme Gradient Boosting (XGB). Across all feature sets, XGB consistently outperformed all other models, exhibiting superior performance with an impressive accuracy of 99.983% and a remarkably low false alarm rate of 0.017%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Performance Analysis of Machine Learning Classifiers on CICIDS2017 Dataset

  • Pochamreddy Mukesh Reddy,
  • Lav Upadhyay,
  • Lanka Rakesh

摘要

An intrusion detection system, commonly referred to as IDS, is a network security tool that continuously monitors network communications. Upon detecting potentially malicious transmissions, it triggers alarms or initiates responsive actions. Numerous researchers have sought to enhance intrusion detection performance by integrating traditional machine learning models with alternative optimization methods. Despite the capability of current intrusion detection models to significantly improve performance, persistent challenges such as inaccurate detection and data preparation activities continue to impact accuracy adversely. In our study, utilizing the CICIDS2017 dataset and the Boruta feature selection algorithm, this paper presents an analytical model featuring various classifiers, including Naive Bayes (NB), Random Forest (RF), K-Nearest Neighbors (KNN), Decision Tree (DT), Stacking (utilizing NB, RF, KNN), Bagging (utilizing DT), and Extreme Gradient Boosting (XGB). Across all feature sets, XGB consistently outperformed all other models, exhibiting superior performance with an impressive accuracy of 99.983% and a remarkably low false alarm rate of 0.017%.