End-to-end threat hunting with a novel multiclass dataset for intelligent intrusion detection
摘要
Traditional rule-based intrusion detection systems are increasingly ineffective against modern and rapidly evolving cyber threats, creating the need for intelligent and scalable intrusion detection frameworks supported by realistic datasets. Existing IDS benchmarks often suffer from limitations such as outdated attack scenarios, limited multiclass coverage, and insufficient realism in traffic generation and monitoring environments. Therefore, this paper aims to develop an intelligent end-to-end threat-hunting framework supported by a novel large-scale multiclass intrusion detection dataset generated within a controlled cybersecurity laboratory environment designed to emulate realistic network conditions. The proposed dataset contains more than 7 million labeled network packets, including benign traffic and 15 modern cyberattack categories such as MITM ARP Spoofing, SSH/FTP brute-force attacks, SQL Injection, XSS, Port Scanning, Remote Code Execution, SYN Flood, and multiple DDoS variants. The proposed framework integrates realistic traffic generation, data acquisition, preprocessing, feature engineering, multiclass labeling, and intelligent intrusion detection using several supervised ML and DL models, including Naïve Bayes, Logistic Regression, Random Forest, Decision Trees, Feedforward Neural Networks, Multi-Layer Perceptron, and Convolutional Neural Networks. Traffic generation and monitoring were performed using real-world attacker tools and security platforms, including Kali Linux, Snort, Suricata, Wireshark, pfSense, and OWASP BWA. Experimental results demonstrate that the Decision Tree model achieved the highest overall performance, with detection accuracy reaching 99.9% and prediction latency as low as 1.1 μs. The findings confirm the effectiveness of the proposed framework for scalable real-time intrusion detection and cyber threat analysis. Compared with traditional IDS benchmarks such as NSL-KDD, UNSW-NB15, and CICIDS2017, the proposed dataset provides more realistic multiclass attack generation and live traffic monitoring within a unified experimental environment.