Darknet, apart from the surface web, is mostly used for privacy and illicit activities over the internet. These are networks that are only accessible to a small number of people rather than the entire internet public and are only accessible through authorization, specific software, and configurations. This includes both benign sites such as academic databases and corporate websites, as well as sites that deal with shadier issues such as black markets, fetish communities, and hacking and piracy. This study proposes a generic machine learning-based technique for identifying Darknet visits, as well as examines and performs statistical preprocessing on the dataset, which provides significant information on Darknet users. Following that, this research looks at various function selection approaches to see which ones are better for detecting and classifying Darknet users. This study employs tuned machine learning (ML) approaches such as Decision Tree (DT), Naive Bayes (NB), Random Forest (RF), and Logistic Regression (LR) based on identified capabilities. Our proposed technique of Deep learning (DL)-based oversampling utilizing Variational Autoencoders (VAE) and Generative Adversarial Networks (GAN) known as Adversarial Autoencoders (AAE) that aids in handling class imbalance is now able to attain 100% accuracy in almost every ML classifier model. To cut down on time, the entire classification model is built on the cloud computing resource, Apache Spark.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Data Balancing in Darknet Traffic Classification Using Deep Learning

  • M. Mahalakshmi,
  • P. Vijay Kumar,
  • S. Navin Kumar,
  • S. Meenakshi Sundaram,
  • M. P. Ramkumar,
  • G. S. R. Emil Selvan

摘要

Darknet, apart from the surface web, is mostly used for privacy and illicit activities over the internet. These are networks that are only accessible to a small number of people rather than the entire internet public and are only accessible through authorization, specific software, and configurations. This includes both benign sites such as academic databases and corporate websites, as well as sites that deal with shadier issues such as black markets, fetish communities, and hacking and piracy. This study proposes a generic machine learning-based technique for identifying Darknet visits, as well as examines and performs statistical preprocessing on the dataset, which provides significant information on Darknet users. Following that, this research looks at various function selection approaches to see which ones are better for detecting and classifying Darknet users. This study employs tuned machine learning (ML) approaches such as Decision Tree (DT), Naive Bayes (NB), Random Forest (RF), and Logistic Regression (LR) based on identified capabilities. Our proposed technique of Deep learning (DL)-based oversampling utilizing Variational Autoencoders (VAE) and Generative Adversarial Networks (GAN) known as Adversarial Autoencoders (AAE) that aids in handling class imbalance is now able to attain 100% accuracy in almost every ML classifier model. To cut down on time, the entire classification model is built on the cloud computing resource, Apache Spark.