Malware detection is a critical task in cybersecurity. Machine learning algorithms have been increasingly utilized for malware detection due to their high accuracy and ability to handle large datasets. However, imbalanced data is a common issue in malware detection, as the number of malicious samples is often significantly lower than the number of benign samples. This can result in biased models that prioritize accuracy on the majority class while neglecting the minority class. In this paper, we investigate the impact of imbalanced data on machine learning-based malware detection. We explore different techniques for handling imbalanced data, including oversampling, under-sampling, and hybrid sampling methods. We evaluate the performance of these techniques on a real-world dataset of malware and benign files using several popular machine-learning algorithms. Our results demonstrate that the choice of sampling method can have a significant impact on the performance of the machine learning models. We also show that oversampling methods, such as SMOTE, generally outperform undersampling and hybrid methods in terms of accuracy and F1 score.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring the Impact of Imbalanced Data on Machine Learning-Based Malware Detection

  • Snehal Kulkarni,
  • Akshat Gaurav,
  • Francisco José García Peñalvo,
  • Konstantinos Psannis

摘要

Malware detection is a critical task in cybersecurity. Machine learning algorithms have been increasingly utilized for malware detection due to their high accuracy and ability to handle large datasets. However, imbalanced data is a common issue in malware detection, as the number of malicious samples is often significantly lower than the number of benign samples. This can result in biased models that prioritize accuracy on the majority class while neglecting the minority class. In this paper, we investigate the impact of imbalanced data on machine learning-based malware detection. We explore different techniques for handling imbalanced data, including oversampling, under-sampling, and hybrid sampling methods. We evaluate the performance of these techniques on a real-world dataset of malware and benign files using several popular machine-learning algorithms. Our results demonstrate that the choice of sampling method can have a significant impact on the performance of the machine learning models. We also show that oversampling methods, such as SMOTE, generally outperform undersampling and hybrid methods in terms of accuracy and F1 score.