Malware is a key weapon for APT attackers, which cause huge losses and damages in cyberspace. To effectively classify real-world APT malware, we propose a method that combines prior knowledge features with AutoML. We apply graph pattern clustering to analyze the community structure of threat actor groups and conduct a hierarchical classification based on these communities. Our model incorporates real-world APT samples autonomously collected from open-source intelligence and analyzes the multidimensional impact on the classification of threat actor groups. We obtain an average AUC score of 98.6% on a real-world imbalanced dataset with 20,127 malware instances from 68 threat actor groups. Moreover, it substantially enhances the categorization performance, attaining an average accuracy of 87.4% and a notable 9.72% increase in the Macro-F1 score in the community-based hierarchical model. By leveraging AutoML and graph pattern clustering, we efficiently categorize malware to deepen the understanding of threat actor groups.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automated Mining of Multi-Dimensional Information from APT Malware for Effective Feature Analysis and Threat Actor Attribution

  • Rongqi Jing,
  • Zhengwei Jiang,
  • Qiuyun Wang,
  • Shuwei Wang,
  • Hao Li,
  • Xiao Chen

摘要

Malware is a key weapon for APT attackers, which cause huge losses and damages in cyberspace. To effectively classify real-world APT malware, we propose a method that combines prior knowledge features with AutoML. We apply graph pattern clustering to analyze the community structure of threat actor groups and conduct a hierarchical classification based on these communities. Our model incorporates real-world APT samples autonomously collected from open-source intelligence and analyzes the multidimensional impact on the classification of threat actor groups. We obtain an average AUC score of 98.6% on a real-world imbalanced dataset with 20,127 malware instances from 68 threat actor groups. Moreover, it substantially enhances the categorization performance, attaining an average accuracy of 87.4% and a notable 9.72% increase in the Macro-F1 score in the community-based hierarchical model. By leveraging AutoML and graph pattern clustering, we efficiently categorize malware to deepen the understanding of threat actor groups.