Binary authorship analysis is a crucial step in malware reverse engineering, but the volume and complexity of the code exacerbate the challenge of this manually intensive task. Consequently, efforts have been made to develop reliable automated tools to facilitate malware authorship analysis; however, many challenges are associated with automated approaches. For instance, the compilation process may remove stylistic features present in the source code. This paper evaluates the features used in existing approaches by utilizing various datasets, including programs written for the Google Code Jam programming competition, student projects from programming courses at multiple universities, and content from GitHub repositories. Additionally, we examined the impact of statistical features on precision, recall, and the false positive rate of these methodologies. The evaluation results reveal that the accuracy of these approaches varies across different application domains and datasets, and some of the selected features appear unrelated to the author’s style, indicating that careful consideration is needed when applying this approach. Finally, using statistical features enhanced the precision and recall of existing approaches while reducing the false positive rate by 10–15%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Revisiting Binary Code Authorship Analysis

  • Saed Alrabaee,
  • Mousa Al-kfairy,
  • Mohammad Bany Taha,
  • Omar Alfandi,
  • Fatma Taher,
  • Jie Tang

摘要

Binary authorship analysis is a crucial step in malware reverse engineering, but the volume and complexity of the code exacerbate the challenge of this manually intensive task. Consequently, efforts have been made to develop reliable automated tools to facilitate malware authorship analysis; however, many challenges are associated with automated approaches. For instance, the compilation process may remove stylistic features present in the source code. This paper evaluates the features used in existing approaches by utilizing various datasets, including programs written for the Google Code Jam programming competition, student projects from programming courses at multiple universities, and content from GitHub repositories. Additionally, we examined the impact of statistical features on precision, recall, and the false positive rate of these methodologies. The evaluation results reveal that the accuracy of these approaches varies across different application domains and datasets, and some of the selected features appear unrelated to the author’s style, indicating that careful consideration is needed when applying this approach. Finally, using statistical features enhanced the precision and recall of existing approaches while reducing the false positive rate by 10–15%.