Facing the Challenge of Leveraging Untrained Humans in Malware Analysis
摘要
Software binary analysis, tools and machine learning aid security analysts in interpreting data, by automated means that filter, prioritize, and arrange pertinent information for skilled analysts. In this work, we revisit cooperative human-machine teams and evaluate the possibility of enabling untrained humans to assist machines and skilled analysts in their analysis of software binaries. Specifically, we propose a pipeline to transform a complex input domain into facial images on which untrained individuals make similarity decisions. Our faces include realistic human, animal, artistic, and anime faces that preserve inherent distances between data points of the input domain. Our approach is evaluated through a human study, where untrained respondents with minimal training successfully flag machine misclassifications. The untrained human does not replace the machine or skilled analyst, instead, utilized in a triage setting, to identify samples without historical precedence, deferring the decision to the skilled analyst for deeper inspection.