<p>The arms race between malware authors and detection frameworks is marked by continuous malware mutations and corresponding model retraining efforts. With over 1.5 million new samples reported daily&#xa0;(VirusTotal Statistics, 2025 <a href="https://www.virustotal.com/en/statistics/">https://www.virustotal.com/en/statistics/</a>), retraining has become the <i>de facto</i> response to evolving threats. In this paper, we question the effectiveness of this approach by exposing key limitations: while retraining offers only marginal improvements in detecting malicious samples, it often degrades performance on benign samples. To address these challenges, we evaluate multiple retraining strategies that enable the timely detection of emerging malware families while tracking mutation patterns. Our analysis reveals that retraining can unintentionally aid adversaries by allowing the reuse of old malware samples, which online detectors often discard. Additionally, we uncover labeling inconsistencies−such as family renaming−across online detection engines, which obscure shared malicious capabilities and weaken family-based detection efforts.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

One step forward, two steps back: ML-based malware detection under concept drift

  • Ahmed Abusnaina,
  • Afsah Anwar,
  • Muhammad Saad,
  • Abdulrahman Alabduljabbar,
  • Rhongho Jang,
  • Saeed Salem,
  • David Mohaisen

摘要

The arms race between malware authors and detection frameworks is marked by continuous malware mutations and corresponding model retraining efforts. With over 1.5 million new samples reported daily (VirusTotal Statistics, 2025 https://www.virustotal.com/en/statistics/), retraining has become the de facto response to evolving threats. In this paper, we question the effectiveness of this approach by exposing key limitations: while retraining offers only marginal improvements in detecting malicious samples, it often degrades performance on benign samples. To address these challenges, we evaluate multiple retraining strategies that enable the timely detection of emerging malware families while tracking mutation patterns. Our analysis reveals that retraining can unintentionally aid adversaries by allowing the reuse of old malware samples, which online detectors often discard. Additionally, we uncover labeling inconsistencies−such as family renaming−across online detection engines, which obscure shared malicious capabilities and weaken family-based detection efforts.