Trojan Vulnerabilities in Host-Based Intrusion Detection Systems
摘要
Host-based intrusion detection systems (HIDS) play a critical role in cybersecurity, yet they remain vulnerable to stealthy adversarial attacks. In particular, Trojan attacks—where hidden triggers are embedded into training data—can manipulate models to misclassify malicious activity while remaining undetected. In this paper, we investigate these vulnerabilities using the DARPA OpTC dataset, a large-scale benchmark simulating real-world cyber operations. We present a trigger identification framework leveraging random forests and n-gram feature importance analysis, followed by targeted poisoning strategies that embed highly effective triggers at low injection rates. Through extensive experiments on DeepLog and LogBERT models, we demonstrate that carefully crafted trigger injections can significantly reduce detection confidence on malicious sequences without degrading performance on benign data. Additionally, we provide discernibility analysis showing that poisoned models are nearly indistinguishable from clean models based on weight inspection alone. Our findings underscore the urgent need for proactive defenses and contribute practical methodologies for trigger discovery, poisoning assessment, and model vulnerability evaluation. These results offer actionable insights for security practitioners and lay the foundation for more robust Trojan detection in operational environments.