<p>Effective malware detection is a key priority for Security Operations Centers (SOC). Machine learning (ML) has emerged as a very powerful tool, widely adopted by many malware detection systems. ML models require extensive and high-quality data to perform well. SOCs are often dependent on their proprietary datasets for training and face challenges to obtain sufficient data, due to privacy and Intellectual Property (IP) concerns, limiting their malware detection capabilities. To address these challenges, this paper introduces the adoption of Cross-Silo Federated Learning (FL), a ML technique that enables different participating SOCs to collaboratively train ML malware detection models without explicitly sharing their private data. The deployment of two distinct FL setups, namely Horizontal Federated Learning (HFL) and Vertical Federated Learning (VFL), is explored to address data sample and feature sharing limitations, respectively. The effectiveness of the proposed architectures is evaluated against two large openly available benchmark datasets and compared to conventional Centralized Learning. Finally, a concrete compelling incentive for SOCs to participate in such federations is provided.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cross-Silo Federated Learning in Security Operations Centers for effective malware detection

  • Georgios Xenos,
  • Dimitrios Serpanos

摘要

Effective malware detection is a key priority for Security Operations Centers (SOC). Machine learning (ML) has emerged as a very powerful tool, widely adopted by many malware detection systems. ML models require extensive and high-quality data to perform well. SOCs are often dependent on their proprietary datasets for training and face challenges to obtain sufficient data, due to privacy and Intellectual Property (IP) concerns, limiting their malware detection capabilities. To address these challenges, this paper introduces the adoption of Cross-Silo Federated Learning (FL), a ML technique that enables different participating SOCs to collaboratively train ML malware detection models without explicitly sharing their private data. The deployment of two distinct FL setups, namely Horizontal Federated Learning (HFL) and Vertical Federated Learning (VFL), is explored to address data sample and feature sharing limitations, respectively. The effectiveness of the proposed architectures is evaluated against two large openly available benchmark datasets and compared to conventional Centralized Learning. Finally, a concrete compelling incentive for SOCs to participate in such federations is provided.