Data Composition for Continual Learning in Application of Cyberattack Detection
摘要
Continual learning (CL) focuses on enabling machine learning algorithms to learn from a series of tasks without forgetting previously acquired knowledge. The use of continual learning has not been widely explored in cybersecurity and network safety applications, partially due to the lack of proper datasets. Besides, the benchmark datasets used in CL methods are often relatively restrictive in terms of data distribution shift among the tasks. In this work, we present a CL benchmark framework to construct datasets for CL in cybersecurity applications. For the cybersecurity applications, the proposed framework can generate datasets for CL under distribution shifts in data inputs (e.g., features of internet traffic flow), distribution shifts in data output (e.g., intrusion types), and distribution shifts in both data inputs and outputs, respectively. Moreover, we propose several distance-based and model-based metrics to meticulously quantify the magnitude of distribution shift between datasets of the tasks. We elaborate the construction of benchmark datasets and evaluate the quality of the constructed datasets by applying several existing CL methods and investigating their performance.