Supervised machine learning has demonstrated the ability to detect cyber attacks. However, model performance can degrade in the face of changes in the cyber environment such as the introduction of a new attack. This study explores two approaches to augmenting the supervised model in response to a new attack entering the environment. The first approach is to perform transfer learning to update the model after the new attack is detected. This approach requires the collection of data after the new attack has been initiated, but does not require labeled examples when implementing domain adaptation techniques. The second approach implements online learning to periodically retrain the model. This approach requires labeled data and can be computationally more expensive, though a decision on the question when to apply the algorithm is not required. We compare these two approaches on an open-source data set and find that some domain adaptation techniques can recover the performance of the model after the new attack is initiated. However, the best performing domain adaptation technique tends to be attack-specific and several techniques exhibit negative transfer. The online learning approach responds quicker to the new attack but the variance of the performance of the model before the new attack can be higher due to overfitting the model.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Online Learning and Transfer Learning for Cyber Attack Detection

  • Stephen Adams,
  • Jared Byers,
  • Bryn Totah,
  • Jeremy West,
  • Tyler Cody,
  • Peter A. Beling

摘要

Supervised machine learning has demonstrated the ability to detect cyber attacks. However, model performance can degrade in the face of changes in the cyber environment such as the introduction of a new attack. This study explores two approaches to augmenting the supervised model in response to a new attack entering the environment. The first approach is to perform transfer learning to update the model after the new attack is detected. This approach requires the collection of data after the new attack has been initiated, but does not require labeled examples when implementing domain adaptation techniques. The second approach implements online learning to periodically retrain the model. This approach requires labeled data and can be computationally more expensive, though a decision on the question when to apply the algorithm is not required. We compare these two approaches on an open-source data set and find that some domain adaptation techniques can recover the performance of the model after the new attack is initiated. However, the best performing domain adaptation technique tends to be attack-specific and several techniques exhibit negative transfer. The online learning approach responds quicker to the new attack but the variance of the performance of the model before the new attack can be higher due to overfitting the model.