We study the impact and stability of the neighborhood parameter for a selection of popular outlier detection algorithms: kNN, LOF, ABOD, LoOP and SDO. We conduct a sensitivity analysis with data undergoing controlled changes related to: cardinality, dimensionality, global outliers ratio, local outliers ratio, layers of density, density differences between inliers and outliers, and zonification. Experiments reveal how each type of data variation affects algorithms differently in terms of accuracy and runtimes, and discloses the performance dependence on the neighborhood parameter. This serves not only to know how to select its value, but also for assessing accuracy robustness against common data phenomena, as well as algorithms’ tolerance to adjustment variations. kNN, ABOD and SDO stand out, with kNN being the most accurate, ABOD the most suitable for both global and local outliers at the same time, and SDO the most stable in the parameterization. The findings of this work are key to understanding the intrinsic behavior of algorithms based on distance and density estimations, which remain the most efficient and reliable in anomaly detection applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Impact of the Neighborhood Parameter on Outlier Detection Algorithms

  • Félix Iglesias,
  • Conrado Martínez,
  • Tanja Zseby

摘要

We study the impact and stability of the neighborhood parameter for a selection of popular outlier detection algorithms: kNN, LOF, ABOD, LoOP and SDO. We conduct a sensitivity analysis with data undergoing controlled changes related to: cardinality, dimensionality, global outliers ratio, local outliers ratio, layers of density, density differences between inliers and outliers, and zonification. Experiments reveal how each type of data variation affects algorithms differently in terms of accuracy and runtimes, and discloses the performance dependence on the neighborhood parameter. This serves not only to know how to select its value, but also for assessing accuracy robustness against common data phenomena, as well as algorithms’ tolerance to adjustment variations. kNN, ABOD and SDO stand out, with kNN being the most accurate, ABOD the most suitable for both global and local outliers at the same time, and SDO the most stable in the parameterization. The findings of this work are key to understanding the intrinsic behavior of algorithms based on distance and density estimations, which remain the most efficient and reliable in anomaly detection applications.