The effectiveness of ranking methods in detecting feature drifts in data streams has been evaluated in this work. The FBDD (Feature-Based Drift Detector) method was used for feature drift detection, as creating a feature ranking is a key component of this method. The study was conducted on artificial and real datasets, representing various drifts, such as abrupt, gradual, incremental, and recurring changes. Ten widely used ranking methods were evaluated, including LASSO, the Laplacian Score, and the Kolmogorov-Smirnov method. The analysis focused on key metrics such as classification accuracy (ACC), Matthews correlation coefficient (MCC), and computational efficiency, providing a comprehensive overview of the strengths and weaknesses of each method. The results revealed significant differences in the performance of the methods depending on the nature of the data and the type of drift. This work provides insights into the practical applications of ranking methods for drift detection. It highlights the trade-offs between accuracy, computational efficiency, and the ability to handle different drifts. The findings aim to support researchers and practitioners in selecting the most suitable methods for specific data streams.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluation of the Effectiveness of Ranking Methods in Detecting Feature Drift in Artificial and Real Data

  • Krzysztof Wrobel,
  • Piotr Porwik,
  • Tomasz Orczyk

摘要

The effectiveness of ranking methods in detecting feature drifts in data streams has been evaluated in this work. The FBDD (Feature-Based Drift Detector) method was used for feature drift detection, as creating a feature ranking is a key component of this method. The study was conducted on artificial and real datasets, representing various drifts, such as abrupt, gradual, incremental, and recurring changes. Ten widely used ranking methods were evaluated, including LASSO, the Laplacian Score, and the Kolmogorov-Smirnov method. The analysis focused on key metrics such as classification accuracy (ACC), Matthews correlation coefficient (MCC), and computational efficiency, providing a comprehensive overview of the strengths and weaknesses of each method. The results revealed significant differences in the performance of the methods depending on the nature of the data and the type of drift. This work provides insights into the practical applications of ranking methods for drift detection. It highlights the trade-offs between accuracy, computational efficiency, and the ability to handle different drifts. The findings aim to support researchers and practitioners in selecting the most suitable methods for specific data streams.