This paper focuses on outlier detection within a non-periodic time series data set. The domain is challenging because the data is usually irregular, skewed, and not normalized. A crucial performance metric is analyzed with techniques for anomaly detection through the Box Plot, Z-test, and Isolation Forest. Each method is perceived to perform differently in terms of detecting anomalies, and their performance is quite dependent on the particular characteristics of the data involved. After multiple iterations, the Box Plot and Isolation Forest have proved to be the most effective. The Z-test is not suitable because the data is neither normalized nor symmetrical. Among these methods, the Isolation Forest has taken the lead with a wider range of anomaly detection as compared to others and included all the outliers recognized by the Box Plot. Considering its capacity to handle the complexity and volume of time series data, Isolation Forest has been deemed the best for anomaly detection in this sector. Although several other time series models such as ARIMA can detect outliers for such data, we will still resort to the conventional methods of outlier detection. The study makes mention of context-appropriate algorithm selection as relevant to precision outlier detection, especially in industries with vast and complex data.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Outlier Detection in a Time Series Data

  • Archita Dasgupta,
  • Aritra Mukhopadhyay,
  • Jyotiraditya Ray,
  • Sayanti Ghosh,
  • Amit Kumar Das

摘要

This paper focuses on outlier detection within a non-periodic time series data set. The domain is challenging because the data is usually irregular, skewed, and not normalized. A crucial performance metric is analyzed with techniques for anomaly detection through the Box Plot, Z-test, and Isolation Forest. Each method is perceived to perform differently in terms of detecting anomalies, and their performance is quite dependent on the particular characteristics of the data involved. After multiple iterations, the Box Plot and Isolation Forest have proved to be the most effective. The Z-test is not suitable because the data is neither normalized nor symmetrical. Among these methods, the Isolation Forest has taken the lead with a wider range of anomaly detection as compared to others and included all the outliers recognized by the Box Plot. Considering its capacity to handle the complexity and volume of time series data, Isolation Forest has been deemed the best for anomaly detection in this sector. Although several other time series models such as ARIMA can detect outliers for such data, we will still resort to the conventional methods of outlier detection. The study makes mention of context-appropriate algorithm selection as relevant to precision outlier detection, especially in industries with vast and complex data.