Deep Learning-Based Data Mining Techniques for Anomaly Detection: A Comprehensive Study
摘要
Analysis of anomaly detection (AD) is an important part of data mining, which is used to find events in a dataset that are very different from the norm, and its functions are extended with deep learning. One of the main problems lies, in inherent imbalance between normal and abnormal instances, leading to skewed datasets and potential biases in model performance. Moreover, the complexity and high dimensionality of modern datasets exacerbate the difficulty of accurately identifying anomalies, necessitating robust techniques capable of handling diverse data types and structures. Furthermore, interpretability and scalability emerge as significant concerns, particularly in the context of large-scale datasets and real-time applications. Hence, this review delves into the application of various classifiers, such as data mining techniques and unsupervised clustering algorithms, to identify anomalies within sensor data. The evaluation revealed significant variability in accuracy rates among these classifiers. A matrix profile (MP) detecting system, a type of data mining technique, had the highest success rate (99.95%) among all the methods reviewed. This expresses that data mining techniques have high efficacy for finding anomalies. Conversely, unsupervised clustering algorithms showed the lowest accuracy rate, achieving a minimum of 80%. The study highlights that the prime of classifiers significantly impacts the success of anomaly detection methods, necessitating careful consideration of the unique demands of each application like cyber-intrusion detection, fraud detection health care, Intelligent Road Transmission, etc.