A comprehensive survey on GNN-based anomaly detection: taxonomy, methods, and the role of large language models
摘要
With the rapid growth of data volumes in real-world applications, anomaly detection has become a crucial task across various scenarios. Anomalies are generally defined as data points that constitute a small proportion yet exhibit significantly different patterns. Numerous detection methods have been proposed and applied, ranging from statistical analysis to the recently extensively studied graph representation learning and large language models (LLMs). Existing surveys on the best-performing detection methods based on graph neural networks (GNNs) or graph structures often attempt to classify and summarize these methods based on anomalous structures or categories. However, these studies frequently conflate the data distribution characteristics with detection approaches, failing to clarify the adaptability of detection methods to different data contexts. To address this issue, we propose a more practical taxonomy for GNN-based anomaly detection methods from the perspective of data characteristics and assumptions regarding anomaly distribution. Specifically, we summarize common characteristics and assumptions, discussing the corresponding detection approaches and methods with a focus on their key aspects. We also analyze the attention given to key subspaces of data, clarify the embedding and classification methods that are often conflated in existing surveys, and summarize decision-based detection methods that are frequently overlooked. Additionally, we discuss the application of LLMs in this field, providing insights for future research.