<p>With the rapid growth of data volumes in real-world applications, anomaly detection has become a crucial task across various scenarios. Anomalies are generally defined as data points that constitute a small proportion yet exhibit significantly different patterns. Numerous detection methods have been proposed and applied, ranging from statistical analysis to the recently extensively studied graph representation learning and large language models (LLMs). Existing surveys on the best-performing detection methods based on graph neural networks (GNNs) or graph structures often attempt to classify and summarize these methods based on anomalous structures or categories. However, these studies frequently conflate the data distribution characteristics with detection approaches, failing to clarify the adaptability of detection methods to different data contexts. To address this issue, we propose a more practical taxonomy for GNN-based anomaly detection methods from the perspective of data characteristics and assumptions regarding anomaly distribution. Specifically, we summarize common characteristics and assumptions, discussing the corresponding detection approaches and methods with a focus on their key aspects. We also analyze the attention given to key subspaces of data, clarify the embedding and classification methods that are often conflated in existing surveys, and summarize decision-based detection methods that are frequently overlooked. Additionally, we discuss the application of LLMs in this field, providing insights for future research.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A comprehensive survey on GNN-based anomaly detection: taxonomy, methods, and the role of large language models

  • Ziqi Yuan,
  • Qingyun Sun,
  • Haoyi Zhou,
  • Minglai Shao,
  • Xingcheng Fu

摘要

With the rapid growth of data volumes in real-world applications, anomaly detection has become a crucial task across various scenarios. Anomalies are generally defined as data points that constitute a small proportion yet exhibit significantly different patterns. Numerous detection methods have been proposed and applied, ranging from statistical analysis to the recently extensively studied graph representation learning and large language models (LLMs). Existing surveys on the best-performing detection methods based on graph neural networks (GNNs) or graph structures often attempt to classify and summarize these methods based on anomalous structures or categories. However, these studies frequently conflate the data distribution characteristics with detection approaches, failing to clarify the adaptability of detection methods to different data contexts. To address this issue, we propose a more practical taxonomy for GNN-based anomaly detection methods from the perspective of data characteristics and assumptions regarding anomaly distribution. Specifically, we summarize common characteristics and assumptions, discussing the corresponding detection approaches and methods with a focus on their key aspects. We also analyze the attention given to key subspaces of data, clarify the embedding and classification methods that are often conflated in existing surveys, and summarize decision-based detection methods that are frequently overlooked. Additionally, we discuss the application of LLMs in this field, providing insights for future research.