Similarity Attribute-Based Categorical Attribute Grouping for Outlier Detecting
摘要
Attribute grouping is one of the effective steps in high-dimensional outlier detection – alleviating the interference of “curse of dimensionality”. Existing attribute grouping methods, however, fail to simultaneously reflect local and global similarity among attributes. Therefore, the attribute grouping may not be viable to optimize the performance of outlier detection. In this paper, we propose a novel attribute grouping approach accompanied by an outlier detection algorithm for categorical data using similarity attribute vectors to characterize the local and global similarity of attribute grouping. After defining the first-order and second-order attribute similarities of categorical attributes through attribute subgraph, we construct the similarity attribute vectors characterizing categorical attributes to effectively resemble the local and global similarity among attributes. Next, we propose an automatic mechanism for electing the number of attribute groups and an attribute grouping approach using the similarity attribute vector and silhouette coefficient. We undertake the experiments driven by the UCI and synthetic datasets, demonstrating that the attribute grouping scheme has prominent intra-group compactness and inter-group sparsity. Importantly, compared with the competing methods, the algorithm bolsters the AUC index and the detection efficiency by averages of 6.01% and 41.42% respectively.