<p>In recent times, with the advancement of data analysis in various fields, the distance- and density-based methodologies are used in the construction of an outlier detection methodology, as the distance-based methodologies effectively detect the data points as outliers, which are far from their neighbourhoods, whereas the density-based methodologies effectively detect the data points as outliers in the low-density areas. However, the usual distance-based methodologies struggle with capturing local density variations, and density-based methodologies face challenges in identifying patterns in the regions of low density. Moreover, most of the existing distance- and density-based outlier detection methodologies face challenges with parameter selection, such as determining the neighborhood size that often compromises the accuracy. In this paper, by combining the distance- and density-based methodologies, a hybrid approach using the concepts of mutual neighbors is proposed. The proposed methodology effectively detects the outliers and handles the above-mentioned challenges. The concept of mutual neighbors with a distance-based methodology effectively detects the data points as outliers, which are far from their neighborhoods. Additionally, the concept of mutual neighbors with a density-based methodology effectively detects the data points as outliers, which are present in the low-density region. Then, the outliers detected by both methodologies are combined as a set of final outliers. Experimental results for the evaluation measures; Area Under Curve (AUC), Recall, Precision, and F1-Score highlight the performance of the proposed methodology for the 12 synthetic data sets (the average values are close to 1, as .9454 for AUC, .8983 for Recall, .8675 for Precision, and .8733 for F1-Score) and 10 publicly available real-world data sets (the average values are close to 1, as .8102 for AUC, .906 for Recall, .877 for Precision, and .883 for F1-Score) in comparison with the existing prominent methodologies.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ODMN: A hybrid approach for outlier detection using mutual neighbors

  • Rakhi Rakhi,
  • Shailendra Kumar Tripathi,
  • Bhupendra Gupta,
  • Subir Singh Lamba

摘要

In recent times, with the advancement of data analysis in various fields, the distance- and density-based methodologies are used in the construction of an outlier detection methodology, as the distance-based methodologies effectively detect the data points as outliers, which are far from their neighbourhoods, whereas the density-based methodologies effectively detect the data points as outliers in the low-density areas. However, the usual distance-based methodologies struggle with capturing local density variations, and density-based methodologies face challenges in identifying patterns in the regions of low density. Moreover, most of the existing distance- and density-based outlier detection methodologies face challenges with parameter selection, such as determining the neighborhood size that often compromises the accuracy. In this paper, by combining the distance- and density-based methodologies, a hybrid approach using the concepts of mutual neighbors is proposed. The proposed methodology effectively detects the outliers and handles the above-mentioned challenges. The concept of mutual neighbors with a distance-based methodology effectively detects the data points as outliers, which are far from their neighborhoods. Additionally, the concept of mutual neighbors with a density-based methodology effectively detects the data points as outliers, which are present in the low-density region. Then, the outliers detected by both methodologies are combined as a set of final outliers. Experimental results for the evaluation measures; Area Under Curve (AUC), Recall, Precision, and F1-Score highlight the performance of the proposed methodology for the 12 synthetic data sets (the average values are close to 1, as .9454 for AUC, .8983 for Recall, .8675 for Precision, and .8733 for F1-Score) and 10 publicly available real-world data sets (the average values are close to 1, as .8102 for AUC, .906 for Recall, .877 for Precision, and .883 for F1-Score) in comparison with the existing prominent methodologies.