This study aimed to adapt the Mahalanobis method to use a box plot and Euclidean distance to identify outliers in multivariate data. For each variable in the multivariate data, the box plot was first used to achieve this. Records where all variables were within the fences of the box plot will be classified as non-outliers, and the other records will be included in the unsure dataset. In the modified Euclidean distance, the cut-off point was then determined using the normal data. We simulate multivariate data with and without contamination using multivariate skew-t and multivariate normal data to evaluate our improved approaches. In order to accomplish this, we could generate random multivariate data for each distribution that has a zero mean with uncontaminated data and a non-zero mean in a situation of contamination. We did a comparison between our modified approaches and the minimum vector variance, minimum covariance determinant, minimum volume ellipsoid, and Mahalanobis method. Except for the Mahalanobis approach, our modified methods might perform better than those methods in our simulation. Our improved approaches may even equal the Mahalanobis method in situations of outlier levels of zero, 2%, and 5% of the data; however, in all cases, our modified methods perform better than the Mahalanobis method in the cases of 10% and 20%. Furthermore, we use the proposed method to identify the outlier in an actual dataset of hepatitis C patients and blood donors, where the proposed method can detect most of all severe cases.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Modifying the Mahalanobis Method Using Boxplots and Euclidean Distance to Identify Outliers in Multivariate Data

  • Theeraphat Thanwiset,
  • Nawapon Nakharutai,
  • Manad Khamkong,
  • Parichart Pattarapanitchai

摘要

This study aimed to adapt the Mahalanobis method to use a box plot and Euclidean distance to identify outliers in multivariate data. For each variable in the multivariate data, the box plot was first used to achieve this. Records where all variables were within the fences of the box plot will be classified as non-outliers, and the other records will be included in the unsure dataset. In the modified Euclidean distance, the cut-off point was then determined using the normal data. We simulate multivariate data with and without contamination using multivariate skew-t and multivariate normal data to evaluate our improved approaches. In order to accomplish this, we could generate random multivariate data for each distribution that has a zero mean with uncontaminated data and a non-zero mean in a situation of contamination. We did a comparison between our modified approaches and the minimum vector variance, minimum covariance determinant, minimum volume ellipsoid, and Mahalanobis method. Except for the Mahalanobis approach, our modified methods might perform better than those methods in our simulation. Our improved approaches may even equal the Mahalanobis method in situations of outlier levels of zero, 2%, and 5% of the data; however, in all cases, our modified methods perform better than the Mahalanobis method in the cases of 10% and 20%. Furthermore, we use the proposed method to identify the outlier in an actual dataset of hepatitis C patients and blood donors, where the proposed method can detect most of all severe cases.