In pattern recognition research, model (algorithm) performance measure is a very important research direction, because the performance measure index has been used to evaluate the model performance throughout the whole process of model estimation, evaluation and selection. Currently, AUC (Area under the ROC (Receiver Operating Characteristic) Curve) measure has became a benchmark performance measure index for classification algorithm. In practical applications, the confidence interval technique of AUC measure is always used to measure the performance of classification algorithm. As we confirmed through simulated experiments, however, those widely used symmetrical confidence intervals with the form of Mean ± SD (Standard Deviation) based on normal distribution assumption may be inappropriate and often exhibit low accuracy, this is because the distribution of AUC measure is actually non-symmetrical. Thus, a new non-symmetrical confidence interval of AUC measure based on K-fold cross-validation is presented by theoretically analyzing its approximate distribution in this paper. Extensive simulated and real data experiments show that the proposed non-symmetrical confidence interval has higher degrees of confidence and shorter interval lengths than the benchmark symmetrical confidence intervals of AUC measure based on K-fold cross-validated t and corrected K-fold cross-validated t distributions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Non-symmetrical Confidence Interval of AUC Measure Based on Cross-Validation

  • Yu Wang,
  • Xiaoyan Zhao,
  • Xingli Yang

摘要

In pattern recognition research, model (algorithm) performance measure is a very important research direction, because the performance measure index has been used to evaluate the model performance throughout the whole process of model estimation, evaluation and selection. Currently, AUC (Area under the ROC (Receiver Operating Characteristic) Curve) measure has became a benchmark performance measure index for classification algorithm. In practical applications, the confidence interval technique of AUC measure is always used to measure the performance of classification algorithm. As we confirmed through simulated experiments, however, those widely used symmetrical confidence intervals with the form of Mean ± SD (Standard Deviation) based on normal distribution assumption may be inappropriate and often exhibit low accuracy, this is because the distribution of AUC measure is actually non-symmetrical. Thus, a new non-symmetrical confidence interval of AUC measure based on K-fold cross-validation is presented by theoretically analyzing its approximate distribution in this paper. Extensive simulated and real data experiments show that the proposed non-symmetrical confidence interval has higher degrees of confidence and shorter interval lengths than the benchmark symmetrical confidence intervals of AUC measure based on K-fold cross-validated t and corrected K-fold cross-validated t distributions.