Instance hardness measures allow us to assess and understand why some observations from a dataset are difficult to classify. With this information, one may curate and cleanse the training dataset for improved data quality. However, these measures require data to be labeled. This limits their usage in the deployment stage when data is unlabeled. This paper investigates whether it is possible to identify observations that will be hard to classify despite their label. For such, two approaches are tested. The first adapts known instance hardness measures to the unlabeled scenario. The second learns regression meta-models to estimate the instance hardness of new data observations. In experiments, both approaches were better at identifying instances lying in borderline regions of the dataset, which pose a greater difficulty when the label is unknown.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Instance Level Analysis of Classification Difficulty for Unlabeled Data

  • Patricia S. M. Ueda,
  • Adriano Rivolli,
  • Ana Carolina Lorena

摘要

Instance hardness measures allow us to assess and understand why some observations from a dataset are difficult to classify. With this information, one may curate and cleanse the training dataset for improved data quality. However, these measures require data to be labeled. This limits their usage in the deployment stage when data is unlabeled. This paper investigates whether it is possible to identify observations that will be hard to classify despite their label. For such, two approaches are tested. The first adapts known instance hardness measures to the unlabeled scenario. The second learns regression meta-models to estimate the instance hardness of new data observations. In experiments, both approaches were better at identifying instances lying in borderline regions of the dataset, which pose a greater difficulty when the label is unknown.