<p>Data characteristics that reveal the complexity of the dataset, such as overlap, lack of density, the presence of noisy points, etc. are key factors for the deterioration of the classification task of imbalanced datasets. The class size imbalance is not the main problem, but its combination with the above mentioned characteristics. Despite this, when applying sampling methods for preprocessing imbalanced data, the sampling ratio is selected based on the class sizes. In this paper, we propose a methodology to select the sampling ratio by seeking for a balance in the complexity of the classes instead of in their sizes. Our proposed methodology, called Hostility-Aware Ratio for Sampling (HARS), tracks how the complexity of the classes changes when a sampling method is applied for different ratios of minority to majority instances, and recommends the ratio for which there is a balance between the class complexities. Complexity is gauged through the <i>hostility measure</i>, a complexity measure that estimates the probability of misclassifying an instance, a class, or the entire dataset. The proposal is assessed on a total of 66 real datasets and compared with the state-of-the-art (SOTA) methods providing satisfactory classification results that validate the use of the complexity for the task of choosing the sampling ratio and that a balance in complexity favors a more balanced learning process of classifiers.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Selecting sampling ratios in imbalanced datasets through class complexity

  • Carmen Lancho,
  • Marina Cuesta,
  • Isaac Martín De Diego,
  • Víctor Aceña,
  • Javier M. Moguerza

摘要

Data characteristics that reveal the complexity of the dataset, such as overlap, lack of density, the presence of noisy points, etc. are key factors for the deterioration of the classification task of imbalanced datasets. The class size imbalance is not the main problem, but its combination with the above mentioned characteristics. Despite this, when applying sampling methods for preprocessing imbalanced data, the sampling ratio is selected based on the class sizes. In this paper, we propose a methodology to select the sampling ratio by seeking for a balance in the complexity of the classes instead of in their sizes. Our proposed methodology, called Hostility-Aware Ratio for Sampling (HARS), tracks how the complexity of the classes changes when a sampling method is applied for different ratios of minority to majority instances, and recommends the ratio for which there is a balance between the class complexities. Complexity is gauged through the hostility measure, a complexity measure that estimates the probability of misclassifying an instance, a class, or the entire dataset. The proposal is assessed on a total of 66 real datasets and compared with the state-of-the-art (SOTA) methods providing satisfactory classification results that validate the use of the complexity for the task of choosing the sampling ratio and that a balance in complexity favors a more balanced learning process of classifiers.