<p>Classification is a vital practice in data mining and knowledge discovery from datasets. A major obstacle within this area involves dealing with classification issues, specifically concerning the varying quantities of samples within each category. Therefore, in this study, a new under-sampling technique utilizing a data-level approach is introduced by combining subtractive clustering and fuzzy similarity measures to remove the majority of class samples to resolve the imbalanced data problem. In the proposed SC-FSM framework, first, the majority of samples are split into many clusters using subtractive clustering, and then the samples are sorted using fuzzy similarity criteria, and many samples are selected from each cluster according to their size. In this research, the unbalanced binary data set of Keel software was used, and after the preprocessing steps, SVM and C4.5 classification algorithms were used for classification. The empirical outcomes on the benchmark datasets showed that the recommended SC-FSM framework had a better output compared to previous under-sampling techniques. Statistical tests were used for more detailed analysis. The statistical analysis results showed the superiority of the Koczy fuzzy similarity measure with subtractive clustering compared to other fuzzy similarity measures and other under-sampling techniques in unbalanced data classification.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SC-FSM: a new hybrid framework based on subtractive clustering and fuzzy similarity measures for imbalanced data classification

  • Hua Ren,
  • Shuying Zhai,
  • Xiaowu Wang

摘要

Classification is a vital practice in data mining and knowledge discovery from datasets. A major obstacle within this area involves dealing with classification issues, specifically concerning the varying quantities of samples within each category. Therefore, in this study, a new under-sampling technique utilizing a data-level approach is introduced by combining subtractive clustering and fuzzy similarity measures to remove the majority of class samples to resolve the imbalanced data problem. In the proposed SC-FSM framework, first, the majority of samples are split into many clusters using subtractive clustering, and then the samples are sorted using fuzzy similarity criteria, and many samples are selected from each cluster according to their size. In this research, the unbalanced binary data set of Keel software was used, and after the preprocessing steps, SVM and C4.5 classification algorithms were used for classification. The empirical outcomes on the benchmark datasets showed that the recommended SC-FSM framework had a better output compared to previous under-sampling techniques. Statistical tests were used for more detailed analysis. The statistical analysis results showed the superiority of the Koczy fuzzy similarity measure with subtractive clustering compared to other fuzzy similarity measures and other under-sampling techniques in unbalanced data classification.