Imbalance class problem : an analytical mapping using spreadsheet, VOSviewer, and large language models
摘要
Imbalance class distribution is a buzzword today. In the imbalance class, the samples of one of the target class labels are less than the others. The imbalanced data can remarkably skew the classifier’s performance, leading to a prophecy bias favoring the class having more samples. Although the minority class is of more significance, algorithms based on classification are incredibly accurate for mostly majority classes. Data generated by applications like medical, risk management, fault diagnosis, fraud detection, face recognition, predictive maintenance, etc., are naturally imbalanced. The revolutionary paradigm perspective on imbalance class has propelled scientific advancement and research. Therefore, an extremely vital exploration is desired to excerpt scientific progress paths. The proposed work supports the notion by a Scientometric analysis and literature review outlining the state of research on the imbalance class. This study actively utilizes Spreadsheets, VOSviewer, and Large Language Models for analysis. The Scientometric study represents analytical aspects and provides deep insight into publication, citation patterns, keyword co-occurrence analysis, geographical distribution analysis, and most prolific journal analysis of class imbalance. Large Language Models portray a democratic viewpoint on specific terms of imbalance classification. The literature offers a preliminary taxonomy that suggests various approaches to handling imbalanced data. Overall, the literature on class imbalance portrays prospects for future research and insightful recommendations to the academic community.