Proposing a Genetic Algorithms-Based Data Selection Method for Imbalanced Medical Datasets
摘要
The class imbalance issue is evident in medical datasets, posing a hurdle for accurate predictive modeling. Maintaining the natural characteristics of medical data while mitigating this issue is paramount in ensuring the success of clinical decision support systems. Therefore, this study proposes a genetic algorithm-based data selection method (GA-DS) for imbalanced medical data reserving its distribution and characteristics. It optimizes the assignment of samples into train and test sets by improving the recognition of rare samples on test data. We presented three versions of the GA-DS, one unconstrained and two constrained (train-set \(>=\) 50%, train-set = 70%) to ensure the reliability of the results. Evaluation of the GA-DS methods on three commonly used imbalanced medical datasets and comparison with SMOTE, random oversampling, and stratified random sampling demonstrated the superior performance of GA-DS by improving the sensitivity score on test data. Future work will investigate its performance on large and highly imbalanced data.