Telemarketing is currently the most effective method for banks to sell long-term deposits and enhance profitability. Leveraging machine learning techniques can optimize profits by accurately targeting the most promising potential customers, particularly addressing the challenge of imbalanced classification inherent in this domain. This paper proposes an algorithm to address the imbalance problem by using the minority instance’s contribution to the variance explained by principal components to train a weighted random forest after clustering the minority class with K-means. Using the UCI bank marketing dataset with an 80:20 training-evaluation split, the proposed algorithm achieves the best compromise between sensitivity and precision without altering the data distribution. This performance is benchmarked against three classifiers-Random Forest (RF), Logistic Regression (LR), and Support Vector Machine (SVM)-and oversampling techniques: SMOTE, K-means SMOTE, and ADASYN.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving Imbalanced Data Classification in Bank Marketing Using Cluster PCA-Based Weighted Random Forest

  • Dalia ATIF

摘要

Telemarketing is currently the most effective method for banks to sell long-term deposits and enhance profitability. Leveraging machine learning techniques can optimize profits by accurately targeting the most promising potential customers, particularly addressing the challenge of imbalanced classification inherent in this domain. This paper proposes an algorithm to address the imbalance problem by using the minority instance’s contribution to the variance explained by principal components to train a weighted random forest after clustering the minority class with K-means. Using the UCI bank marketing dataset with an 80:20 training-evaluation split, the proposed algorithm achieves the best compromise between sensitivity and precision without altering the data distribution. This performance is benchmarked against three classifiers-Random Forest (RF), Logistic Regression (LR), and Support Vector Machine (SVM)-and oversampling techniques: SMOTE, K-means SMOTE, and ADASYN.