Credit risk assessment plays a critical role in ensuring responsible lending and reducing financial losses in the banking sector. With the growing availability of financial data, machine learning (ML) has emerged as a powerful tool for evaluating creditworthiness by uncovering hidden patterns. However, the highly imbalanced nature of credit datasets where non-default cases vastly outnumber defaults poses a significant challenge to predictive accuracy. To address this, our study examines and compares six resampling techniques: Synthetic Minority Over-sampling Technique (SMOTE), Random Over Sampling, Random Under Sampling, Adaptive Synthetic Sampling (ADASYN), Tomek Links, and a hybrid SMOTE-RUS approach. These methods are applied across various ML models, including Decision Tree, Random Forest, XGBoost, Neural Network, AdaBoost, CatBoost, and K-Nearest Neighbors. The hybrid SMOTE-RUS technique demonstrated the most balanced outcomes, enhancing accuracy while reducing overfitting risks. SMOTE and ADASYN also showed notable improvements in detecting minority-class instances. This research offers a detailed comparative analysis of resampling strategies, helping to guide the development of more reliable and equitable credit scoring models for real-world applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Comprehensive Evaluation of Data Balancing Techniques and Machine Learning Models for Credit Risk Assessment

  • Safae Ndama,
  • Oussama Ndama,
  • El Mokhtar En-Naimi

摘要

Credit risk assessment plays a critical role in ensuring responsible lending and reducing financial losses in the banking sector. With the growing availability of financial data, machine learning (ML) has emerged as a powerful tool for evaluating creditworthiness by uncovering hidden patterns. However, the highly imbalanced nature of credit datasets where non-default cases vastly outnumber defaults poses a significant challenge to predictive accuracy. To address this, our study examines and compares six resampling techniques: Synthetic Minority Over-sampling Technique (SMOTE), Random Over Sampling, Random Under Sampling, Adaptive Synthetic Sampling (ADASYN), Tomek Links, and a hybrid SMOTE-RUS approach. These methods are applied across various ML models, including Decision Tree, Random Forest, XGBoost, Neural Network, AdaBoost, CatBoost, and K-Nearest Neighbors. The hybrid SMOTE-RUS technique demonstrated the most balanced outcomes, enhancing accuracy while reducing overfitting risks. SMOTE and ADASYN also showed notable improvements in detecting minority-class instances. This research offers a detailed comparative analysis of resampling strategies, helping to guide the development of more reliable and equitable credit scoring models for real-world applications.