<p>Credit risk assessment plays a key role for financial services institutions to identify the categories (default or non-default) of borrowers. There are many machine learning algorithms for credit risk assessment. However, the credit data often present the phenomenon of unbalance (fewer defaulters and more non-defaulters) and redundant characteristic variables. In credit risk assessment, little attention is paid to the problem of data unbalance and characteristic variables screening. The paper designs a hybrid machine learning algorithm combining clustering algorithm, Random Forest (RF) and Bayesian Network (BN) to reduce redundant characteristic variables in data. Then the data processed by the hybrid machine learning algorithm is sampled (Fusion Sampling), and the sampled data is used to train the supervised learning algorithm. Experiments show that the selection of characteristic variables reduces the cost of computing resources and data storage, while also reducing the “noise” (the negative impact of variables on the classifier) in the data. The method of data sampling is simple and easy to implement, and greatly improves the prediction performance of the classifier. The paper applies the scheme to the analysis of three bank credit data, proving that the proposed scheme can ensure high predictive performance while deeply reducing characteristic variables and is superior to the state-of-the-art hybrid machine learning algorithms (the combination of supervised learning algorithm and unsupervised learning algorithm).</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Variable Selection and Fusion Sampling of Unbalanced Data in Credit Risk Assessment

  • Shujie Zou,
  • Zhiming Cai,
  • Chiawei Chu,
  • Zefeng Zhao,
  • Haohao Cai,
  • Ning Shen,
  • Jie Ren

摘要

Credit risk assessment plays a key role for financial services institutions to identify the categories (default or non-default) of borrowers. There are many machine learning algorithms for credit risk assessment. However, the credit data often present the phenomenon of unbalance (fewer defaulters and more non-defaulters) and redundant characteristic variables. In credit risk assessment, little attention is paid to the problem of data unbalance and characteristic variables screening. The paper designs a hybrid machine learning algorithm combining clustering algorithm, Random Forest (RF) and Bayesian Network (BN) to reduce redundant characteristic variables in data. Then the data processed by the hybrid machine learning algorithm is sampled (Fusion Sampling), and the sampled data is used to train the supervised learning algorithm. Experiments show that the selection of characteristic variables reduces the cost of computing resources and data storage, while also reducing the “noise” (the negative impact of variables on the classifier) in the data. The method of data sampling is simple and easy to implement, and greatly improves the prediction performance of the classifier. The paper applies the scheme to the analysis of three bank credit data, proving that the proposed scheme can ensure high predictive performance while deeply reducing characteristic variables and is superior to the state-of-the-art hybrid machine learning algorithms (the combination of supervised learning algorithm and unsupervised learning algorithm).