Credit risk assessment is critical for financial institutions to balance profitability and risk management. Traditional credit scoring models, such as logistic regression, are limited by assumptions of linearity and an inability to address the complexities of borrower behavior. Machine learning has introduced advanced methods to enhance predictive accuracy, but challenges persist in handling missing data, particularly for rejected loan applicants whose repayment behavior is unobservable. This issue, known as the reject inference problem, often introduces selection bias, as models rely exclusively on data from approved applicants. This study proposes the Non-Random Missing Parameter Iterative (NRMPI) method, a novel framework designed to address the reject inference problem under non-random missing data assumptions. By incorporating an adjustment parameter, NRMPI accounts for differences in default probabilities between accepted and rejected applicants, leveraging a selection modeling framework to improve parameter estimation and predictive performance. Through simulations and empirical tests on the Lending Club dataset, NRMPI outperforms traditional methods in recall, especially under high rejection rates and strong selection bias, while maintaining comparable performance on other metrics. The findings highlight NRMPI’s ability to mitigate selection bias, offering a robust solution for credit risk modeling. The study also provides practical insights into optimizing model parameters, such as the adjustment parameter \(\gamma \) , to balance interpretability and accuracy. These advancements enable financial institutions to enhance credit decision-making, reduce non-performing loans, and optimize resource allocation in complex lending environments. Future work will explore the integration of temporal dynamics and economic factors to further refine the model’s applicability.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Inference of Rejection in Credit Scoring Using the NRMPI Method: A Non-random Missing Data Approach

  • Baoshan Li,
  • Renhua Gao,
  • Zhao Chen,
  • Christina Dan Wang

摘要

Credit risk assessment is critical for financial institutions to balance profitability and risk management. Traditional credit scoring models, such as logistic regression, are limited by assumptions of linearity and an inability to address the complexities of borrower behavior. Machine learning has introduced advanced methods to enhance predictive accuracy, but challenges persist in handling missing data, particularly for rejected loan applicants whose repayment behavior is unobservable. This issue, known as the reject inference problem, often introduces selection bias, as models rely exclusively on data from approved applicants. This study proposes the Non-Random Missing Parameter Iterative (NRMPI) method, a novel framework designed to address the reject inference problem under non-random missing data assumptions. By incorporating an adjustment parameter, NRMPI accounts for differences in default probabilities between accepted and rejected applicants, leveraging a selection modeling framework to improve parameter estimation and predictive performance. Through simulations and empirical tests on the Lending Club dataset, NRMPI outperforms traditional methods in recall, especially under high rejection rates and strong selection bias, while maintaining comparable performance on other metrics. The findings highlight NRMPI’s ability to mitigate selection bias, offering a robust solution for credit risk modeling. The study also provides practical insights into optimizing model parameters, such as the adjustment parameter \(\gamma \) , to balance interpretability and accuracy. These advancements enable financial institutions to enhance credit decision-making, reduce non-performing loans, and optimize resource allocation in complex lending environments. Future work will explore the integration of temporal dynamics and economic factors to further refine the model’s applicability.