This study investigates the application of Explainable Artificial Intelligence (XAI) techniques in credit risk modeling, focusing on predicting home loan defaults using the Home Credit Default Risk dataset, with the aim of enhancing financial institutions’ risk management strategies through predictive analytics while ensuring model interpretability; the research compares the performance and interpretability of Logistic Regression, Random Forest, CatBoost, and LightGBM under imbalanced data conditions, utilizing rigorous exploratory data analysis, resampling techniques, new financial feature creation, and Information Value (IV) analysis to identify key predictors; results indicate that CatBoost and LightGBM consistently outperform other models, with LightGBM achieving a Receiver Operating Curve-Area Under Curve (ROCAUC) value of 0.76 and an F1-score of 0.74. Furthermore, the incorporation of SHAP improved the transparency of the models by offering essential insights into the impact of individual features.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

XAI-Driven Credit Risk Modeling: A Comparative Analyses of Model Performance and Interpretability

  • Abdelilah Haji,
  • Badr Hssina

摘要

This study investigates the application of Explainable Artificial Intelligence (XAI) techniques in credit risk modeling, focusing on predicting home loan defaults using the Home Credit Default Risk dataset, with the aim of enhancing financial institutions’ risk management strategies through predictive analytics while ensuring model interpretability; the research compares the performance and interpretability of Logistic Regression, Random Forest, CatBoost, and LightGBM under imbalanced data conditions, utilizing rigorous exploratory data analysis, resampling techniques, new financial feature creation, and Information Value (IV) analysis to identify key predictors; results indicate that CatBoost and LightGBM consistently outperform other models, with LightGBM achieving a Receiver Operating Curve-Area Under Curve (ROCAUC) value of 0.76 and an F1-score of 0.74. Furthermore, the incorporation of SHAP improved the transparency of the models by offering essential insights into the impact of individual features.