<p>Automated decisions provided by machine learning algorithms are rapidly gaining traction and shaping lending markets, affecting businesses’ performance and consumers’ well-being. Consequently, financial authorities are adapting the regulation, requiring that credit decisions are explainable. Although there are post hoc interpretability techniques capable of fulfilling this task, there is discussion about their reliability. In this article we propose a novel framework to test it. Our work is based on generating datasets intended to resemble typical credit settings, in which we define the importance of the variables. We then use XGBoost and Deep Learning on these datasets, and explain their predictions using SHapley Additive exPlanations (SHAP) and permutation Feature Importance. Finally, we calculate to what extent these explanations match the underlying important variables. Our results suggest that SHAP is better at capturing relevant variables, although the explanations may vary significantly depending on the characteristics of the dataset and model used.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Should We Trust the Credit Decisions Provided by Machine Learning Models?

  • Andrés Alonso-Robisco,
  • José Manuel Carbó

摘要

Automated decisions provided by machine learning algorithms are rapidly gaining traction and shaping lending markets, affecting businesses’ performance and consumers’ well-being. Consequently, financial authorities are adapting the regulation, requiring that credit decisions are explainable. Although there are post hoc interpretability techniques capable of fulfilling this task, there is discussion about their reliability. In this article we propose a novel framework to test it. Our work is based on generating datasets intended to resemble typical credit settings, in which we define the importance of the variables. We then use XGBoost and Deep Learning on these datasets, and explain their predictions using SHapley Additive exPlanations (SHAP) and permutation Feature Importance. Finally, we calculate to what extent these explanations match the underlying important variables. Our results suggest that SHAP is better at capturing relevant variables, although the explanations may vary significantly depending on the characteristics of the dataset and model used.