<p>This study aims to enhance default prediction processes for companies by applying machine learning (ML) models, addressing challenges in hyperparameter tuning, sample size, and feature selection. By improving the accuracy of default risk assessments, the research contributes to greater financial and economic stability. Using company data from two emerging economies—Poland (10,503 cases) and Taiwan (6,819 cases)—we implement ML techniques including artificial neural networks, support vector machines, naïve Bayes, logistic regression, decision trees, and K-nearest neighbours. The LASSO method is applied for effective variable selection, allowing a comprehensive evaluation of model performance across varying hyperparameters. Findings indicate that naïve Bayes consistently underperforms, while K-nearest neighbours achieves the highest accuracy. Model performance is sensitive to dataset characteristics and tuning, with a notable risk of overfitting in high-dimensional data scenarios. The study uniquely examines how hyperparameter variation and feature diversity affect model reliability across economies, an underexplored dimension in credit risk prediction. The results provide actionable guidance for regulatory bodies, lenders, and decision-makers aiming to optimize credit risk management and reduce non-performing assets through more efficient data-driven strategies.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Machine learning for credit risk management through cross-economy evidence in default prediction

  • Sunaina Kanojia,
  • Anubhav Arora

摘要

This study aims to enhance default prediction processes for companies by applying machine learning (ML) models, addressing challenges in hyperparameter tuning, sample size, and feature selection. By improving the accuracy of default risk assessments, the research contributes to greater financial and economic stability. Using company data from two emerging economies—Poland (10,503 cases) and Taiwan (6,819 cases)—we implement ML techniques including artificial neural networks, support vector machines, naïve Bayes, logistic regression, decision trees, and K-nearest neighbours. The LASSO method is applied for effective variable selection, allowing a comprehensive evaluation of model performance across varying hyperparameters. Findings indicate that naïve Bayes consistently underperforms, while K-nearest neighbours achieves the highest accuracy. Model performance is sensitive to dataset characteristics and tuning, with a notable risk of overfitting in high-dimensional data scenarios. The study uniquely examines how hyperparameter variation and feature diversity affect model reliability across economies, an underexplored dimension in credit risk prediction. The results provide actionable guidance for regulatory bodies, lenders, and decision-makers aiming to optimize credit risk management and reduce non-performing assets through more efficient data-driven strategies.