Financial fraud continues to pose significant threats to the global economy. Traditional rule-based or statistical learning methods suffer from low accuracy and limited adaptability when confronted with the evolving complexity of fraudulent activities. Although machine learning models have demonstrated superior predictive capabilities in recent years, their “black-box” nature restricts their practical application in financial institutions. To address this challenge, this study proposes a fraud detection framework that balances predictive performance and interpretability. Using a publicly available financial transaction dataset, we conduct systematic data preprocessing and construct four classifiers: XGBoost, Support Vector Machine (SVM), k-Nearest Neighbors (KNN), and Naïve Bayes. Experimental results on a dataset of 10,000 samples with a 1% fraud ratio show that XGBoost outperforms the other models, achieving an AUC of 0.912 and an F1-score of 89.4%. SHAP (Shapley Additive Explanations) is applied to interpret the XGBoost model, revealing risk score, failed transaction count, and transaction amount are the most influential features. The proposed framework combines strong predictive power with high interpretability, making it suitable for high-stakes financial risk control scenarios and offering data-driven strategies for optimizing fraud monitoring.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Interpretable Machine Learning-Based Fraud Detection Model and Knowledge Discovery in Financial Transactions

  • Weixing Xu

摘要

Financial fraud continues to pose significant threats to the global economy. Traditional rule-based or statistical learning methods suffer from low accuracy and limited adaptability when confronted with the evolving complexity of fraudulent activities. Although machine learning models have demonstrated superior predictive capabilities in recent years, their “black-box” nature restricts their practical application in financial institutions. To address this challenge, this study proposes a fraud detection framework that balances predictive performance and interpretability. Using a publicly available financial transaction dataset, we conduct systematic data preprocessing and construct four classifiers: XGBoost, Support Vector Machine (SVM), k-Nearest Neighbors (KNN), and Naïve Bayes. Experimental results on a dataset of 10,000 samples with a 1% fraud ratio show that XGBoost outperforms the other models, achieving an AUC of 0.912 and an F1-score of 89.4%. SHAP (Shapley Additive Explanations) is applied to interpret the XGBoost model, revealing risk score, failed transaction count, and transaction amount are the most influential features. The proposed framework combines strong predictive power with high interpretability, making it suitable for high-stakes financial risk control scenarios and offering data-driven strategies for optimizing fraud monitoring.