The term “credit card fraud” describes the actual loss of a credit card or the theft of sensitive credit card information. In this study, the Credit Card Fraud Detection dataset was utilized. Due to the used financial dataset extreme imbalance, this study also handled it by using three techniques: (i) class weight (ratio)—balancing, and (ii) synthetic minority oversampling technique (SMOTE), and (iii) undersampling. For detection, a variety of methods utilizing machine learning are available. This study presents a number of approaches that may be applied to the classification of transactions as genuine or fraud. The experiment employed the following approaches: logistic regression, random forest, decision tree, support vector machine, and naive Bayes. The experiment applied these methods on imbalanced, balanced, oversampled, and undersampled dataset. Based on the results, it can be concluded that all algorithms are highly accurate in detecting credit card fraud. Based on handling the imbalanced dataset, the paper shows that the random forest method gives the best results when class weight and oversampling techniques are applied to the dataset. Whereas, naive Bayes model gives the best results after applying undersampling technique.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Credit Card Fraud Detection Using Machine Learning with Handling Financial Imbalanced Dataset

  • Amjad Qtaish,
  • Fadi Herzallah

摘要

The term “credit card fraud” describes the actual loss of a credit card or the theft of sensitive credit card information. In this study, the Credit Card Fraud Detection dataset was utilized. Due to the used financial dataset extreme imbalance, this study also handled it by using three techniques: (i) class weight (ratio)—balancing, and (ii) synthetic minority oversampling technique (SMOTE), and (iii) undersampling. For detection, a variety of methods utilizing machine learning are available. This study presents a number of approaches that may be applied to the classification of transactions as genuine or fraud. The experiment employed the following approaches: logistic regression, random forest, decision tree, support vector machine, and naive Bayes. The experiment applied these methods on imbalanced, balanced, oversampled, and undersampled dataset. Based on the results, it can be concluded that all algorithms are highly accurate in detecting credit card fraud. Based on handling the imbalanced dataset, the paper shows that the random forest method gives the best results when class weight and oversampling techniques are applied to the dataset. Whereas, naive Bayes model gives the best results after applying undersampling technique.