Unlocking Healthcare Fraud Detection Using Innovations of Machine Learning Strategies
摘要
The identification of fraud in health insurance claims is important because of the huge cost implications involved. This paper aims to determine the applicability of random forest classifiers (RFC), support vector machines (SVM), and K-nearest neighbors (KNNs) in detecting fraudulent claims. Data cleaning and data preprocessing were done, and new feature engineering was done from various datasets such as patient dataset, inpatient dataset, and outpatient dataset. Nevertheless, if the first was equipped with a pretend AUC of 0 and an accuracy of 71%, other comparisons indicated that the Random Forest Classifier outperformed the other two models, with achieved accuracies of 82% for SVM and 80% for KNN. These were the diagnosis index, and the Chronic Disease Index. Using threefold cross-validation, the method involving the use of the precision/recall graph, ROC curve, and calibration curve was used to establish the model’s capacity to differentiate between genuine and fake claims. Based on the findings of this study, it is recommended that Random Forest Classifiers are suitable for this task. Subsequent studies should aim at using higher levels of analysis and develop these models while using larger samples. The two objectives of this research are to design advanced, timely fraud detection instruments, and control mechanisms that are instrumental in cutting costs and enhancing healthcare’s effectiveness. Consequently, this research signifies that machine learning classifiers, particularly Random Forest, hold promise in identifying and eradicating healthcare fraud to foster improved and less expensive medical service delivery.