Prediction of Auto Insurance Claims as an Imbalance Data Distribution Problem
摘要
Claim accuracy is one of the major concerns of both clients and insurance companies. Claim accuracy is directly associated with predictive modeling in the insurance business. Predictive modeling applications have played a vital role in the insurance business in areas such as fraud detection and claim prediction. Predictive modeling heavily depends on the data distribution. The present work addresses the challenges of an imbalanced dataset in the context of insurance claim prediction for autos. Prediction of the occurrence of claims in auto insurance has been employed in the present work. XGBoost, LR, SVM, RF, KNN, DT, and GNV have been used to calculate the claim prediction. The predictive model and random oversampling, random undersampling, and SMOTE have been applied, respectively. It has been observed that XGBoost is giving good results in all three cases.