Regional Determinants of Domestic Violence in India: Machine Learning Approach
摘要
In this paper, we analyze the factors affecting the likelihood of an Indian woman experiencing domestic violence. In addition, we also analyze how these factors vary across six regions in India. Using a nationally representative survey, we employ machine learning (ML) techniques such as logistic regression, Least Absolute Shrinkage and Selection Operator (LASSO), Random Forest (RF), Extreme Gradient Boosting (XGBoost), and SHapley Additive exPlanations (SHAP), along with dimension reduction methods like autoencoder. Among logistic regression, LASSO, and XGBoost, RF consistently outperforms the others in both the full sample and regional analyses. While different models prioritize features differently, key predictors of spousal abuse — husband’s control issues, woman’s age at first cohabitation, woman’s family background, and her physical stature — consistently emerge as significant both in the overall sample and in the subsamples divided across the six regions in India. However, at the regional level, cluster analysis reveals significant intra-regional variation, reinforcing the need for localized interventions to effectively address domestic violence in India.