Integrating GIS and ensemble learning models to predict landslide-prone zones in Chamoli District, India
摘要
Landslides are among the most hazardous geomorphological processes, especially in mountainous terrains where they threaten lives, disrupt infrastructure, and degrade ecosystems. The Chamoli district in Uttarakhand, India, with its rugged topography, erratic rainfall, and anthropogenic disturbances, is highly susceptible to such events. This research focuses on generating a detailed landslide susceptibility map for Chamoli using four machine learning algorithms Naïve Bayes (NB), K-Nearest Neighbors (KNN), Random Forest (RF), and Extreme Gradient Boosting (XGBoost). A total of sixteen causative factors were selected and processed using Geographic Information System (GIS) techniques. To avoid multicollinearity and ensure the reliability of input variables, statistical validation was performed. The landslide inventory, consisting of 778 past events, was split into 70% training and 30% testing datasets for model development and evaluation. The models were assessed using statistical indicators such as sensitivity, specificity, precision, accuracy, F1-score, Matthews Correlation Coefficient (MCC), and the Area Under the Curve (AUC) of the Receiver Operating Characteristic (ROC). Among the models, XGBoost showed the highest performance (AUC = 0.95), followed by RF (0.83), KNN (0.79), and NB (0.78). The susceptibility analysis revealed that XGBoost categorized 17.53% of the area as highly vulnerable, outperforming RF (14.08%), KNN (14.00%), and NB (4.55%). The enhanced accuracy of XGBoost and RF stems from their ensemble learning approach, which effectively captures nonlinear relationships and mitigates overfitting. The resulting susceptibility maps are crucial tools for risk management, infrastructure planning, and sustainable development in hazard-prone Himalayan regions. This study reinforces the importance of machine learning in natural hazard assessment and suggests that future efforts should incorporate real-time data, finer-resolution remote sensing, and deep learning methods to further refine prediction accuracy.