<p>Landslides in mountainous regions require advanced predictive frameworks to mitigate escalating risks to communities and infrastructure. This study has proposed a novel hybrid ensemble Machine Learning (ML) models integrating Extreme Gradient Boosting (XGB), Stochastic Gradient Boosting (SGB), and Rotated Random Forest (RRF) with wrapper-based Recursive Feature Elimination (RFE) algorithm for landslide susceptibility mapping (LSM) in Uttarakhand, India, a Himalayan region chronically vulnerable to slope failures. Addressing critical gaps in conventional approaches, the framework incorporates landslide typology-specific modeling (soil, debris, rock) and infrastructure vulnerability quantification. A geospatial database of 5,200 historical landslides and 16 conditioning factors (CgFs) was optimized through hyper-parameter tuning via the Random Search (RS) method, enhancing model generalizability. The XGB-RFE model achieved superior predictive accuracy, validated through repeated cross-validation, with a peak area under the curve (AUC) of 0.996 for total landslides and 0.917–0.990 for typology-specific assessments, identifying slope, land use/land cover (LULC), topographic wetness index (TWI), and road proximity as dominant predictors. Geospatial analysis classified 38%–51% of the study area as Very High susceptibility, concentrated in the northern and northwestern zones of the study area characterized by steep slopes and dense infrastructure. Integration of Google Open Buildings data with landslide hazard assessments enabled the development of Uttarakhand first landslide vulnerability-building map, showing that 30.06% of 372,412 structures (112,000+ buildings) are located in high-risk zones. These results offer practical insights for disaster risk reduction and infrastructure planning, supporting policymakers in formulating proactive, data-driven strategies to enhance resilience in landslide-prone mountain regions.</p> Graphical Abstract <p>This study aims to determine landslide susceptibility (LS), analyse the evolution of terrain instability, and assess infrastructure vulnerability across high-risk zones in Uttarakhand, India. The graphical abstract illustrates the integration of hybrid Machine Learning (ML) models including Extreme Gradient Boosting (XGB), Stochastic Gradient Boosting (SGB), and Rotated Random Forest (RRF), with wrapper-based Recursive Feature Elimination (RFE) to develop high-precision Landslide Susceptibility Maps (LSMs). A comprehensive geospatial information systems (GIS) database of historical landslide occurrences was compiled, and 16 conditioning factors (CgFs) were extracted, including altitude, slope, aspect, curvature, geology, distance to roads and streams, precipitation, land use/land cover (LULC), normalized difference vegetation index (NDVI), soil type, topographic wetness index (TWI), sediment power index (SPI), terrain ruggedness index (TRI), and lineament density. These factors underwent correlation and multicollinearity analysis, followed by RFE with k-fold cross-validation to identify the most relevant features. The selected features were then used to train hybrid ensemble models through adaptive cross-validation, with hyperparameter optimization conducted via the Random Search (RS) method. Model performance was assessed using AUC metrics, where the XGB-RFE model achieved the highest predictive accuracy (AUC = 0.996 for total landslides and 0.917–0.990 for typology-specific models). Key predictors included slope, LULC, TWI, and road proximity. Spatial analysis revealed that 38%–51% of the study area falls under Very High susceptibility, concentrated mainly in the northern and northwestern zones of the study area characterized by steep slopes and dense infrastructure. The study also integrated Google Open Buildings data to produce the first vulnerability-building map for the region, identifying that over 30% of 372,412 structures are located within high-risk zones. These findings offer critical perceptions for land-use planning, infrastructure resilience, and disaster risk reduction in Himalayan environments, contributing to data-driven decision-making for slope stability and regional safety.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

GeoRisk Intelligence: Hybrid Ensemble Data-Driven Models with Recursive Feature Elimination for Landslide Susceptibility and Infrastructure Vulnerability in Uttarakhand

  • Alireza Habibi Khouzani,
  • Chiranjit Singha,
  • Armin Moghimi,
  • Mahmoud Reza Delavar

摘要

Landslides in mountainous regions require advanced predictive frameworks to mitigate escalating risks to communities and infrastructure. This study has proposed a novel hybrid ensemble Machine Learning (ML) models integrating Extreme Gradient Boosting (XGB), Stochastic Gradient Boosting (SGB), and Rotated Random Forest (RRF) with wrapper-based Recursive Feature Elimination (RFE) algorithm for landslide susceptibility mapping (LSM) in Uttarakhand, India, a Himalayan region chronically vulnerable to slope failures. Addressing critical gaps in conventional approaches, the framework incorporates landslide typology-specific modeling (soil, debris, rock) and infrastructure vulnerability quantification. A geospatial database of 5,200 historical landslides and 16 conditioning factors (CgFs) was optimized through hyper-parameter tuning via the Random Search (RS) method, enhancing model generalizability. The XGB-RFE model achieved superior predictive accuracy, validated through repeated cross-validation, with a peak area under the curve (AUC) of 0.996 for total landslides and 0.917–0.990 for typology-specific assessments, identifying slope, land use/land cover (LULC), topographic wetness index (TWI), and road proximity as dominant predictors. Geospatial analysis classified 38%–51% of the study area as Very High susceptibility, concentrated in the northern and northwestern zones of the study area characterized by steep slopes and dense infrastructure. Integration of Google Open Buildings data with landslide hazard assessments enabled the development of Uttarakhand first landslide vulnerability-building map, showing that 30.06% of 372,412 structures (112,000+ buildings) are located in high-risk zones. These results offer practical insights for disaster risk reduction and infrastructure planning, supporting policymakers in formulating proactive, data-driven strategies to enhance resilience in landslide-prone mountain regions.

Graphical Abstract

This study aims to determine landslide susceptibility (LS), analyse the evolution of terrain instability, and assess infrastructure vulnerability across high-risk zones in Uttarakhand, India. The graphical abstract illustrates the integration of hybrid Machine Learning (ML) models including Extreme Gradient Boosting (XGB), Stochastic Gradient Boosting (SGB), and Rotated Random Forest (RRF), with wrapper-based Recursive Feature Elimination (RFE) to develop high-precision Landslide Susceptibility Maps (LSMs). A comprehensive geospatial information systems (GIS) database of historical landslide occurrences was compiled, and 16 conditioning factors (CgFs) were extracted, including altitude, slope, aspect, curvature, geology, distance to roads and streams, precipitation, land use/land cover (LULC), normalized difference vegetation index (NDVI), soil type, topographic wetness index (TWI), sediment power index (SPI), terrain ruggedness index (TRI), and lineament density. These factors underwent correlation and multicollinearity analysis, followed by RFE with k-fold cross-validation to identify the most relevant features. The selected features were then used to train hybrid ensemble models through adaptive cross-validation, with hyperparameter optimization conducted via the Random Search (RS) method. Model performance was assessed using AUC metrics, where the XGB-RFE model achieved the highest predictive accuracy (AUC = 0.996 for total landslides and 0.917–0.990 for typology-specific models). Key predictors included slope, LULC, TWI, and road proximity. Spatial analysis revealed that 38%–51% of the study area falls under Very High susceptibility, concentrated mainly in the northern and northwestern zones of the study area characterized by steep slopes and dense infrastructure. The study also integrated Google Open Buildings data to produce the first vulnerability-building map for the region, identifying that over 30% of 372,412 structures are located within high-risk zones. These findings offer critical perceptions for land-use planning, infrastructure resilience, and disaster risk reduction in Himalayan environments, contributing to data-driven decision-making for slope stability and regional safety.