The influence of non-landslide sample selection on susceptibility result: a case study in the Hanbing District, Shaanxi Province, China
摘要
Landslide susceptibility research is critical for risk management, but current research predominantly emphasizes advanced intelligent algorithms and overlooks non-landslide sample quality. This study focuses on Hanbing District as the study area where 524 landslides were documented. Building on the susceptibility map generated by the frequency ratio (FR) model, an innovative non-landslide sampling strategy was developed, where the sample quantity inversely correlates with the area of FR-derived susceptibility levels. Meanwhile, three comparative approaches were established: (1) random selection across the entire study area; (2) similar to our proposed method, but applied to the landslide distribution density map; and (3) random selection in low susceptibility zones. Subsequently, the landslide susceptibility evaluation datasets were established and randomly divided into training and validation datasets with a 70:30 ratio, and the extreme gradient boosting (XGBoost) algorithm was introduced to build susceptibility models. Results demonstrated that our proposed sampling method achieved superior performance, with the area under the receiver operating characteristic curve (AUC) reaching 0.873. Additionally, the results also identified that nearly 90% of existing landslides are distributed in the Yuehe Basin and Hanjiang River areas. Finally, the Shapley Additive exPlanations (SHAP) revealed that regions characterized by paddy fields, low altitude (below 655 m), and sparse vegetation (normalized differential vegetation index below 0.65) are prone to landslides. This study provides valuable insight for non-landslide sample selection, and the susceptibility map can help landslide risk management for the study area.