Machine learning-based soil depth estimation from the diverse landscapes of Nagaland, North-Eastern Himalayas
摘要
Understanding soil depth is fundamental for evaluating a soil’s ability to support vegetation, retain water and nutrients, sequester carbon, and sustain infrastructure. Despite its importance, the spatial variability of soil depth remains poorly understood, especially in mountainous terrains with limited sampling. Digital Soil Mapping (DSM), empowered by Machine Learning (ML) algorithms, offers a practical solution for estimating soil properties across complex landscapes. This study aimed to develop a predictive soil depth map for Nagaland, a hilly state in Northeastern India, using ML techniques. A dataset comprising soil depth observations from 70 locations and 37 environmental covariates was analysed. Four ML algorithms Random Forest (RF), Support Vector Machine (SVM), Extreme Gradient Boosting (XGB), and K-Nearest Neighbours (KNN) were employed. Covariates derived from MODIS, Worldclim, and digital elevation models (DEM) were used to model soil-landscape relationships. Among all predictors, BIO15 (precipitation seasonality) and NDVI (Normalized Difference Vegetation Index) emerged as key variables influencing soil depth variation. The soil depth predictions across models ranged between 44 and 158 cm, with XGB showing the best performance (R2 = 0.32 ± 0.24; RMSE = 22.63 ± 4.74 cm; MAE = 18.59 ± 3.99 cm), followed by RF, while SVM and KNN were comparatively less accurate. The spatial distribution revealed that deeper soils were concentrated in the central and northern plains and valleys, while shallower soils were predominantly found in the steep, dissected Southeastern Hill Region. The findings highlight the efficacy of ML-driven DSM in mountainous, data-scarce environments, offering a valuable tool for land use planning, soil conservation, and ecological management in the Eastern Himalayas.