Comprehensive evaluation of tree-based machine learning algorithms for soil classification from investigative drilling data
摘要
Investigative drilling (ID) is a measurement while drilling technique increasingly applied in site investigations in Australia, yet its use for geotechnical soil classification remains underexplored compared to its application in predicting rock formations in mining and petroleum engineering. Leveraging the efficacy of machine learning in handling complex data, this study evaluates the performance of three tree-based algorithms—Random Forest (RF), XGBoost (XGB), and LightGBM (LGB)—for soil classification using ID data. A comprehensive assessment was conducted across binary, three-class, and four-class classification tasks, considering predictive accuracy, computational cost, feature importance, and imbalance-handling techniques. Results show that tuned XGB models achieved the highest predictive accuracy, while default RF models delivered nearly comparable performance with much lower complexity. LGB demonstrated superior runtime efficiency, but its advantage diminished as the number of classes increased. Among all the drilling parameters, penetration rate and rotation pressure were consistently found to be the most influential, whereas feed pressure and borehole diameter contributed the least to soil-type prediction. Application to two unseen boreholes confirmed the reliability of the default RF model, particularly in fine-dominant ground conditions, with undersampling outperforming oversampling methods in addressing data imbalance. These findings demonstrate the potential of ID data for effective soil classification using tree-based models and highlight opportunities for developing more robust, real-world-ready models through advanced data augmentation, broader datasets, and expanded hyperparameter optimization.