Impact of data size on ML-based prediction of shear and compressional slowness
摘要
Accurate prediction of compressional (P-wave) and shear (S-wave) velocity is vital for structural, geomechanical, and petrophysical analyses of subsurface formations. Since velocity is typically measured as its reciprocal, slowness is the standard parameter recorded in sonic logs and is the focus of this study. This work investigates the predictive performance of three machine learning (ML) models—Artificial Neural Network (ANN), Support Vector Regression (SVR), and Adaboost regressor—using well-log data of varying sizes. The dataset includes true resistivity (RT), density (RHOB), neutron porosity (NPHI), gamma-ray (GR), P-wave slowness (DTC), S-wave slowness (DTS), and photoelectric (PEF) logs from nineteen wells in the Sarvak Formation, a well-known carbonate reservoir in southern Iran. Results indicate that ANN consistently outperformed SVR and Adaboost in DTS prediction, achieving the highest accuracy with R2 values of 0.8956–0.9192 and RMSE values of 2.752–2.352 µs/ft when using DTC, NPHI, GR, RHOB, and PEF as input features for a dataset containing 3654 data points from three wells. SVR demonstrated competitive performance for larger datasets (R2 = 0.9185, RMSE = 2.497 µs/ft) but showed high sensitivity to dataset size, underperforming with limited data. Adaboost, while improving steadily with increasing data, remained the least accurate (R2 = 0.875, RMSE = 3.0 µs/ft). For DTC prediction, ANN again showed superior accuracy, with R2 values ranging from 0.798 to 0.9101 and RMSE decreasing from 2.82 µs/ft (2,896 data points) to 1.841 µs/ft (16,550 data points), demonstrating the positive impact of data availability on model accuracy. SVR exhibited competitive performance for larger datasets (R2 = 0.798–0.9101, RMSE = 3.442–2.171 µs/ft), while Adaboost remained the least effective across all cases (R2 < 0.88, RMSE = 3.546–2.461 µs/ft). A threshold of ~ 12,650 data points was identified, beyond which additional data yielded diminishing returns in model performance. Additionally, incorporating optimal input parameters, such as PEF and RT logs, significantly improved prediction accuracy, while less critical features (e.g., resistivity in DTS prediction) contributed minimally. These findings highlight the critical role of dataset size, input parameter selection, and model choice in optimizing ML applications for subsurface characterization, providing practical guidelines for future geophysical studies.