Leveraging Pre-existing Geological Model to Generate Multiple Realizations of Geological Domain Through Machine Learning Algorithms
摘要
Spatial modeling of geological domains and the assessment of uncertainty in the resulting models are essential steps in mineral resource estimation. Geostatistical simulation algorithms, such as plurigaussian simulation and sequential indicator simulation, have been widely used for modeling geological domains. However, despite their widespread acceptance, these algorithms face limitations, particularly when applied to non-stationary geological domains and datasets with limited sample sizes. Additionally, the successful application of these approaches requires expert knowledge, as numerous parameters must be inferred to ensure accurate algorithm execution. Machine learning (ML) models present an interesting alternative for modeling geological domains as they can manage complex relationships between variables and do not require extensive parameter inference. However, the performance of ML models depends on the amount of available data. In this study, a pre-existing interpreted model of geological domains is used to address the issue of limited data availability. The interpreted model is considered as soft data, providing prior geological knowledge that guides the generation of different datasets. These datasets are then individually used to produce realizations of geological domains using a data-driven approach. Three ML algorithms—K-nearest neighbors (KNN), random forest (RF), and artificial neural networks (ANN)—were applied to classify five geological domains in a copper porphyry deposit. With ANN and KNN, soft data integration enhanced classification accuracy for all geological domains. For RF models, incorporating soft data improved the classification of all domains except dikes. Apart from the classification of dikes using the RF model, the results demonstrate that even a small proportion of soft data can improve the classification accuracy of ML models. The cross-validation process demonstrates that incorporating 1% of the interpreted model (treated as soft data) alongside hard data provides an optimal proportion of additional data, yielding better results compared to other tested proportions. Furthermore, the same cross-validation process shows that the KNN algorithm outperforms other methods in capturing the uncertainty associated with the spatial layout of domains. Overall, the proposed approach offers a robust alternative to geostatistical simulation algorithms for modeling geological domains, especially in early stages when data, such as drill hole information, are scarce, and an interpreted model of geological domains is available.