Machine learning–based characterization of shale formations using core analysis data
摘要
Reliable productivity prediction in shale reservoirs is frequently limited by the scarcity of comprehensive reservoir property data. This study presents a data-driven framework that utilizes group-specific characteristics derived from core analysis data to support production-related assessment, rather than direct estimation of reservoir properties. Core analysis data comprising 26,402 samples from 714 Montney shale wells were analyzed to identify geological variables exhibiting nonlinear relationships. Principal component analysis was employed to reduce dimensionality and to select five key factors: easting, northing, average sample depth, porosity, and permeability. These features were subsequently used to classify the core samples into three distinct groups through clustering analysis. Because porosity and permeability data are often unavailable in field-scale applications, eight supervised machine learning models were developed to classify samples into the identified groups using only well information (easting, northing, and average sample depth). Model performance was evaluated using independent training and test datasets constructed at the well level. The ensemble bagging tree model demonstrated robust classification performance, achieving average accuracies exceeding 90%. The results indicate that quantitatively defined group-specific characteristics derived from core analysis data can be systematically incorporated as surrogate inputs for prediction models based on publicly available well information. In addition, the compiled core analysis database provides a quantitative basis for subsequent studies on shale productivity assessment.