<p>The Sembar Formation is a key petroliferous unit in the Lower Indus Basin of Pakistan; however, its resource potential remains uncertain due to limited and discontinuous Total Organic Carbon (TOC) and Rock-Eval pyrolysis data. Accurate assessment of TOC, volatile hydrocarbons (S<sub>1</sub>), remaining hydrocarbons (S<sub>2</sub>), carbon dioxide yield (S<sub>3</sub>), maximum pyrolysis temperature (T<sub>max</sub>), and Vitrinite Reflectance (%R<sub>o</sub>) is essential but challenging because laboratory analyses are costly, time-consuming, and lack continuous data coverage. This study proposes a machine learning (ML)-based approach to predict TOC and Rock-Eval pyrolysis parameters from conventional well logs in the Sembar Formation. Six ML algorithms, Decision Tree (DT), Random Forest (RF), K-Nearest Neighbors (KNN), Gradient Boosting Regressor (GBR), Adaptive Gradient Boosting (AGB), and Extreme Gradient Boosting (XGB), were applied to data from 21 wells to predict the target parameters. The results demonstrate strong predictive performance, with the AGB model showing the best results and achieving correlation coefficients of 0.92 for TOC, 0.89 for S<sub>1</sub>, 0.894 for S<sub>2</sub>, 0.89 for S<sub>3</sub>, and 0.94 for T<sub>max</sub>. The TOC values range from 1.05 to 4.15 wt%, indicating good hydrocarbon generation potential. Kerogen analysis shows a predominance of Type III organic matter, indicating gas-prone characteristics. Vitrinite reflectance (%R<sub>o</sub>: 0.40–1.44%) indicates that the Sembar Formation is predominantly within the oil to wet gas window, while T<sub>max</sub> (420–480&#xa0;°C) supports this thermal maturity level. These findings confirm that ML models reliably replicate laboratory-derived parameters, enabling continuous source rock evaluation. The proposed approach provides a cost-effective alternative to conventional log-based methods, offering improved prediction accuracy for identifying favorable resource zones in similar data-limited settings.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comprehensive evaluation of the source rock potential of the Sembar Formation: a machine learning-aided approach in the Lower Indus Basin, Pakistan

  • Muhammad Abid,
  • Zeeshan Tariq,
  • Feng Zhu,
  • Jing Ba,
  • Muhsan Ehsan,
  • Uti Ikitsombika Markus,
  • Muhammad Mudasir

摘要

The Sembar Formation is a key petroliferous unit in the Lower Indus Basin of Pakistan; however, its resource potential remains uncertain due to limited and discontinuous Total Organic Carbon (TOC) and Rock-Eval pyrolysis data. Accurate assessment of TOC, volatile hydrocarbons (S1), remaining hydrocarbons (S2), carbon dioxide yield (S3), maximum pyrolysis temperature (Tmax), and Vitrinite Reflectance (%Ro) is essential but challenging because laboratory analyses are costly, time-consuming, and lack continuous data coverage. This study proposes a machine learning (ML)-based approach to predict TOC and Rock-Eval pyrolysis parameters from conventional well logs in the Sembar Formation. Six ML algorithms, Decision Tree (DT), Random Forest (RF), K-Nearest Neighbors (KNN), Gradient Boosting Regressor (GBR), Adaptive Gradient Boosting (AGB), and Extreme Gradient Boosting (XGB), were applied to data from 21 wells to predict the target parameters. The results demonstrate strong predictive performance, with the AGB model showing the best results and achieving correlation coefficients of 0.92 for TOC, 0.89 for S1, 0.894 for S2, 0.89 for S3, and 0.94 for Tmax. The TOC values range from 1.05 to 4.15 wt%, indicating good hydrocarbon generation potential. Kerogen analysis shows a predominance of Type III organic matter, indicating gas-prone characteristics. Vitrinite reflectance (%Ro: 0.40–1.44%) indicates that the Sembar Formation is predominantly within the oil to wet gas window, while Tmax (420–480 °C) supports this thermal maturity level. These findings confirm that ML models reliably replicate laboratory-derived parameters, enabling continuous source rock evaluation. The proposed approach provides a cost-effective alternative to conventional log-based methods, offering improved prediction accuracy for identifying favorable resource zones in similar data-limited settings.