Leveraging petrophysical and geological constraints for AI-driven predictions of total organic carbon (TOC) and hardness in unconventional reservoir prospects
Bibliographic record
Abstract
Key parameters for evaluating shale reservoirs include total organic carbon (TOC), thermal maturity, and hardness, the latter influencing fracture development and being crucial for managing ultralow permeability reservoirs. These parameters are often derived from costly, time-consuming core sample analyses and may be limited in availability. Recently, machine learning (ML) and deep learning (DL) have effectively predicted TOC and hardness from well logs but often require large datasets and lack integration with petrophysical and geological constraints. This study examines the impact of incorporating these constraints on prediction accuracy using four manually fine-tuned ML algorithms: Random Forest (RF), Support Vector Regression (SVR), XGBoost (XGB), and Artificial Neural Network (ANN). Data from five wells in the Horn River Basin (HRB) comprising 6366 data points were analyzed, with TOC and hardness values for 612 and 3492 points, respectively. Petrophysical constraints were derived from triple combo well logs (gamma ray, bulk density, neutron porosity), while geological constraints included stratigraphic data or spatial distance between training and target wells—petrophysical constraints most improved predictions, while stratigraphic and spatial constraints had progressively less impact. Our optimized models achieved R2 (coefficient of determination) of 0.89 and RMSE (root-mean-square error) of 0.47 for TOC predictions and 0.90 and 34.8 for hardness predictions, reducing RMSE by up to 13.52% compared to the unconstrained model. The XGB algorithm emerged as the best choice, and integrating domain knowledge transforms a data-driven method into a scientifically driven one, enhancing prediction accuracy and aligning model predictions with petrophysical and geological intricacies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".