Hierarchical automated machine learning (AutoML) for advanced unconventional reservoir characterization
Bibliographic record
Abstract
Abstract Recent advances in machine learning (ML) have transformed the landscape of energy exploration, including hydrocarbon, CO2 storage, and hydrogen. However, building competent ML models for reservoir characterization necessitates specific in-depth knowledge in order to fine-tune the models and achieve the best predictions, limiting the accessibility of machine learning in geosciences. To mitigate this issue, we implemented the recently emerged automated machine learning (AutoML) approach to perform an algorithm search for conducting an unconventional reservoir characterization with a more optimized and accessible workflow than traditional ML approaches. In this study, over 1000 wells from Alberta’s Athabasca Oil Sands were analyzed to predict various key reservoir properties such as lithofacies, porosity, volume of shale, and bitumen mass percentage. Our proposed workflow consists of two stages of AutoML predictions, including (1) the first stage focuses on predicting the volume of shale and porosity by using conventional well log data, and (2) the second stage combines the predicted outputs with well log data to predict the lithofacies and bitumen percentage. The findings show that out of the ten different models tested for predicting the porosity (78% in accuracy), the volume of shale (80.5%), bitumen percentage (67.3%), and lithofacies classification (98%), distributed random forest, and gradient boosting machine emerged as the best models. When compared to the manually fine-tuned conventional machine learning algorithms, the AutoML-based algorithms provide a notable improvement on reservoir property predictions, with higher weighted average f1-scores of up to 15–20% in the classification problem and 5–10% in the adjusted-R2 score for the regression problems in the blind test dataset, and it is achieved only after ~ 400 s of training and testing processes. In addition, from the feature ranking extraction technique, there is a good agreement with domain experts regarding the most significant input parameters in each prediction. Therefore, it is evidence that the AutoML workflow has proven powerful in performing advanced petrophysical analysis and reservoir characterization with minimal time and human intervention, allowing more accessibility to domain experts while maintaining the model’s explainability. Integration of AutoML and subject matter experts could advance artificial intelligence technology implementation in optimizing data-driven energy geosciences.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".