Applicability of decision tree-based machine learning models in the prediction of core-calibrated shale facies from wireline logs in the late Devonian Duvernay Formation, Alberta, Canada
Bibliographic record
Abstract
Abstract Well logs provide insight into stratigraphically compartmentalized rock properties and are a cost-effective alternative to core. The identification of reservoir (and nonreservoir) facies in core, and their calibration to well-log response has traditionally relied on expert domain knowledge and is inherently inconsistent. Such analyses are time-consuming, tedious, error prone, and often biased due to a lack of objectivity. Automated lithologic interpretations from wireline logs appear to be a promising solution for identifying and understanding depositional complexity within a reservoir. Using the Duvernay Formation in the Western Canada Sedimentary Basin as a case study, the authors evaluate the applicability of decision tree-based machine learning (ML) methods in the prediction of core-calibrated facies and/or facies association distributions within wireline logs. The authors use three independent decision tree-based ML models to predict (1) facies (FACM), (2) facies associations (FAM), and (3) reservoir rock (RESM) from wireline logs. Model accuracies are 60.3%, 88.1%, and 88.1% for FACM, FAM, and RESM, respectively, but individual class F1 scores range from 0 to 0.92. The authors attribute discrepancies in individual class performance to interval thickness, sample proportion of training data, and distinguishability of the output class. Classes thicker than 3 m and encompassing at least 16% of the training data set have F1 scores greater than 0.60. The authors attribute exceptions to these general cutoffs to the ability to recognize diagnostic sedimentologic features observed in core. Results from this study help in understanding stratigraphic complexity in the absence of core aiding in subsurface characterization of reservoirs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".