Automatic Fracture Identifications From Image Logs With Machine-Learning Approaches: A Contest Summary
Bibliographic record
Abstract
Borehole image logs are essential for characterizing subsurface formations, particularly in identifying fractures that influence reservoir behavior and productivity. Manual interpretation of such logs, however, remains time consuming and susceptible to subjectivity and inconsistency. To address these challenges, the SPWLA Petrophysical Data-Driven Analytics special interest group (PDDA SIG) launched its 4th Annual Machine-Learning Competition, aimed at developing automated methods for accurate, efficient, and reproducible fracture detection. The competition utilized resistivity image logs and conventional quadruple-combo logs from eight wells in the Western Canadian Sedimentary Basin (WCSB), accompanied by expert-labeled fracture annotations for training. A separate blind test data set from two additional wells in the same basin was reserved for final evaluation. Participants were provided with a Jupyter Notebook containing preprocessed data and a baseline framework to facilitate model development. Submissions were evaluated based on F1 score and averaged root mean squared error (RMSE) on the blind test set, reflecting both classification accuracy and predictive reliability. This paper reviews the top five approaches submitted, highlighting key methodologies, feature engineering strategies, and model architectures that led to improved fracture detection performance. Our results demonstrate that advanced machine-learning techniques can substantially enhance the consistency and accuracy of fracture identification from borehole image logs. These findings support the integration of data-driven solutions into petrophysical workflows, offering scalable and objective tools to augment or replace manual interpretation, ultimately improving decision making in exploration and production operations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".