Application of Machine Learning for Shale Oil and Gas “Sweet Spots” Prediction
Bibliographic record
Abstract
With the continuous improvement of shale oil and gas recovery technologies and achievements, a large amount of geological information and data have been accumulated for the description of shale reservoirs, and it has become possible to use machine learning methods for “sweet spots” prediction in shale oil and gas areas. Taking the Duvernay shale oil and gas field in Canada as an example, this paper attempts to build recoverable shale oil and gas reserve prediction models using machine learning methods and geological and development big data, to predict the distribution of recoverable shale oil and gas reserves and provide a basis for well location deployment and engineering modifications. The research results of the machine learning model in this study are as follows: ① Three machine learning methods were applied to build a prediction model and random forest showed the best performance. The R2 values of the built recoverable shale oil and gas reserves prediction models are 0.7894 and 0.8210, respectively, with an accuracy that meets the requirements of production applications; ② The geological main controlling factors for recoverable shale oil and gas reserves in this area are organic matter maturity and total organic carbon (TOC), followed by porosity and effective thickness; the main controlling factor for engineering modifications is the total proppant volume, followed by total stages and horizontal lateral length; ③ The abundance of recoverable shale oil and gas reserves in the central part of the study area is predicted to be relatively high, which makes it a favorable area for future well location deployment.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".