Well GOR Prediction from Surface Gas Composition in Shale Reservoirs
Bibliographic record
Abstract
Abstract There are increased development activities in shale reservoirs with ultra-low permeability thanks to the advances in drilling and fracking technology. However, representative reservoir fluid samples are still difficult to acquire. The challenge leads to limited reservoir fluid data and large uncertainties for shale play evaluation, field development, and production optimization. In this work, we built a large unconventional reservoir fluid database with more than 2400 samples from shale reservoirs in Canada, Argentina, and the USA, comprising early production surface gas data and traditional PVT data from selected shale assets. A machine learning approach was applied to the database to predict gas to oil ratio (GOR) in shale reservoirs. To enhance regional correlations and obtain a more accurate GOR prediction, we developed a machine learning model focused on Canada shale plays data, intended for wells with limited reservoir fluid data available and located within the same region. Both surface gas compositional data and well location and are input features to this model. In addition, we developed an additional machine learning model for the objective of a generic GOR prediction model without shale dependency. The database includes Canada shale data and Argentina and USA shale data. The GOR predictions obtained from both models are good. The machine learning model circumscribed to the Canada shale reservoirs has a mean percentage error (MAPE) of 4.31. In contrast, the generic machine learning model, which includes additional data from Argentina and USA shale assets, has a MAPE of 4.86. The better accuracy of the circumscribed Canada model is due to the introduction of the geospatial well location to the model features. This study confirms that early production surface gas data can be used to predict well GOR in shale reservoirs, providing an economical alternative for the sampling challenges during early field development. Furthermore, the GOR prediction offers access to a complete set of reservoir fluid properties which assists the decision-making process for shale play evaluation, completion concept selection, and production optimization.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".