A deep learning framework for production forecasting
Bibliographic record
Abstract
Over the last decades conventional decline curve analysis has been a common method to forecast hydrocarbon well production. However, this technique is often faced with inefficiencies in reliably predicting hydrocarbon production. This is partially due to its inability to predict multiple phases in a single model, and to incorporate operational changes and subsurface fluid flow mechanisms. In this paper, we propose a deep learning framework that addresses the limitations of traditional parametric DCA. This was achieved by using an Encoder-Decoder Long Short-Term Memory Networks architecture. The model combines three types of inputs consisting of time-variant information (i.e., historical production), static well features (e.g., geology and spacing parameters), or known-in-advance control variables (e.g., artificial lift type as a function of time), and the output is a multi-step forecasts for oil, gas, and water rates. We applied the model to a dataset of 213 gas wells from the Eagle Ford Basin, with the goal of predicting the gas and water rates. Production history ranges from 11 to 52 timesteps (mean 31) and is supplemented with 18 static features comprising geology, spacing, and operational attributes. Our proposed framework demonstrated the ability to; 1) forecast future production with limited historical production, 2) predict production behavior under different control regimes and quantify the impact of future planned activities, 3) forecast multiple phases (oil, gas, and water) simultaneously with the same model (i.e., multi-target prediction).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".