Integrated Data Mining and Optimization in Hydraulic Fractured Tight Oil Reservoirs
Bibliographic record
Abstract
Unconventional tight oil reservoirs have emerged as one of most important hydrocarbon resources due to the advanced horizontal-well drilling and multi-stage stimulation techniques. Optimizing well placement and fracture treatment design is critical to maximize well productivity and net present value (NPV) in such reservoirs. Using reservoir simulation to optimize performance of stimulation strategies is very computationally expensive due to extensive complex processes of hydraulic fracturing. With rapid development of unconventional reservoirs, the amount of data related to hydraulic fracturing process and well production is accumulating rapidly. Therefore, it is of practical interest for reservoir engineers to develop data-driven models to optimize the well performance in tight formations. In this study, a comprehensive data mining process is developed and successfully applied to evaluate the well performance in the Montney Formation, Western Canada. 6 features are identified as the most important variables by using the recursive feature elimination with cross validation (RFECV) method. Based on these features, four commonly used supervised learning approaches including random forest (RF), adaptive boosting (AdaBoost), support vector machine (SVM), and neural network (NN) are evaluated to predict the first-year oil production in the Montney Formation. It is found that the RF provides an accurate and robust production forecasting model in comparison with other three methods. Then, the applicability of the deep neural networks (DNNs) is evaluated to predict the early production in the Bakken tight Formation. An optimum DNN model is obtained by optimizing the hyperparameters of the DNN models. In addition, the recurrent neural networks (i.e., long short-term memory LSTMs) are applied to learn the pressure transient behavior in the stress-sensitive reservoirs. It is found that the developed LSTM models are capable of discovering relationships between the flow rate and pressure through data mining process without prior knowledge of physical models. Finally, a global optimization framework based on generalized differential evolution (GDE) algorithm is developed and successfully applied to optimize the production performance of multi-well pad in the Cardium tight oil formation. The optimization process integrates the available field data into the optimization framework to obtain a practical optimum scenario for the multi-well pad development.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".