Proxy modeling for CO2 injection in tight gas condensate reservoirs using active learning-based artificial intelligence
Bibliographic record
Abstract
This study proposes active learning-based artificial intelligence application to efficiently build a proxy model for carbon dioxide (CO2) injection scenarios in tight gas condensate reservoirs. In gas condensate reservoirs, as production progresses and reservoir pressure decreases, condensate accumulation in the reservoir pores leads to a relative reduction in gas permeability, thereby lowering gas productivity. Injecting CO2 into the target gas condensate reservoirs to maintain pressure can mitigate condensate banking while simultaneously enabling CO2 geological storage. However, multiple variables influence the performance of such CO2 injection strategies. In this research, proxy modeling for a tight gas condensate reservoir mimicking the Montney region in Canada was performed using active learning, which optimizes the acquirement of additional training data. The proxy model was constructed with a random forest algorithm trained on reservoir simulation results generated using Petrel, Eclipse, and MEPO software from SLB. Initially, simulations were conducted for a limited number of scenarios, and additional data were iteratively acquired by identifying input scenarios with high uncertainty in predictions from the previous proxy model. This active learning process improves the efficiency when adding extra training dataset, enhancing the model's performance while reducing the need for exhaustive simulations. The input parameters for CO2 injection included the timing of switching a production well to an injection well, the bottomhole pressure of an injection well, and the maximum production rate. Output parameters included CO2 molar injection and production rates, field gas and oil production totals, field oil saturation averages, field gas injection cumulative total, CO2 storage total, and field average pressure. Experiments analyzed the minimum additional data required to achieve an R2 score of 0.95, with initial datasets of 30, 40, 50, and 60 simulations. For these initial dataset sizes, the active learning method saved an average of 4, 6, 3, and 1 reservoir simulations, respectively. Considering that each reservoir simulation requires an average of 45 minutes, the computational cost savings are significant. This efficiency is expected to be even greater for more complex reservoir simulations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".