Development and comparison of reduced-order models for CO2-enhanced oil recovery predictions
Bibliographic record
Abstract
CO 2 -enhanced oil recovery (CO 2 -EOR) has been a mature and promising technology since the 1970s, offering a dual solution for energy production and carbon sequestration. Recent advances in reduced-order models (ROMs) using empirical analysis and artificial intelligence (AI) tools handle complex data efficiently. However, existing ROMs for CO 2 -EOR often lack validation due to data confidentiality or are too case-specific for broader application. This paper introduces a framework to close these gaps, enabling the development and consistent comparison of generalized ROMs for CO 2 -EOR with carbon capture and storage (CCS), even with traditional tools. A synthesis dataset (∼3000 runs) was established to develop ROMs, which were validated using field data from both EOR and CCS perspectives. Three key findings are revealed. First, normalizing outputs with respect to CO 2 utilization illustrated a direct relationship between CCS and EOR. Second, generalized statistics-based ROMs reduced input complexity and validated field data but predicted fewer outputs. Machine learning-based ROMs predicted more outputs, supporting field operational decision-makings. Last, ROMs were particularly suitable for early-stage, large-scale CO 2 -EOR assessments. This study extended the boundaries of developing generalized ROMs for CO 2 -EOR and identified pros and cons across modeling approaches, contributing to net-zero goals and advancing sustainable and affordable energy future. • The stats- and ML-ROMs are developed and compared for CO 2 -EOR with validation on Weyburn oil field production profile. • Stats-ROM predictions reduced input complexity, while ML-ROMs can incorporate geological properties. • ROMs offer time-efficient, scalable solutions for early-stage CO 2 -EOR with fewer inputs. • ROMs can support early-stage economic and environmental assessments at regional and national levels.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".