A Data Mining Approach to Assess Field Scale CO2 Enhanced Oil Recovery and Sequestration Performance Correlated to Geological and Reservoir Characteristics
Bibliographic record
Abstract
Summary Carbon dioxide (CO2) injection has gained popularity in the petroleum industry as a dual-purpose method for enhanced oil recovery (EOR) and long-term carbon sequestration. However, assessing the performance of CO2 EOR and its storage potential across large-scale fields is a complex task, primarily due to the heterogeneous geological characteristics of reservoirs and the dynamic behavior of injected CO2. Traditional methods for evaluating CO2 injection often rely on manual interpretations or computationally expensive reservoir simulations, both of which can be biased, time-intensive, and less effective for fieldwide analyses involving extensive data sets. In this study, a data mining-driven methodology was developed and applied to one of the most prominent CO2 injection projects in the world. More than 2,000 wells with decades-long production histories were analyzed using advanced statistical and geostatistical approaches, including spatial and temporal normalization of production data. By correlating key production metrics with geological features inferred from the data, fracture-dominated and matrix-dominated regions within the field were identified. The analysis further highlighted zones with differing CO2 injection efficiency and oil displacement behavior, providing a comprehensive understanding of reservoir performance in terms of oil recovery and CO2 sequestration. A critical aspect of the methodology involved combining multiple production metrics—such as gas/oil ratio (GOR), water cut (WCT), time to peak production, and CO2 breakthrough patterns—using Z-score-based normalization across both spatial and temporal domains. This approach enabled localized trend interpretation while maintaining consistency with physical reservoir behavior. Zones where CO2 injection was successful in both enhancing oil recovery and sequestering carbon were differentiated from areas where CO2 rapidly broke through without effective oil displacement, primarily due to fracture orientations and density (less vertically oriented fractures or matrix system dominated reservoir sections). Additionally, regions dominated by vertical fractures, which contributed to long-term CO2 storage, were identified. The results of this work provide valuable insights for optimizing CO2 injection strategies and improving sweep efficiency, ultimately aiding in better decision-making for both enhanced recovery and greenhouse gas sequestration. This novel approach bridges the gap between data-driven analysis and traditional reservoir engineering principles, offering a scalable framework for CO2 EOR operations in fields with complex geologies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".