Updated Screening Criteria for Steam Flooding Based on Oil Field Projects Data
Bibliographic record
Abstract
Abstract Enhanced oil recovery (EOR) screening is considered the first step in evaluating potential EOR techniques for candidate reservoirs. Therefore, as new technologies are developed, it is important to update the screening criteria. Many of the screening criteria for steam flooding that have been described in the literature were based on data collected from EOR surveys biennially published in the Oil & Gas Journal. However, these datasets contain some problems, including outliers, missing data, inconsistent data and duplicate data, that could affect the accuracy of the results. Despite the importance of ensuring the quality of a dataset before running analyses, data quality has not been addressed in previous research related to EOR screening criteria. The objective of this current work was to update the screening criteria for steam flooding by using a database that had been cleaned. The original dataset included 1, 785 steam flooding field projects from around the world (Brazil, Canada, China, Colombia, Congo, France, Germany, Indonesia, Trinidad, U.S. and Venezuela). These projects had been reported in the Oil and Gas Journal from 1980 to 2012. After detecting and deleting the duplicate projects, only 626 field projects remained. To analyze and describe the results of the dataset, both graphical and statistical methods were used. A box plot and cross plots were used to detect and identify data problems, allowing for the removal of outliers and inconsistent data. Histogram distributions and box plots were used to show the distribution of each parameter and present the range of the dataset. New screening criteria were developed based on these statistics and the defined data parameters. The developed criteria were com-pared with previously published criteria, and their differences are explained in this paper.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.057 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.016 | 0.009 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".