Computational Methodology to Study Heterogeneities in Petroleum Reservoirs
Bibliographic record
Abstract
Abstract Characterization of hydrocarbon reservoirs is strategically important to define the productivity of oil and/or gas fields. It involves many challenges such as appropriate identification, classification and interpretation of diagenetic processes that directly affect the quality of the reservoirs. Proper studies demand integration and analysis of very large amounts of data, usually presenting high-dimensional feature spaces. Current methods have many manual steps leading to a limited exploration of the data. These challenges are being intensified due to the need of knowledge and time dedication from experts. We developed a novel methodology that combines established techniques, such as Principle Component Analysis (PCA), clustering methods, parallel coordinates and scatter plots, with features such as dynamic (magic) lenses – filter and shadow lenses –, axes reordering and color maps, to automatically perform reservoir characterization in order to assist the identification, validation and interpretation of petrofacies. Petrofacies is a set of petrographic characteristics of microscopic order which allow the analyst to understand the diagenetic processes, aiding in the evaluation of the potential for hydrocarbon storage in the reservoir. We have applied our methodology on several databases from different sedimentary basins – Espirito Santo and Parana basins (Brazil), Talara Basin (Peru) and Niger Delta Basin (Nigeria). We conclude that our method allows the analyst to gain insights about the entire database in a manner that is faster than the analysis using a manual method. It also allows validation of the results because it is a powerful tool that can qualitatively and quantitatively support the analyst in the identification, interpretation and validation of petrofacies. This new methodology can optimize data analysis of similar databases, accelerating the analysis and reducing the committed work by the experts.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.007 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".