MétaCan
Menu
Back to cohort
Record W1126901406 · doi:10.2118/174460-ms

Practical Data Mining and Artificial Neural Network Modeling for SAGD Production Analysis

2015· article· en· W1126901406 on OpenAlexaff
Zhiwei Ma, Yaqi Liu, Juliana Y. Leung, Stefan Zanon

Bibliographic record

VenueSPE Canada Heavy Oil Technical Conference · 2015
Typearticle
Languageen
FieldEngineering
TopicReservoir Engineering and Simulation Methods
Canadian institutionsNexen (Canada)University of Alberta
Fundersnot available
KeywordsArtificial neural networkData miningComputer scienceOil fieldPrincipal component analysisCluster analysisField (mathematics)PetrophysicsProduction (economics)Reservoir simulationPetroleum engineeringEngineeringMachine learningArtificial intelligenceMathematics

Abstract

fetched live from OpenAlex

Abstract Quantitative evaluation of steam-assisted gravity drainage (SAGD) performance of heterogeneous reservoir is important for reservoir management and optimization of development strategies for oil sand operations. Although conventional commercial simulators are capable for detailed appraisal SAGD recovery performance, they are usually deterministic and computationally-demanding. Artificial intelligence approaches can be employed as a complementary tool for production forecast and pattern recognition of highly non-linear relationships between system variables. In this paper, a comprehensive dataset, consisting of petrophysical log measurements, production and injection profiles is assembled from various publicly available sources, encompassing ten different SAGD operating fields with approximately two hundred well pairs. Only fields with complete data records are selected. Artificial neural network (ANN) is employed to facilitate the production performance analysis. Predicting (input) variables that are descriptive of reservoir heterogeneities and operating constraints, including log-derived petrophysical parameters, dimensionless shale index, effective numbers of producers and injectors for a given well pair, total production time and cumulative steam injection, are formulated, while parameters pertaining to cumulative production and steam-to-oil ratio are considered as prediction (output) variables. Principal components analysis (PCA) is performed to reduce the dimensionality of the input variables, improve prediction quality and limit over-fitting. Clustering analysis is integrated to identify internal groupings among data. Finally, statistical analysis is conducted to study the influences of data uncertainty because of limited size of field dataset and imprecise log-interpretation criteria, together with model parameter uncertainty due to learning algorithm and initialization on the final ANN predictions. Workflows involving Monte Carlo and bootstrapping methods are applied successfully. A comprehensive uncertainty analysis using an actual SAGD dataset is a novel contribution. The modeling results are demonstrated to be both reliable and acceptable. This paper demonstrates the combination of artificial-intelligence approaches and data-mining analysis can be implemented in a practical manner to analyze large amount of field data, which is often prone to uncertainties and errors, with high reliability and feasibility. Considering that many important variables such as bottom-hole pressures, PVT properties, permeability measurements, multi-phase flow functions and thermal conductivities are typically unavailable in the public domain and, hence, are missing in the dataset, this work demonstrates how practical data-driven analysis approaches can be tailored to construct models capable of predicting SAGD recovery performance from only log-derived and operational variables. Another advantage of the proposed approach is that it can be updated when new information is obtained.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.003
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: Simulation or modeling
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.992
Threshold uncertainty score0.016

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.003
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.001
Science and technology studies0.0000.000
Scholarly communication0.0010.001
Open science0.0010.000
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.198
GPT teacher head0.354
Teacher spread0.156 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations9
Published2015
Admission routes1
Has abstractyes

Explore more

Same venueSPE Canada Heavy Oil Technical ConferenceSame topicReservoir Engineering and Simulation MethodsFrench-language works237,207