MétaCan
Menu
Back to cohort
Record W3033502499 · doi:10.2118/200621-ms

Data Driven Machine Learning Models for Shale Gas Adsorption Estimation

2020· article· en· W3033502499 on OpenAlexaboutno aff
Lei Wang, Mingliang Liu, Arlybek Altazhanov, Bekassyl Syzdykov, Jiang Yan, Xin Meng, Kai Jin

Bibliographic record

Venuenot available
Typearticle
Languageen
FieldEngineering
TopicHydrocarbon exploration and reservoir analysis
Canadian institutionsnot available
Fundersnot available
KeywordsLangmuirSupport vector machineAdsorptionOil shaleArtificial neural networkVolume (thermodynamics)Petroleum engineeringComputer scienceChemistryMachine learningGeologyThermodynamicsPhysicsPhysical chemistry

Abstract

fetched live from OpenAlex

Abstract Accurate calculation of adsorbed shale gas content is critical for gas reserve evaluation and development. However, gas adsorption and desorption experiments are expensive and time-consuming, while physics-based models and empirical correlations are unable to accurately capture the adsorption characteristics for different shales. Langmuir adsorption is one of the most commonly used model for calculating the adsorbed gas content in shale gas reservoirs. However, most existing correlations for the Langmuir pressure and Langmuir volume in the model are oversimplified based on limited experimental data points. Thus they are not representative of key geological parameters and are far from accurate for prediction in many cases. We developed a variety of machine learning models that are multivariable controlled to quantify shale gas adsorption. The data-driven method subdivides into two procedures: data compilation and machine learning regression. Over 700 data entries, composed of reservoir temperature (T, °C), total organic carbon (TOC, wt%), vitrinite reflectance (Ro,%), Langmuir pressure, and Langmuir volume are compiled from shale gas plays mainly in USA, Canada, and China. Data have been consistently curated, then machine learning approaches, including multiple linear regression (MLR), support vector machine (SVM), random forest (RF) and artificial neural network (ANN), have been built, trained and tested by partitioning the data into 75%:25%. For SVM, RF and NN models, 1000 simulations were run and averaged for performance comparison. MLR identifies non-negligible parameters and general trends for shale gas adsorption. Nonetheless, the correlation coefficients from MLR are far from satisfactory. For Langmuir pressure, RF models fit best to the data entries and the other models follow the order of SVM > ANN > MLR. Particularly, RF models show the highest performance stability with the averaged R-squared value of 0.84 and the maximum of 0.87, indicating a very strong relationship constructed for these 213 data entries. For 485 Langmuir volume data entries, RF models also perform best while the other three regression methods are comparable. It should be noted that altering machine learning model structure and parameters could significantly affect the regression results. Robust and universal machine learning models for estimating adsorbed shale gas content with high confidence level are established, which not only provide more accurate estimation and broader parameter adaptation than physics-based and empirical models, but also circumvent the high-cost and time-consuming deficiency of experimental measurements. These machine learning models can be used to estimate adsorbed gas content for shale plays with limited experimental measurements. Moreover, they can be incorporated into reservoir simulators to improve the simulation performance.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.005
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: Simulation or modeling
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.015
Threshold uncertainty score0.030

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0020.005
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.001
Science and technology studies0.0000.000
Scholarly communication0.0010.001
Open science0.0010.001
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0020.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.073
GPT teacher head0.261
Teacher spread0.188 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations18
Published2020
Admission routes1
Has abstractyes

Explore more

Same topicHydrocarbon exploration and reservoir analysisFrench-language works237,207