Prediction of Mineralogical Composition in Heterogeneous Unconventional Reservoirs: Comparisons Between Data-Driven and Chemistry-Based Models
Notice bibliographique
Résumé
Abstract Prediction of mineralogical compositions along multi-fractured horizontal wells (MFHWs) using indirect methods, for the purpose of characterizing lithological and rock brittleness heterogeneity, is appealing due to the challenges associated with direct mineralogical evaluation. This study aims to 1) develop predictive machine learning models for indirect estimation of mineralogical compositions from elemental compositions, 2) compare mineralogical compositions obtained from data-driven and chemistry-based approaches, and 3) provide practical recommendations for fine-tuning and training of data-driven models. Leveraging recent advances in deep learning, an attention-based gated recurrent unit (AttnGRU) with a "feature extractor-post processor" architecture was developed for predicting compositions of ten primary minerals based on elemental data. For comparison, classic regression-based and ensemble learning models including support vector regression (SVR), random forest (RF), and a feedforward neuron network (FFNN) were utilized. Data-driven models were trained and tested using XRD data measured on 217 samples from the Montney Formation, and the outcomes were compared to those derived from stoichiometric material balance equations (a previously-developed chemistry-based model) to evaluate the effectiveness and capabilities of different predictive approaches. The data-driven models consistently outperformed the chemistry-based method with significantly lower mean absolute error (MAE) and higher R2. The predictive performance order was FFNN ≥ AttnGRU > RF > SVR >> chemistry-based model, with MAE = 1.05, 1.09, 1.24, 1.35, and 2.46 wt.%, respectively. Importantly, FFNN, AttnGRU and RF offered more accurate predictions of chlorite and illite, which are known to negatively affect reservoir quality. This indicates the superior performance of the three models for reservoir characterization applications. Furthermore, AttnGRU exhibited greater robustness than the other two models, with less sensitivity to overfitting issues. Data-driven models displayed different levels of performance when decreasing training dataset size. It is recommended that, in order to achieve reasonable predictions for the studied reservoir with data-driven approaches, more than 50 training samples be used. It is further observed that data-driven models exhibited limited predictive capability (MAEs ranging from 3.02-3.45 wt.%) when applied to a synthetic "global dataset" comprised of samples from various formations. Through the comparison of multiple independent datasets (XRF-derived chemistry-based, XRF-derived data-driven, XRD) collected on identical samples, this work highlights the strengths, limitations, and capabilities of different machine learning techniques for along-well estimation of mineralogical composition to assist with reservoir characterization.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,001 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,001 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,001 |
| Science ouverte | 0,001 | 0,000 |
| Intégrité de la recherche | 0,001 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».