Insights into predicting equilibrium conditions of clathrate hydrates of methane + water-soluble hydrate former
Bibliographic record
Abstract
• Equilibrium conditions of clathrate hydrates of CH 4 and formers are determined. • Inputs involve molecular descriptors, mole fraction and pressure. • Data splitting uses a former-based approach, not traditional sample-based splits. • The developed ML approaches demonstrate acceptable prediction accuracy. • Shapley Additive Explanations (SHAP) approach is used to interpret results. This study aims to improve the prediction of equilibrium conditions in methane hydrate systems by incorporating diverse water-soluble hydrate formers and applying advanced machine learning techniques. Methane hydrates, which naturally form under high pressure and low temperature, can be more efficiently formed or dissociated by altering thermodynamic conditions using these hydrate formers. Accurate prediction of these conditions is crucial for optimizing gas storage and energy applications. In this research, molecular descriptors and operational parameters, such as mole fraction and pressure, are used as input variables to predict equilibrium temperature. Machine learning methods, including Decision Trees (DT), Random Forests (RF), Support Vector Machines (SVM), and Multi-Layer Perceptron (MLP), were employed with a novel data-splitting approach based on hydrate formers rather than traditional sample-based methods. Among these models, the RF achieved the highest performance, with a coefficient of determination (R 2 ) of 0.930, a root mean square error (RMSE) of 1.71, and an average absolute relative deviation (AARD) of 0.48%. Feature selection, preprocessing, and Shapley Additive Explanations (SHAP) provided valuable insights into the influence of specific variables on model predictions. Additionally, a supplementary examination, termed the reduced model, highlights the critical role of proper feature selection, with certain features regarded as less important yet essential for the functionality of distance-based models, particularly for models like SVM and MLP. This work advances methane hydrate research by offering a more accurate and interpretable framework for predicting hydrate equilibrium, addressing key gaps in previous studies, and extending its applicability to a broader range of systems. Moreover, the introduction of a former-based data-splitting method improves generalization across different hydrate formers, while the use of SHAP values for model interpretability offers deeper insights into the relationships between molecular descriptors and hydrate equilibrium conditions. This study paves the way for improved selection of hydrate formers in hydrate systems.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".