Leveraging advanced ensemble learning techniques for methane uptake prediction in metal organic frameworks
Bibliographic record
Abstract
Energy and environmental policy agencies have been looking for suitable adsorbent materials to promote the use of adsorbed natural gas (ANG). Various candidate adsorbent materials have been developed and tested for methane adsorption. Metal-Organic Frameworks (MOFs) have shown a promising performance in methane adsorption and are of particular interest due to their power in adsorption and separation of gases, chemical tunability, ease of synthesis, and high surface area. Accurate calculation of the theoretical adsorption potential of methane in MOFs and its validation through experiments brings about significant challenges. A growing number of researchers are adopting soft-computing approaches, particularly machine-learning (ML) algorithms, to tackle these challenges. Although ML algorithms have been applied in assessing methane uptake capacity of MOFs, the majority of these efforts have primarily focused on feature selection or the criteria for MOF screening. This communication, however, mainly focuses on the implementation of ensemble-based ML paradigms, including gradient boosting (GBoost), extreme gradient boosting (XGBoost), light gradient boosting machine (LightGBM), and gradient boosting with categorical features support (CatBoost) in accurate estimation of methane uptake capacity of experimentally synthesized MOFs based on some readily available features including temperature, pressure, and MOF’s pore volume and surface area, for the first time. To this end, a database containing almost 2600 datapoints was attained. The results indicated the high performance of the XGBoost algorithm in estimating the methane uptake capacity of MOFs with a correlation coefficient (R 2 ) of 0.9955. Moreover, further analyses revealed that the developed predictive model can reliably estimate the physical trend of CH 4 capacity variations with changing pressure. Also, further analysis indicated the large impact of pressure value on the predicted values. The employed outlier detection technique showed that almost 95% of the collected data points were valid.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".