A comparative analysis of machine learning predictive models for the oxidative coupling of methane reaction
Bibliographic record
Abstract
The amalgamation of catalytic and electronic characteristics, along with empirical data, furnishes enhanced insights into catalyst analysis, thereby advancing innovation and the design of heterogeneous catalytic reactions. In this research, we juxtaposed the catalysts' electronic properties, including the Fermi energy, bandgap energy, and magnetic moment of catalyst components with available high-throughput OCM experimental data, to prognosticate catalytic efficacy and the resultant reaction outcomes, encompassing methane conversion and yields of ethylene, ethane, and carbon dioxide yields. Comparative evaluation of diverse machine learning models indicates that the extreme gradient boost regression model stands out for its superior predictive accuracy in evaluating catalytic performance with an average R 2 of 0.91. The order of performance of the modeling techniques was XGBR > RFR > DNN > SVR. The MSE and MAE of the XGBR models appeared to be lower than those of the other modeling techniques, with numbers ranging from 0.26 to 0.08 for MSE and 1.65–0.17 for MAE. The MSE and MAE of the models trained with a particular dataset generally aligned with those of an external dataset not seen by the model at the time of training or testing, as well as the MSE from bootstrap sampling, confirming the high generalizability of the models. The accuracy of the models was in the order of C 2 H 6 y > CH 4 _conv > CO 2 y > C 2 y > C 2 H 4 y. Comparing the quantification of uncertainty of the RCCEP feature-based predictive models of the different ML techniques based on the prediction band suggests that the ML techniques rank in the order of XGBR > RFR > DNN > SVR. In analyzing the impact of the model features, the combined ethylene and ethane yield increases with an increase in dataset features, including the number of moles of the alkali/alkali-earth metal in the catalyst, the atomic number of the catalyst promoter and the Fermi energy of the metal, and just relatively in the case of temperature, suggesting a highly non-linear relationship between the combined ethylene and ethane yield and temperature. The extent of the impact of these features on the predictive model for the combined ethylene and ethane yield was 5.91 %, 13.28 % and 33.76 % for the atomic number of the promoter, the number of moles of the alkali/alkali-earth metal in the catalyst and the reaction temperature, respectively. Other features, including the bandgap of the active metal oxide and the support, as well as the Fermi energy of the catalyst support, were also seen to have a relatively modest impact on the predictive models for the combined ethylene and ethane yield and methane conversion. • A comparative evaluation of different ML models in the study of the OCM reaction. • The order of model performance is XGBR > RFR > DNN > SVR. • XGBR models have an average R 2 of 0.91; MSE and MAE ranging from 0.26 to 0.08 and 1.65–0.17 respectively. • The catalyst's promoter fermi energy and atomic number impact ethylene and ethane. • The catalyst's oxide and support bandgap moderately affect methane-to-ethylene conversion.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".