Modeling of H2S solubility in ionic liquids: comparison of white-box machine learning, deep learning and ensemble learning approaches
Bibliographic record
Abstract
Abstract In the context of gas processing and carbon sequestration, an adequate understanding of the solubility of acid gases in ionic liquids (ILs) under various thermodynamic circumstances is crucial. A poisonous, combustible, and acidic gas that can cause environmental damage is hydrogen sulfide (H2S). ILs are good choices for appropriate solvents in gas separation procedures. In this work, a variety of machine learning techniques, such as white-box machine learning, deep learning, and ensemble learning, were established to determine the solubility of H2S in ILs. The white-box models are group method of data handling (GMDH) and genetic programming (GP), the deep learning approach is deep belief network (DBN) and extreme gradient boosting (XGBoost) was selected as an ensemble approach. The models were established utilizing an extensive database with 1516 data points on the H2S solubility in 37 ILs throughout an extensive pressure and temperature range. Seven input variables, including temperature (T), pressure (P), two critical variables such as temperature (Tc) and pressure (Pc), acentric factor (ω), boiling temperature (Tb), and molecular weight (Mw), were used in these models; the output was the solubility of H2S. The findings show that the XGBoost model, with statistical parameters such as an average absolute percent relative error (AAPRE) of 1.14%, root mean square error (RMSE) of 0.002, standard deviation (SD) of 0.01, and a determination coefficient (R2) of 0.99, provides more precise calculations for H2S solubility in ILs. The sensitivity assessment demonstrated that temperature and pressure had the highest negative and highest positive affect on the H2S solubility in ILs, respectively. The Taylor diagram, cumulative frequency plot, cross-plot, and error bar all demonstrated the high effectiveness, accuracy, and reality of the XGBoost approach for predicting the H2S solubility in various ILs. The leverage analysis shows that the majority of the data points are experimentally reliable and just a small number of data points are found beyond the application domain of the XGBoost paradigm. Beyond these statistical results, some chemical structure effects were evaluated. First, it was shown that the lengthening of the cation alkyl chain enhances the H2S solubility in ILs. As another chemical structure effect, it was shown that higher fluorine content in anion leads to higher solubility in ILs. These phenomena were confirmed by experimental data and the model results. Connecting solubility data to the chemical structure of ILs, the results of this study can further assist to find appropriate ILs for specialized processes (based on the process conditions) as solvents for H2S.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".