Estimation of Speciation Data for Hydrocarbons using Data Science
Bibliographic record
Abstract
Strict regulations on air pollution motivates clean combustion research for fossil fuels. To numerically mimic real gasoline fuel reactivity, surrogates are proposed to facilitate advanced engine design and predict emissions by chemical kinetic modelling. However, chemical kinetic models could not accurately predict non-regular emissions, e.g. aldehydes, ketones and unsaturated hydrocarbons, which are important air pollutants. In this work, we propose to use machine-learning algorithms to achieve better predictions. Combustion chemistry of fuels constituting of 10 neat fuels, 6 primary reference fuels (PRF) and 6 FGX surrogates were tested in a jet stirred reactor. Experimental data were collected in the same setup to maintain data uniformity and consistency under following conditions: residence time at 1.0 second, fuel concentration at 0.25%, equivalence ratio at 1.0, and temperature range from 750 to 1100K. Measured species profiles of methane, ethylene, propylene, hydrogen, carbon monoxide and carbon dioxide are used for machine-learning model development. The model considers both chemical effects and physical conditions. Chemical effects are described as different functional groups, viz. primary, secondary, tertiary, and quaternary carbons in molecular structures, and physical conditions as temperature. Both the Machine-learning models used in this study showed a good prediction accuracy with a test set regression score of 97.75 for support vector regression and 91.07 for random forest regression. This finding shows the great potential of machine learning application on combustion chemistry. By expanding the experimental database, machine-learning models can be further applied to many other hydrocarbons in future work.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".