MétaCan
Menu
Back to cohort
Record W4411946786 · doi:10.1021/acsomega.5c04463

Machine Learning-Based Predictive Modeling of Infrared Spectroscopic Data from Thermal Conversion of Athabasca Bitumen

2025· article· en· W4411946786 on OpenAlexfundno aff
Noora Al Mansoori, Munawar A. Shaik, Kaushik Sivaramakrishnan

Bibliographic record

VenueACS Omega · 2025
Typearticle
Languageen
FieldChemistry
TopicPetroleum Processing and Analysis
Canadian institutionsnot available
FundersUnited Arab Emirates UniversityUniversity of AlbertaAbu Dhabi University
KeywordsAsphaltThermal infraredInfraredThermalEnvironmental scienceMaterials scienceComputer scienceArtificial intelligenceComposite materialOpticsThermodynamicsPhysics

Abstract

fetched live from OpenAlex

Abstract This study explores the use of machine learning (ML) techniques to predict Fourier-transform infrared (FTIR) intensities of products from the thermal cracking of Athabasca bitumen, aiming to develop a reliable soft-sensor. The ultimate goal is to obtain the FTIR spectra of the thermally cracked products online to reduce process time from slow physical measurements. Various ML models, including Linear Regression (LinR), partial least squares regression (PLSR), support vector regression (SVR), K-nearest neighbors (k-NN), random forest (RF), and gradient boosting regression (GBR), were implemented to enhance the predictive accuracy and efficiency of FTIR spectroscopy, aiming to reduce the need for traditional physical measurements which are often slow compared to the rapid predictions offered by ML techniques. To assess the model’s generalization capabilities, with respect to model predictions, the models were trained and tested across four different scenarios with varying temperature data obtained from visbreaking experiments performed on Athabasca Bitumen at temperatures ranging from 25 to 420 °C with reaction times ranging from 15 min to 27 h. Scenario 1 included all 61,740 data points utilizing an 80/20 train-test split with 10-fold cross-validation (CV). Scenario 2 involved training on temperatures of 25, 350, and 400 °C and testing on 300, 380, and 420 °C. Scenario 3 involved training on temperatures of 350, 380, and 400 °C and testing on 25, 300, and 420 °C. Finally, Scenario 4 involved training on temperatures of 25, 300, 350, and 380 °C and testing on 400 and 420 °C. Bayesian optimization was employed for hyperparameter tuning to identify the optimal configurations for each model. The results indicate that ensemble methods, particularly GBR, consistently achieved the highest predictive accuracy (R2) and lowest root mean squared error (RMSE) across all scenarios. In Scenario 1, GBR achieved a prediction accuracy of 99.66%. Scenario 2 highlighted the models’ ability to generalize across varying temperatures, with both RF and GBR achieving similar performance with high prediction accuracies of around 94%. Scenario 3, characterized by significant temperature variability, demonstrated the robustness of GBR, which outperformed RF and k-NN with a predictive accuracy of 92.15%. Scenario 4, focusing on high-temperature predictions from low-temperature training data, showed that GBR still performed robustly with a predictive accuracy of 80.40%. The study concludes that GBR models, particularly those with well-tuned hyperparameters, are highly effective in predicting FTIR intensities, outperforming other techniques like RF, k-NN, LinR, and PLSR. The integration of advanced ML techniques and Bayesian optimization significantly enhances the capability to predict FTIR spectra, providing a reliable soft-sensor as an alternative to traditional physical experimentation methods. This approach not only saves time and resources but also ensures consistent and high-quality predictive performance in chemical analysis and monitoring.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.001
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: Simulation or modeling
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.016
Threshold uncertainty score0.032

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.001
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0000.001
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0010.000
Open science0.0010.000
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.015
GPT teacher head0.251
Teacher spread0.237 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueACS OmegaSame topicPetroleum Processing and AnalysisFrench-language works237,207