Prediction of Thermogravimetric Data for Asphaltenes Extracted from Deasphalted Oil Using Machine Learning Techniques
Bibliographic record
Abstract
Thermogravimetric analysis (TGA) has been extensively used in the bitumen literature to investigate its thermal stability and various stages of thermal decomposition. The primary aim of these studies has been to calculate the kinetic parameters, such as activation energy and the pre-exponential factor of each thermal event. However, in our current paper, we explore the application of three machine learning (ML) techniques, namely, support vector regression (SVR), random forest (RF), and gradient booster regression (GBR), to predict the TGA data for the asphaltenes extracted from the feed and products of visbreaking of three types of materials: (i) deasphalted oil (DAO), (ii) DAO doped with 5.55 wt % indene, and (iii) DAO doped with 11.11 wt % indene. The addition of indene was shown to significantly affect the free-radical chemistry of DAO in a previous work, and the key contribution of our work in this paper was to minimize the requirement of the TGA instrument to obtain the mass loss curves by employing ML techniques on available experimental data. This will reduce the human errors involved in sample preparation and data collection as well as decrease the process time in obtaining the TGA data as compared to experimentation. We observed that the regression techniques based on decision trees, i.e., RF and GBR, showed the best performance and highest prediction accuracy of >0.99 for predicting the TGA data of the asphaltenes extracted from the feed and products obtained by reacting the feedstocks at visbreaking reaction times of 30, 45, and 60 min. A number of inputs were considered for the ML models, such as the temperature of the TGA chamber and sample, heat supplied to the sample, visbreaking time, and time spent inside the TGA chamber. The novelty of our work lies in the fact that no previous study has reproduced the TGA data for asphaltenes extracted from DAO and indene-added DAO and their visbroken products through ML approaches, and we believe that the results of this work will help in fastening the process times in the heavy oil industry by eliminating the need for offline measuring instruments.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".