An open-access computational fingerprinting workflow for source classifications of neat gasoline using GC × GC-TOFMS and Machine Learning
Bibliographic record
Abstract
Advances in sensitivity and selectivity of multidimensional chromatography have enhanced our ability to better characterize and identify sources of neat gasoline used in arson cases. However, the large and complex chemical datasets generated present a significant challenge for data management and interpretation, requiring robust computational analysis techniques. In this study, we present a novel, open-access computational fingerprinting workflow to develop regional database of gasoline profiles and source tracking of gasoline samples from local gas stations for arson investigations. The computational workflow included data reduction, normalization, clustering analyses, feature selection and supervised machine learning (ML) to explore the differentiation between gasoline sources. Chromatographic features (n = 25,415) from multidimensional gas chromatography-time of flight mass spectrometry (GC × GC-TOFMS) analysis of 69 neat gasoline samples, collected from 10 gas stations in Alberta (Canada), were used in supervised ML for the classification of neat gasoline samples. Fifty chemical features selected using recursive feature addition (RFA), with associated chemistries of n-alkanes, alkenes, cycloalkanes, and aromatics, were found to differentiate local gas stations. Despite overlapping between gas stations in clustering analyses, an average improvement of 18 % in ML accuracy was achieved by using decision tree-based ML classifiers coupled with RFA as compared to using all features. Our open-source computational workflow ensures transparency and reproducibility in creating a method and regional database for the distinction of gasoline sources commonly used in wildfire arson. The workflow enables forensic analysts to integrate additional chemical features into existing target chemical libraries within the ASTM E1618-19 protocol, enhancing ignitable liquid identification without requiring extensive re-training of computational models or programming expertise.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".