Determination of Trace Organic Contaminant Concentration via Machine Classification of Surface-Enhanced Raman Spectra
Bibliographic record
Abstract
Surface-enhanced Raman spectroscopy (SERS) has been well explored as a highly effective characterization technique that is capable of chemical pollutant detection and identification at very low concentrations. Machine learning has been previously used to identify compounds based on SERS spectral data. However, utilization of SERS to quantify concentrations, with or without machine learning, has been difficult due to the spectral intensity being sensitive to confounding factors such as the substrate parameters, orientation of the analyte, and sample preparation technique. Here, we demonstrate an approach for predicting the concentration of sample pollutants from SERS spectra using machine learning. Frequency domain transform methods, including the Fourier and Walsh-Hadamard transforms, are applied to spectral data sets of three analytes (rhodamine 6G, chlorpyrifos, and triclosan), which are then used to train machine learning algorithms. Using standard machine learning models, the concentration of the sample pollutants is predicted with >80% cross-validation accuracy from raw SERS data. A cross-validation accuracy of 85% was achieved using deep learning for a moderately sized data set (∼100 spectra), and 70-80% was achieved for small data sets (∼50 spectra). Performance can be maintained within this range even when combining various sample preparation techniques and environmental media interference. Additionally, as a spectral pretreatment, the Fourier and Hadamard transforms are shown to consistently improve prediction accuracy across multiple data sets. Finally, standard models were shown to accurately identify characteristic peaks of compounds via analysis of their importance scores, further verifying their predictive value.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".