Reconstruction of Raman Spectra of Biochemical Mixtures Using Group and Basis Restricted Non-Negative Matrix Factorization
Bibliographic record
Abstract
Raman spectroscopy is a useful tool for obtaining biochemical information from biological samples. However, interpretation of Raman spectroscopy data in order to draw meaningful conclusions related to the biochemical make up of cells and tissues is often difficult and could be misleading if care is not taken in the deconstruction of the spectral data. Our group has previously demonstrated the implementation of a group- and basis-restricted non-negative matrix factorization (GBR-NMF) framework as an alternative to more widely used dimensionality reduction techniques such as principal component analysis (PCA) for the deconstruction of Raman spectroscopy data as related to radiation response monitoring in both cellular and tissue data. While this method provides better biological interpretability of the Raman spectroscopy data, there are some important factors which must be considered in order to provide the most robust GBR-NMF model. We here evaluate and compare the accuracy of a GBR-NMF model in the reconstruction of three mixture solutions of known concentrations. The factors assessed include the effect of solid versus solutions bases spectra, the number of unconstrained components used in the model, the tolerance of different signal to noise thresholds, and how different groups of biochemicals compare to each other. The robustness of the model was assessed by how well the relative concentration of each individual biochemical in the solution mixture is reflected in the GBR-NMF scores obtained. We also evaluated how well the model can reconstruct original data, both with and without the inclusion of an unconstrained component. Overall, we found that solid bases spectra were generally comparable to solution bases spectra in the GBR-NMF model for all groups of biochemicals. The model was found to be relatively tolerant of high levels of noise in the mixture solutions using solid bases spectra. Additionally, the inclusion of an unconstrained component did not have a significant effect on the deconstruction, on the condition that all biochemicals in the mixture were included as bases chemicals in the model. We also report that some groups of biochemicals achieve a more accurate deconstruction using GBR-NMF than others, likely due to similarity in the individual bases spectra.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".