Evaluating Low-Intensity Unknown Signals in Quantitative Proton NMR Mixture Analysis
Bibliographic record
Abstract
Analytical analyses of highly complex mixtures, such as biofluids or liquid food products, often give rise to signals for unknown compounds, particularly for compounds at low concentration. Here we compare two conventional chemometric approaches for NMR spectral analysis ("spectral binning" and "high-resolution analysis") with a novel library-based method ("targeted profiling of unknowns", TPU). The three methods were applied to a proton NMR spectral data set of ultrafiltered mouse serum typical of those examined in metabolomics/metabonomics studies. The advantages of high-resolution analysis of typical NMR peaks have been well described previously, and as a result we examined low intensity unknowns peaks (LIUPs). A total of 25 LIUPs were assessed based on their significance to multivariate statistical analysis of the data set using the TPU method. The linearity of NMR signals at low incremental concentration changes (< 10 microM) was determined by titration of endogenously occurring metabolites into filtered mouse serum. Carbon-13 decoupling of the NMR spectra was used to ensure isotope-satellite peaks were eliminated. Four peaks were noted as significant to separation between arthritic and diseased animals. The conventional spectral methods were hampered by baseline noise or overlap with high concentration metabolites and were not able to identify these LIUPs reliably. In general, conventional methods, particularly high-resolution analysis, are recommended for peaks with moderate signal-to-noise. The TPU method is recommended for peaks with low signal-to-noise or when compression of spectral data with high fidelity is desirable, such as integration of NMR data into cross-platform studies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".