Decision-tree-based non-resonant background removal for enhancing chemical-selective contrast in hyperspectral CARS microscopy
Bibliographic record
Abstract
Coherent anti-Stokes Raman scattering (CARS) is a nonlinear optical process used for spectroscopy and chemical imaging. CARS signals can be orders of magnitude stronger than those of its incoherent counterpart, spontaneous Raman scattering, thus enabling substantially faster acquisition speeds. This attribute has positioned CARS as a desirable alternative to spontaneous Raman scattering as a contrast mechanism for label-free hyperspectral chemical imaging due to its traditionally long acquisition times. The presence of a non-resonant background (NRB) that distorts the shapes and intensities of resonant peaks and introduces spurious signal to non-resonant spectral regions, however, complicates spectral analysis and degrades chemical-selective image contrast. The NRB has and continues to hinder the widespread adoption of CARS despite its clear advantages. NRB removal techniques that retrieve Raman-like signals from CARS spectra have long been a central focus of CARS research, with "deep" machine learning approaches being most recently explored. Here, we present an original "shallow" machine learning approach to NRB removal based on gradient-boosted decision trees obtained using the open-source gradient-boosting framework XGBoost. We find that the tree-based model accurately retrieves Raman-like spectra, both in simulated and experimental CARS spectra. When applied to experimental hyperspectral CARS microscopy, the tree-based model significantly improves chemical-selective contrast, thereby facilitating the spatiospectral analysis of samples within the region of interest. This work establishes tree-based gradient boosting as an effective and viable tool for NRB removal for enhancing chemical-selective contrast in hyperspectral CARS microscopy.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".