Methodology adjusting for least squares regression slope in the application of multiplicative scatter correction to near‐infrared spectra of forage feed samples
Bibliographic record
Abstract
Abstract Scatter corrections are commonly applied to refine near‐infrared (NIR) spectra. The aim of this study is to assess the impact of measurement errors when using ordinary least squares (OLS) for multiplicative scatter correction (MSC). Any measurement errors attached to the set‐mean spectrum may attenuate the OLS slope and that in turn will affect the estimate of the intercept and the adjustment of the spectra when using MSC methods to mitigate scattering. A corrected least squares slope may be used instead to prevent this problem, although the impact of this approach on the final outcome will depend on the relative size of the measurement errors in the individual spectra and the set‐mean spectrum. The errors‐in‐variables or type II regression model (also known as Deming regression) and its special cases, major axis (MA) and reduced major axis (RMA), are discussed and illustrated. The extent of OLS slope bias or attenuation is demonstrated as is the resulting MSC spectral distortion. Further modification to the MSC transformation method is also suggested. The influence of scattering correction (by MSC, standard normal variate (SNV) and detrending) and of using the maximum likelihood estimate of the slope for MSC on the prediction of chemical composition of Lucerne herbage from NIR spectra was assessed. The predictive performance was slightly improved by the use of scattering corrections with fairly minor differences among methods. Nonetheless, it seems well worth considering the use of type II regression models for assessing MSC application aiming at improving the goodness of prediction from NIR spectra.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.009 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".