MétaCan
Menu
Back to cohort
Record W4307042788 · doi:10.1177/09670335221124612

Before reliable near infrared spectroscopic analysis - the critical sampling proviso. Part 1: Generalised theory of sampling

2022· article· en· W4307042788 on OpenAlexaff
Kim H. Esbensen, Nawaf Abu‐Khalaf

Bibliographic record

VenueJournal of Near Infrared Spectroscopy · 2022
Typearticle
Languageen
FieldChemistry
TopicSpectroscopy and Chemometric Analyses
Canadian institutionsUniversité du Québec à Chicoutimi
Fundersnot available
KeywordsSampling (signal processing)InfraredStatisticsRemote sensingMaterials scienceMathematicsOpticsPhysicsGeologyDetector

Abstract

fetched live from OpenAlex

Non-representative sampling of materials, lots and processes intended for near infrared (NIR) analysis is often contributing hidden additions to the full Measurement Uncertainty (MU total = TSE + TAE NIR ). The Total Sampling Error (TSE) can dominate over the Total Analytical Error (TAE NIR ) by factors ranging from 5 to 10 to even 25 times, depending on material heterogeneity and the specific sampling procedures employed to produce the minuscule aliquot, which is the only material analysed. This review (Parts 1 and 2), extensively referenced with easily available complementing literature, presents a brief of all sampling uncertainty elements in the “lot-to-aliquot” pathway, which must be identified and correctly managed (eliminated or maximally reduced) in order to achieve, and to be able to document, fully minimised MU total . The more irregular and pervasive the heterogeneity, the higher the number of increments needed to reach ‘fit-for-purpose representativity’. A particular focus is necessary regarding the sampling bias, which is fundamentally different from the well-known analytical bias. Whereas the latter can easily be subjected to bias correction, the sampling bias is non-correctable by any posteori means, notably not by chemometrics, nor statistics. Instead, all sampling operations must be designed to exclude the so-called Incorrect Sampling Errors (ISE), which are the hidden bias-generating agents. The key element in this endeavour is representative sampling and sub-sampling before analysis, as laid out by the Theory of Sampling (TOS), which is presented here in a novel compact fashion along with a complement of selected examples and demonstrations. TOS includes a safeguard facility, termed the Replication Experiment (RE), which enables estimation of the total sampling- plus-analysis uncertainty level (MU total ) associated with NIR analysis (the RE is, for practical and logistical reasons, found in Part 2). Neglecting the TSE effects from the before-analysis domain is lack of due diligence. TOS to the fore!

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.036
metaresearch head score (Gemma)0.046
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: Theoretical or conceptual
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.036
Threshold uncertainty score0.192

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0360.046
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0020.001
Bibliometrics0.0020.002
Science and technology studies0.0010.019
Scholarly communication0.0060.010
Open science0.0030.004
Research integrity0.0050.009
Insufficient payload (model declined to judge)0.0030.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.032
GPT teacher head0.311
Teacher spread0.278 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designTheoretical or conceptual
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations5
Published2022
Admission routes1
Has abstractyes

Explore more

Same venueJournal of Near Infrared SpectroscopySame topicSpectroscopy and Chemometric AnalysesFrench-language works237,207