MétaCan
Menu
Back to cohort
Record W4280552350 · doi:10.1255/tosf.171

Estimating the heterogeneity invariant using size-density classes – the case of contaminated soil and complex materials

2022· article· en· W4280552350 on OpenAlexaff
Jean‐Sébastien Dubé, Kim H. Esbensen

Bibliographic record

VenueTOS forum · 2022
Typearticle
Languageen
FieldComputer Science
TopicGeochemistry and Geologic Mapping
Canadian institutionsÉcole de Technologie Supérieure
Fundersnot available
KeywordsSampling (signal processing)Variance (accounting)Sampling theoryStatisticsVariance reductionMathematicsField (mathematics)Sample size determinationEconometricsComputer scienceStatistical physicsPhysicsMonte Carlo methodPure mathematics

Abstract

fetched live from OpenAlex

For the important material class of aggregate mixtures comprised by both analyte-enriched and analyte-coated particles, Gy’s classical s2(FSE) formula has often been reported to yield estimates of the fundamental sampling variance greater than the empirically estimated sampling variance. Which, however, is physically impossible according to the tenets of the Theory of Sampling (TOS), both physically and logically, since the fundamental sampling variance is, by definition, the minimum sampling variance remaining after all other sources of sampling errors have been eliminated. This situation has for many decades hindered rational use of the Theory of Sampling for this kind of complex systems. We here focus on contaminated soil as a typical illustrative example of great interest, as well as more generally in the field of environmental site assessment. This uncomfortable situation is exacerbated by the fact that sampling in these fields is still, after 70+ years of TOS, largely conducted by grab sampling, which assuredly lead to significant uncertainty and bias. However, there is a solution to this at first sight“intractable” problem to be found, specifically within TOS. In some of his earlier publications, Gy developed a variant the s2(FSE) formula based on consideration of both size- and density classes, but quickly dismissed this approach as being inapplicable to « the metal, mining, and processing industries […] due to the unusual density contrast between the components » typical of matrices sampled in these fields. This size-density class variant was consequently then left out of sampling awareness and literature for a long time. We revisit herein development of the heterogeneity invariant on this basis and show this to represent a general option which can be adapted to distinct and specific matrices andanalytes beyond the original restricted realm. As an example, a size-density class variant is applied to data from studies on sampling complex contaminated soils for which the use of Gy’s classical formula yielded such “impossible” estimates of s2(FSE) larger than empirical sampling variances by several orders of magnitude. This size-density class s2(FSE) variant now provides estimates for all cases and examples, which are systematically smaller than the empirical sampling variances, and thus in full accordance with the Theory of Sampling. This generalised approach is also applied with similar success to controlled materials, which were made to represent analyte-enriched and/or analyte-coated matrices, as used in recent studies on sampling bias and representativeness. The results in our studies all show that it is not Gy’s classical formula which was at fault when applied outside the traditional domains, e.g., to contaminated soils, it is that the critical assumptions behind the formula were broken, unwittingly, or worse, with blunt carelessness. In analyte-coated materials, or for mixed matrices, the original full size-density class-based formula now provides the proper starting point for developing TOS-compliant matrix-specific approaches on a much broader scale. With this new scope, analysts no longer must forego the revolutionary advantage of Gy’s classical formula, i.e., the capacity to estimate the fundamental sampling variancea priori. Now, only at the cost of a pilot sampling stage, the augmented size-density class formula provides the analyst with the capacity to adapt sampling protocols also to the challenging task of taking on practically all natural systems sesu lato, however complex.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.768
Threshold uncertainty score0.699

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0010.000
Scholarly communication0.0000.000
Open science0.0000.001
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.035
GPT teacher head0.258
Teacher spread0.223 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2022
Admission routes1
Has abstractyes

Explore more

Same venueTOS forumSame topicGeochemistry and Geologic MappingFrench-language works237,207