Bibliographic record
Abstract
Abstract The changeable, variable and fragile nature of snow creates unique sampling challenges. In this paper, we present Star: an efficient, field-usable method for use in point-sampling spatial studies. We validate the accuracy of the Star method using a comparative Monte Carlo simulation of 1024 detailed samples of elevation data. As spatial snow studies generally attempt to find spatial continuity in layers and other properties, we use variogram ranges to compare the ability of four sampling methods to accurately reveal such spatial correlation. The three methods compared to Star represent gridded, gridded-random and pure-random methods; Star can be described as a linear-random method. The simulation shows Star’s accuracy to be comparable to both gridded and gridded-random methods. Following this comparative process we introduce a new measure of appropriateness for sampling methods: the correct convergence on a variogram model, which we call correct spatial correlation detection. This directly measures how many sampled areas become correctly classified with either spatially correlated or non-correlated variance for a given variogram model fit. In this measure, Star performs equivalently to the other methods, and in correct convergence it performs as well as pure-random sampling.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.008 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".