Benefits of fuzzy methods for the evaluation of high-resolution snow models
Bibliographic record
Abstract
<!--!introduction!--> In mountainous areas, accurately resolving snowpack temporal and spatial variability is still challenging for the snow modelling community. The last decades have seen significant advances in snow processes simulations, including wind-induced snow transport. However, the verification/evaluation methods used to confront model results with observations have not improved at the same pace. A direct evaluation of every simulated processes is currently impracticable. Therefore, the evaluation of regional to continental snowpack simulations is mostly done using indirect observations of the physical process of interest. Snow depth measurements from satellite or airborne laser scanner or composite variables of snow absence and presence derived from satellites are usual verification data sources. Yet, using this kind of data in a pixel-to-pixel verification usually shows poor model performance, making it difficult to assess the added value of newly implemented processes, for instance, wind-induced snow transport. Fuzzy verification methods have been developed to account for these difficulties, allowing more spatial tolerance. We will illustrate this challenge with the evaluations of the SnowPappus model, a new simple blowing snow transport model coupled with the Crocus state-of-the-art physical snow model. It is designed to improve the snow spatial variability of the French snowpack simulation system by predicting blowing snow occurrence, transport fluxes and sublimation at 250m resolution. Although pixel-to-pixel comparisons of the Snowpappus model with satellite-retrieved snow depth show low added values of the model, fuzzy verification techniques demonstrate an increased spatial variability with the transport model, in line with observations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.009 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".