Representation of Snow in the Canadian Seasonal to Interannual Prediction System. Part II: Potential Predictability and Hindcast Skill
Bibliographic record
Abstract
Abstract This paper examines potential predictability (PP) and actual skill for snow water equivalent (SWE) in the Canadian Seasonal to Interannual Prediction System (CanSIPS). A significant PP is found for SWE, with potentially predictable variance over 50% of the total variance at up to a 5-month lead in mid- to high latitudes in forecasts initialized after snow onset. Much, though not all, of this PP stems from a tendency for SWE anomalies to persist through the snow season. Although the spring melt acts as a PP barrier regardless of initialization date, in some regions significant PP reemerges in the following snow season. This is due primarily to ENSO teleconnections that are modeled realistically by CanSIPS, particularly in northwestern North America. Actual skill of CanSIPS in forecasting SWE is assessed using several verification datasets. Highest skills are obtained using a blend of five such datasets, consistent with the hypothesis that skill scores are degraded by errors in the verification data as well as by forecast errors, and that observational errors can be reduced by blending multiple datasets, much as forecast errors can be reduced by averaging across different models. Actual skill for SWE is comparable to, though generally lower than, that implied by PP. This is due in part to the similar autocorrelation properties of the forecast and observed SWE anomalies, which provide skill through anomaly persistence, combined with reasonably accurate initialization of SWE by CanSIPS. Long-lead skill across snow seasons is found to be linked to ENSO, particularly in western North America, much as for PP.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".