Spatial and Temporal Variation of Subseasonal-to-Seasonal (S2S) Precipitation Reforecast Skill across the CONUS
Bibliographic record
Abstract
Abstract Precipitation forecasts, particularly at subseasonal-to-seasonal (S2S) time scale, are essential for informed and proactive water resource management. Although S2S precipitation forecasts have been evaluated, no systematic decomposition of the skill, Nash–Sutcliffe efficiency (NSE) coefficient, has been analyzed toward understanding the forecast accuracy. We decompose the NSE of S2S precipitation forecast into its three components—correlation, conditional bias, and unconditional bias—by four seasons, three lead times (1–12, 1–22, and 1–32 days), and three models, European Centre of Medium-Range Weather Forecasts (ECMWF), National Centers for Environmental Prediction’s (NCEP) Climate Forecast System (CFS) model, and Environment and Climate Change Canada (ECCC), over the conterminous United States (CONUS). Application of a dry threshold, removal of grid cells with seasonal climatological precipitation means below 0.01 in. per day, is important as the NSE and correlations are lower across all seasons after masking areas with low precipitation values. Further, a west-to-east gradient in S2S forecast skill exists, and forecast skill was better during the winter months and for areas closer to the coast. Overall, ECMWF’s model performance was stronger than both ECCC and NCEP CFS’s performance, mainly for the forecasts issued during the fall and winter months. However, ECCC and NCEP CFS performed better for the forecast issued during the spring months and for areas further from the coast. Postprocessing using simple model output statistics could reduce both unconditional and conditional biases to zero, thereby offering better skill for regimes with high correlation. Our decomposition results show that efforts should focus on improving model parameterization and initialization schemes for climate regimes with low correlation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".