Improved Analog Ensemble Formulation for 3-Hourly Precipitation Forecasts
Bibliographic record
Abstract
Abstract Analog ensembles (AnEns) traditionally use a single numerical weather prediction (NWP) model to make a forecast, then search an archive to find a number of past similar forecasts (analogs) from that same model, and finally retrieve the actual observations corresponding to those past forecasts to serve as members of an ensemble forecast. This study investigates new statistical methods to combine analogs into ensemble forecasts and validates them for 3-hourly precipitation over the complex terrain of British Columbia, Canada. Applying the past analog error to the target forecast (instead of using the observations directly) reduces the AnEn dry bias and makes prediction of heavy-precipitation events probabilistically more reliable—typically the most impactful forecasts for society. Two variants of this new technique enable AnEn members to obtain values outside the distribution of the finite archived observational dataset—that is, they are theoretically capable of forecasting record events, whereas traditional analog methods cannot. While both variants similarly improve heavier precipitation events, one variant predicts measurable precipitation more often, which enhances accuracy during winter. A multimodel AnEn further improves predictive skill, albeit at higher computational cost. AnEn performance shows larger sensitivity to the grid spacing of the NWP than to the physics configuration. The final AnEn prediction system improves the skill and reliability of point forecasts across all precipitation intensities. Significance Statement The analog ensemble (AnEn) technique is a data-driven method that can improve local weather forecasts. It improves raw model forecasts using past similar model predictions and observations, reducing future forecast errors and providing probabilities for a range of possible outcomes. One limitation of AnEns is that they commonly tend to make rare-event (e.g., heavy precipitation) forecasts appear less extreme. Usually, heavier precipitation events have a higher impact on society and the economy. This study introduces two new AnEn techniques that make operational forecasts of both probabilities and most likely amounts more accurate for heavy precipitation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".