NWP verification against own analysis by using a Data Assimilation confidence mask
Bibliographic record
Abstract
Numerical Model Prediction (NWP) verification against station measurements from a surface network is affected by sub-tile representativeness issues. Moreover, the station network is often not representative of the whole verification domain (e.g. usually coastal stations are predominant) and large unpopulated regions (such as oceans, Polar regions, deserts) are under-sampled. Verification against gridded analyses mitigate these issues, since they partially address the sub-tile representativeness, and sample homogeneously the verification domain. Moreover, gridded analyses merge station network measurements to radar and satellite retrieval estimates, in a physical coherent fashion, over the same NWP grid. Verification against own analysis, despite quite convenient, is however hampered by its dependence on the NWP background model, which renders the verification “incestuous”, further than being affected by the uncertainties introduced by retrieval algorithms and Data Assimilation (DA) procedures. In this study we investigate the use of a gridded NWP own analysis for verification, by applying a mask to reduce the background model contribution. The mask weights the verification scores to account for the amounts of observations assimilated and their associated uncertainty, as estimated from DA. We illustrate the approach by using the Canadian Precipitation Analysis (CaPA), which assimilates station measurements, radar and satellite-based (IMERG) observations. The CaPA confidence (weighting) mask is dynamic and changes depending on the daily available (assimilated) observations, and on their corresponding DA error statistics; it is defined as mask = 1 - var(A-O)/var(B-O) where A=analysis, B=Background, O=observations. Where the analysis is identical to the background model, the weighting mask is zero. We evaluate the Canadian Regional Deterministic Prediction System (RDPS), which is the NWP system used as background model for CaPA. As expected, the verification results obtained by using the weighting mask lay between the verification results obtained verifying against the analysis over the full domain, and the results obtained verifying against station measurements. The effects of sub-tile representativeness are quantified by comparing verification results against station measurements to verification results against CaPA for the grid-points co-located with the stations. Finally, the comparison of the verification results against CaPA over the full domain versus the verification results against CaPA for the grid-points co-located with stations, estimates to which extent the station network is representative of the full domain. The approach aims to propose a simple -yet effective- better practice for verification against own analysis.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.030 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".