Uncertainty and outliers in high‐resolution gridded precipitation products over eastern North America
Bibliographic record
Abstract
Abstract Several observational precipitation products that provide high temporal (≤3 h) and spatial (≤0.25°) resolution gridded estimates are available, although no single product can be assumed worldwide to be closest to the (unknown) “reality.” Here, we propose and apply a methodology to quantify the uncertainty of a set of precipitation products and to identify, at individual grid points, the products that are likely wrong (i.e., outliers). The methodology is applied over eastern North America for the 2015–2019 period for eight high‐resolution gridded precipitation products: CMORPH, ERA5, GSMaP, IMERG, MSWEP, PERSIANN, STAGE IV and TMPA. Four difference metrics are used to quantify discrepancies in different aspects of the precipitation time series, such as the total accumulation, two characteristics of the intensity‐frequency distribution, and the timing of precipitating events. Large regional and seasonal variations in the observational uncertainty are found across the ensemble. The observational uncertainty is higher in Canada than in the United States, reflecting large differences in the density of precipitation gauge measurements. In northern midlatitudes, the uncertainty is highest in winter, demonstrating the difficulties of satellite retrieval algorithms in identifying precipitation in snow‐covered areas. In southern midlatitudes, the uncertainty is highest in summer, probably due to the more discontinuous nature of precipitation. While the best product cannot be identified due to the lack of an absolute reference, our study is able to identify products that are likely wrong and that should be excluded depending on the specific application.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".