Uncertainty of precipitation reference dataset for climate change impact studies
Bibliographic record
Abstract
Abstract. Climate change impact studies typically require a reference climatological dataset providing a baseline period to assess future changes. The reference dataset is also used to perform bias correction of climate model outputs. Various reliable precipitation datasets are now available over regions with a high-density network of weather stations such as over most parts of Europe and in the United States. In many of the world’s regions, the low-density of observation stations (or lack thereof) renders gauge-based precipitation datasets highly uncertain. Satellite, reanalysis and merged products can be used to overcome this limitation. However, each dataset brings additional uncertainty to the reference climate. This study compares ten precipitation datasets over 1091 African catchments to evaluate dataset uncertainty contribution in climate change studies. The precipitation datasets include two gauged-only products (GPCC, CPC), four satellite products (TRMM, CHIRPS, PERSIANN-CDR and TAMSAT) corrected using ground-based observations, three reanalysis products (ERA5, ERA-I, and CFSR) and one merged product of gauge, satellite, and reanalysis (MSWEP). Each of those datasets was used to assess changes in future streamflows. The climate change impact study used a top-down modelling chain using 10 CMIP5 GCMs under RCP8.5. Each climate projection was bias-corrected and fed to a lumped hydrological model to generate future streamflows over the 2071-2100 period. A variance decomposition was performed to compare GCM uncertainty and reference dataset uncertainty for 51 streamflow metrics over each catchment. Results show that dataset uncertainty is much larger than GCM uncertainty for most of the streamflow metrics and over most of Africa. A selection of the best performing reference datasets (credibility ensemble) significantly reduced the uncertainty attributed to datasets, but remained comparable to that of GCMs in most cases. Results show also relatively small differences between datasets over a reference period can propagate to generate large amounts of uncertainty in the future climate.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.026 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.007 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.009 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".