Globally Observed Annual Extreme Daily And Persistent Precipitation Relative Totals
Bibliographic record
Abstract
To provide the most comprehensive analysis of observed global extreme daily and persistent precipitation, we use high-quality daily precipitation data from a number of different sources. Co-authors from fifteen countries contributed daily data, most of which until now were not available for global precipitation studies. The compilation of global daily precipitation data includes the GHCND dataset (https://www.ncdc.noaa.gov/ghcn-daily-description), the ECA&D dataset (https://www.ecad.eu/), the USHCN dataset (http://cdiac.ess-dive.lbl.gov/ftp/ushcn_daily/), and the dataset for Canada (http://climate.weather.gc.ca/), raw data provided by authors from Argentina, Australia, Benin, Brazil, China, India, Japan, Korea, Mongolia, New Zealand, Pakistan, South Africa, Saudi Arabia, Spain, and Russia. In total 12151 stations were collated. After quality control and homogeneity test, 6125 high-quality stations with long-term (data are available at least for 45 years) daily precipitation for the period 1961-2010 are remained. The 95th percentile of daily and persistent precipitation series on wet days (≥ 1 mm) is used to identify daily and persistent extremes, respectively. The base period for percentile calculation is 1961-2010. Considering regional precipitation characteristics, the ‘relative total’ used here is not the simple precipitation amount, but a relative measure (%) associated with the local threshold of ‘extremity’ (i.e. the 95th percentile) and the total extreme precipitation amount. The relative total of extreme precipitation is defined as the mean precipitation amount exceeding the threshold divided by the corresponding threshold. This dataset contains the annual extreme precipitation relative totals for the 6125 stations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.002 | 0.005 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".