On the need for sub-daily data to study changes in extreme rainfall.
Bibliographic record
Abstract
Heavy rainfall is among the most impactful natural events. Our understanding of such events has improved significantly in the last decades, but large uncertainties remain around their recent and future response to a changing climate. At global scales, the frequency and intensity of daily extreme precipitation has increased, the hydrological cycle is becoming faster. However, the response at regional scales and shorter timescales is much more complex. The study of sub-daily or even sub-hourly data has been explored to some extent only, mostly due to the limited availability of data. When using high-resolution models to explore rainfall changes, it is possible to examine much higher frequencies, yet most studies focus on daily rainfall changes. Here, we demonstrate inherent limitations of daily data to study present and future precipitation extremes. Limitations that are not purely a matter of refining our sampling, but do have a physical background because outstanding rainfall rates rarely occur over the course of a day. Our results show that fundamental aspects of rainfall changes are not described with daily data, and the assessment of future changes in daily precipitation likely leads to misrepresentation of causes and impacts. We show that the short-lived and intermittent nature of most rainfall extremes need at least hourly data to be properly characterized, otherwise heavy rainfall is poorly detected. Analyzing higher frequencies also reveals aspects of extremes that cannot be addressed with daily data, such as changes in their intensity and duration. This is particularly relevant for risk and impact assessment studies because a significant part of changes in extremes occur at sub-daily scales. Such changes go unnoticed or, even worse, are misrepresented by daily rainfall amounts.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.011 | 0.038 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.002 | 0.005 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.005 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.011 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".