Is the storminess in the Twentieth Century Reanalysis really inconsistent with observations? A reply to the comment by Krueger et al. (2013b)
Bibliographic record
Abstract
In a recent study of trends and low frequency variability of extra-tropical cyclone activity in the ensemble of Twentieth Century Reanalyses, we concluded that “For the North Atlantic-European region and southeast Australia, the 20CR cyclone trends are in agreement with trends in geostrophic wind extremes derived from in-situ surface pressure observations”. This conclusion has been challenged by Krueger et al. (Clim Dyn, submitted, 2013b), because a recent study (doi: 10.1175/JCLI-D-12-00309.1 , by the same lead author) comparing annual 95th percentiles (P95) of geostrophic wind speed (geo-wind) derived from surface pressure observations and from the 20CR found that “20CR-geostrophic storminess deviates to a large extent from the observation-based curve” in the period prior to 1950. In this reply, we show that our conclusion is valid; and we clarify that several factors contribute to the reported inconsistencies between the 20CR and observation-based geo-wind extremes. These include the choice of index that is used to represent the temporal variation of extremes (e.g., annual vs. seasonal percentiles), the use of different sampling intervals (6-hourly vs. 3-hourly), and the presence of very large errors in the observations that were not identified, corrected, or excluded in any of the previous studies of observation-based geo-wind extremes. We show that the time series of consecutive seasonal P95 geo-winds derived from the observations and from 20CR are in good agreement back to about 1893, with some deviation earlier when the observations (especially digitized data) remain limited and are more uncertain. We find that the correlation between the 20CR and observation-based geo-wind extremes (P95) time series for the full 134-year record is highly significant statistically, with and without the correction or exclusion of the newly identified erroneous SLP values. The agreement between 20CR and observations is further improved after the correction or exclusion of these erroneous values.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".