Anthony C. Davison and Raphaël de Fondeville’s contribution to the Discussion of ‘Inference for extreme spatial temperature events in a changing climate with application to Ireland’ by Healy et al.
Bibliographic record
Abstract
We congratulate the authors on their paper, which puts some of the ideas in de Fondeville and Davison (2018) to good use and suggests innovative approaches to dealing with issues not previously discussed in the r-Pareto context, such as the effect of missing data and the extrapolation of station data to unmonitored locations. Major environmental events due to ‘heat domes’ seem to have become more common. The corresponding heatwaves are spatially large and can smash previous temperature records, as happened in 2021 in North America, when records were set across western Canada. Naive fits of extreme-value models to temperature maxima invariably result in a negative estimated shape parameter, as shown in Table 2 of the paper, which implies that the range of future values has an upper bound. Such a bound based on analysis of the previous data suggested that the 2021 event would be impossible, but of course it happened. One way to deal with this would be to construct a mixture distribution, perhaps with event probabilities dependent on climatic and/or meteorological variables. Major heatwaves are often associated with blocking anticyclones: would it be feasible to use such events to construct a more complex, but perhaps more realistic, model, rather than treating heatwaves as identically distributed? Otherwise we might insert the knowledge that such events might arise, for example using a penalized likelihood when estimating the shape parameter: do the authors think this might be helpful? Heatwaves are typically defined in terms of a succession of particularly hot days. In peaks-over-threshold approaches, temporal dependence can bias the estimation of marginal parameters, potentially creating situations such as those described above. Selecting a risk functional for which extremes correspond to events whose intensity is both marginally significant and proportional to the event’s duration would allow one to focus on specific types of physical processes. With such an approach, the classical modeling strategy of ‘one location, one tail index’ shifts to ‘one physical phenomenon, one tail index’, making the assumption of common shape parameters more defensible. The presence of temporal dependence and a potential mixture of physical processes might make the identically-distributed assumption used for statistical modeling unrealistic. Did the authors find evidence of multiple physical processes governing extreme temperatures in Ireland? If so, do they think their approach could benefit from the suggestions above?
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.103 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.002 | 0.005 |
| Scholarly communication | 0.005 | 0.007 |
| Open science | 0.005 | 0.004 |
| Research integrity | 0.008 | 0.014 |
| Insufficient payload (model declined to judge) | 0.011 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".