Extremal modeling of dependent and non-stationary series and joint modeling of extreme rainfall at multiple sites
Bibliographic record
Abstract
This thesis contributes to extreme value theory; it focuses on temporal dependence of univariate extremes and compares pairwise extremal behavior in a multivariate setting. Temporal dependence in a univariate series can cause a clustering effect for extremes, i.e., extreme observations do not necessarily occur in isolation. In the univariate setting, extreme value theory can be extended to certain stationary series, however these methods do not account for temporal dependence. We consider three extreme models which account for temporal dependence: the Markov chain models of Smith et al. (1997) and Ramos and Ledford (2009), the conditional exceedance model of Heffernan and Tawn (2004), and the M3-Dirichlet model of Süveges and Davison (2012). We begin by investigating the M3-Dirichlet model in more detail through a simulation study and propose improvements related to its implementation. We then fit the three models to a precipitation series collected in Burlington, Vermont, and derive estimates of the return period of the Spring 2011 Richelieu River Valley flood in Québec, Canada. Most results from these models were found to be either unstable or too conservative. This led us to propose an alternative model for extreme cluster sums called the Random Scale Model. This model is a simple extension of the standard Peaks-Over-Threshold method, where the cluster maxima are scaled by a random factor to obtain the cluster sum. Under appropriate conditions, the random factor may be independent of the cluster maxima, which indeed is the case for the precipitation series studied. Bayesian inference is developed for this model. Our results show an adequate model fit and reasonable estimation of the return period. In the multivariate setting, we propose a new method for clustering time series based on pairwise extremal behavior. Our method has less strict assumptions on dependence structures than previous methods. We compare the performance of estimators used by our method to alternatives in a simulation study and find a distinct gain in performance for the case of asymptotic independent pairs. Each of the considered estimators are then incorporated in a clustering analysis on rainfall data from 23 stations around Québec. We found the results of our approach to be intuitive and spatially cohesive
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.011 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".