A nationwide regional flood frequency analysis at ungauged sites using ROI/GLS with copulas and super regions
Bibliographic record
Abstract
Region of influence is a common approach to estimate runoff information at ungauged locations. To estimate flood quantiles from annual maximum discharges, the Generalized Least Squares (GLS) framework has been recommended to account for unequal sampling variance and intersite correlation, which requires a proper evaluation of the sampling covariance structure. Since some jurisdictions do not have clear guidelines to perform this evaluation, a general procedure using copulas and a nonparametric intersite correlation model is investigated to estimate sampling covariance structure in situations where no common at-site distribution is imposed or when some paired sites do not have common periods of record. The investigated methodology is applied on 771 sites in Canada. The Normal copula is verified to be an adequate model that better fit paired observations than other types of extreme copulas. A sensitivity analysis is carried out to evaluate the impact of either ignoring, or considering a simpler form of, intersite correlation. Additionally, super regions are defined based on drainage area and mean annual precipitation to improve the calibration of pooling groups across large territories and a wide range of climate conditions. Performance criteria based on cross-validation revealed that using super regions and a combination of geographic distance and similarity between catchment descriptors improves the calibration of the pooling groups by providing more accurate estimates.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".