Three years of monitoring severe hailswaths across Canada using radar
Bibliographic record
Abstract
One of the primary goals of the Northern Hail Project (NHP) is to generate the first detailed hail climatology for Canada. To do this, we are adopting several independent approaches. One of these is the use of the Maximum Expected Size of Hail (MESH) from radar data. Most of the 33 new S-band dual-polarization radars have been operational since early 2022. Using these radar data allowed us to identify hailstorms at high-spatial temporal resolution wherever we have radar data. In this research, we focused on manually identifying (and digitizing) severe hailswaths using the MESH data from 2022 through 2024. A severe MESH hailswath is one that has a continuous 40 km long or greater track of 10 mm pixels, and must include at least 2 adjacent pixels of 30 mm or greater. A total of almost 2,000 severe hailswaths have been identified by radar to date. Although the regional year-to-year variability is significant, our analysis has identified the Canadian Prairies and far western Ontario as hot spots for long-lived, severe hailstorms. Some of the severe hailswaths in the dataset are impressive, extending over 500 km and lasting up to 6 hours. The widest hailswath in our MESH dataset is approximately 50 km across. Even though most of the MESH hailswaths in our database have occurred near or just to the north of the Canadian/U.S border, some hailswaths have occurred at the edge of our available radar network range, with the most northern MESH hailswath terminating at a latitude of 58.0 degrees north in Saskatchewan. Moving forward, we will continue to monitor severe hailswaths using a semi-automated algorithm that draws on machine vision and machine learning techniques and will be trained on the existing dataset. In 2022, we used the MESH product produced by the National Oceanic and Atmospheric Administration-Multi-Radar Multi-Sensor (NOAA-MRMS) and switched to the Environment and Climate Change Canada (ECCC) MESH product in 2023 and 2024. Although the products from the two groups are generally in good agreement, we noted that there are notable differences in the MESH values at times. Reasons for these discrepancies are being investigated using ground reference data collected by the NHP.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".