Understanding the Drivers of Drought Onset and Intensification in the Canadian Prairies: Insights from Explainable Artificial Intelligence (XAI)
Bibliographic record
Abstract
Abstract Recent advances in artificial intelligence (AI) and explainable AI (XAI) have created opportunities to better predict and understand drought processes. This study uses a machine learning approach for understanding the drivers of drought severity and extent in the Canadian Prairies from 2005 to 2019 using climate and satellite data. The model is trained on the Canadian Drought Monitor (CDM), an extensive dataset produced by expert analysis of drought impacts across various sectors that enables a more comprehensive understanding of drought. Shapley additive explanation (SHAP) is used to understand model predictions during emerging or worsening drought conditions, providing insight into the key determinants of drought. The results demonstrate the importance of capturing spatiotemporal autocorrelation structures for accurate drought characterization and elucidates the drought time scales and thresholds that optimally separate each CDM severity category. In general, there is a positive relationship between the severity of drought and the time scale of the anomalies. However, high-severity droughts are also more complex and driven by a multitude of factors. It was found that satellite-based evaporative stress index (ESI), soil moisture, and groundwater were effective predictors of drought onset and intensification. Similarly, anomalous phases of large-scale atmosphere–ocean dynamics exhibit teleconnections with Prairie drought. Overall, this investigation provides a better understanding of the physical mechanisms responsible for drought in the Prairies, provides data-driven thresholds for estimating drought severity that could improve future drought assessments, and offers a set of early warning indicators that may be useful for drought adaptation and mitigation. Significance Statement This work is significant because it identifies drivers of drought onset and intensification in an agriculturally and economically important region of Canada. This information can be used in the future to improve early warning for adaptation and mitigation. It also uses state-of-the-art machine learning techniques to understand drought, including a novel approach called SHAP probability values to improve interpretability. This provides evidence that machine learning models are not black boxes and should be more widely considered for understanding drought and other hydrometeorological phenomena.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".