A global analysis of the influence of shallow and deep groundwater tables on relationships between environmental parameters and heatwaves
Bibliographic record
Abstract
Heatwaves increasingly impact ecosystems, human health, and economic activities worldwide. As their frequency and intensity rise, understanding the mechanisms driving heatwave dynamics and interactions with land surface processes becomes crucial. While numerous studies have examined atmospheric and land surface variables, the role of groundwater, through its effects on soil moisture and surface evaporative fluxes, remains less understood. Although modeling approaches at various scales have enhanced our understanding of groundwater-atmosphere coupling, machine learning (ML) enables capturing complex, nonlinear interactions and evaluating the relative importance of key drivers globally. We developed pixel-based ML models to estimate global summer heatwave frequency over the past 21 years. For each pixel, we considered data within a 1.5° radius (149 neighboring pixels), identified as the optimal scale through a saturation radius analysis. We used feature importance metrics to identify the dominant drivers among surface fluxes, land characteristics, atmospheric and hydrological variables, and interpreted these results in relation to contrasting groundwater depths (<10 m and >100 m). We ensured robustness using 10-fold cross-validation and confirmed that results were not driven by randomness with two additional validation runs on a subset of the data, with shuffled targets and randomized covariates. Our findings suggest that geopotential height showed the highest relative importance among predictors in regions with deep groundwater tables, while in areas with shallow groundwater, surface fluxes emerge as the key contributor. Incorporating groundwater-related processes may therefore improve understanding of land-atmosphere interactions and support more robust assessments of future heatwave risks.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".