Predicting microcystin concentrations in lakes and reservoirs at a continental scale: A new framework for modelling an important health risk factor
Bibliographic record
Abstract
Abstract Aim Scientists, governments and non‐governmental organizations are increasingly moving towards the collection of large, open‐access data. In aquatic sciences, this effort is expanding the scope of questions and analyses that can be performed to further our knowledge of the global drivers of water quality. Cyanotoxin concentration is one variable that has received considerable attention, and although strong local‐scale models have been described in the literature, modelling cyanotoxin concentrations across broader spatial scales has been more difficult. Commonly used statistical frameworks have not fully captured the complex response of toxic algal blooms to global change, limiting our ability to predict and mitigate the impairment of freshwaters by toxic algae. Here, we advance our understanding of emergent drivers of cyanotoxins across a structured landscape by applying a hierarchical “hurdle” model. Location Lakes and reservoirs in the conterminous United States [ n = 1127]. Methods We studied cyanobacteria and their toxins [microcystins] during the 2007 summer period. We applied a hierarchical zero‐altered model to test the importance of multi‐scale interactions among environmental features in driving microcystin concentrations above the limit of detection. We then used boosted regression trees [BRTs] to identify environmental thresholds associated with severe impairment by microcystins. Results Accounting for numerous non‐detections, spatial heterogeneity and cross‐scale interactions substantially improved continental‐scale predictions of bloom toxicity. Our model accounted for 55% of the variance in the probability of detecting microcystins across the United States, and 26% of the variability in microcystin concentrations once detected. BRTs further showed that although both local and regional drivers were associated with microcystin concentrations at low to intermediate provisional guidelines, only local drivers came into play when predicting higher limits. Main conclusions Identifying the interaction between local and regional processes is key to understanding the heterogeneous responses of microcystins to environmental change. Our framework could increase the effectiveness of continental‐scale analyses for many different water variables.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".