Both biotic and abiotic predictors explain significant variation in cyanobacteria biomass across lakes from temperate to subarctic zones
Bibliographic record
Abstract
Abstract The development of cyanobacteria blooms is of increasing concern in many lakes worldwide, and as a result, modeling their predictors is vital for understanding where and why they occur. In this study, we developed and analyzed a 640‐lake data set that spans Canada and 12 ecozones to identify the drivers of cyanobacteria biomass and of several key toxin‐ and bloom‐forming genera (Microcystis, Aphanizomenon, and Dolichospermum). The database consisted of an exhaustive list of potential predictors (n = 55), including water chemistry, land‐use, and zooplankton variables. We applied a series of empirical modeling approaches to identify significant predictors and thresholds (generalized linear and additive models, mixed effect regression trees), all while accounting for ecozone variability. Across all modeling approaches, and ecozones total phosphorus was identified as the most important predictor of total cyanobacterial and focal genera biomass. In addition, cyanobacteria across Canada showed significant associations with increasing dissolved organic and inorganic carbon, and several ions. Despite the widely held notion that cyanobacteria are often toxic and/or a poor food source for zooplankton, we found a positive relationship between cyanobacteria and zooplankton, particularly with daphnid and copepod biomass. Localized top‐down forces and evolutionary adaptations resulting from long‐term exposure in eutrophic lakes are among the possible explanations for this observed positive association. By considering a suite of complementary modeling approaches, we found that nonlinear models provided greater predictive power and the random ecozone effect was minor due to the overarching importance of local abiotic and biotic factors.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".