Hidden diversity – DNA metabarcoding reveals hyper-diverse benthic invertebrate communities
Bibliographic record
Abstract
Abstract Freshwater ecosystems, such as streams, are facing increasing pressures from agricultural land use. Aquatic insects and other macroinvertebrates have historically been used as indicators of ecological condition and water quality in freshwater biomonitoring programs; however, many of these protocols use coarse taxonomic resolution (e.g., family) when identifying macroinvertebrates. The use of family-level identification can mask species-level diversity, as well as patterns in community composition in response to environmental variables. Recent literature stresses the importance of robust biomonitoring to detect trends in insect decline globally, though most of these studies are carried out in terrestrial habitats. Here, we incorporate molecular identification (DNA metabarcoding) into a stream biomonitoring sampling design to explore the diversity and variability of aquatic macroinvertebrate communities at small spatial scales. We sampled twenty southern Ontario streams in an agricultural landscape for aquatic macroinvertebrates and, using DNA metabarcoding, revealed incredibly rich benthic communities which were largely comprised of rare taxa detected only once per stream despite multiple biological replicates. In addition to numerous rare taxa, our species pool estimates indicated that after 240 samples from twenty streams, there was a large proportion of taxa present which remained undetected by our sampling regime. When comparing different levels of taxonomic resolution, we observed that using OTUs revealed over ten times more taxa than family-level identification. A single insect family, the Chironomidae, contained over one third of the total number of OTUs detected in our study. Within-stream dissimilarity estimates were consistently high for all taxonomic groups (invertebrate families, invertebrate OTUs, chironomid OTUs), indicating stream communities are very dissimilar at small spatial scales. While we predicted that increased land use would homogenize benthic communities, this was not supported as within-stream dissimilarity was unrelated to land use.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".