Environmental <scp>DNA</scp> metabarcoding in the Cape Fold aquatic ecoregion: Opportunities and challenges for <scp>eDNA</scp> uptake in an endemism hotspot
Bibliographic record
Abstract
Abstract Environmental DNA (eDNA) metabarcoding has the potential to significantly improve surveys of biodiversity in freshwater systems. However, this methodology is still infrequently used in global hotspots of endemic species, in part because of two major barriers to the success of eDNA metabarcoding: (1) insufficient regional taxonomic representation in public reference sequence databases; and (2) inconsistent species incidence and abundance estimates when compared to conventional surveys. We sampled eDNA and conducted visual surveys in the headwaters of two rivers in the Cape Fold aquatic ecoregion, South Africa. A reference sequence database was generated for the regional diversity of fishes to improve taxonomic classification of endemic species. We also compared the consistency of incidence data and relative abundance estimates of fishes from eDNA metabarcoding sequencing results and visual surveys (snorkel and underwater cameras) at each site sampled. Only 1% of eDNA metabarcoding reads could be classified to fish species without the supplementation of reference sequence databases for local endemic species. Once regional reference sequences were added, a total of nine species were detected and >99% reads classified. A strong positive relationship (Φ = 0.75) was observed between the patterns of detections using eDNA and visual approaches. However, eDNA metabarcoding detected more species than visual methods. In addition, the relative read frequency and abundance observed in visual surveys was significantly correlated (R2 = 0.85–0.88, p < 0.001) in the majority of the small pools surveyed. Incomplete public reference sequence databases hinder the use of eDNA metabarcoding in regions of high endemicity. Local reference sequence libraries can overcome these challenges. When appropriately implemented, eDNA metabarcoding shows promise, producing results consistent or better than the more frequently used and labour‐intensive visual surveys. Efficient survey methods like eDNA metabarcoding are urgently needed to improve aquatic monitoring efforts globally. Understanding how eDNA metabarcoding sequencing results relate to conventional survey methods is a key step in its implementation in ecoregions with high endemicity.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".