Optimization of environmental DNA-based methods: A case study for detecting brook trout (Salvelinus fontinalis).
Bibliographic record
Abstract
The utility of eDNA for fish species and community monitoring is well-established using targeted amplification (i.e., qPCR and ddPCR) and passive sequencing approaches (i.e., metabarcoding). However, the lack of optimized and standardized methods reduces the sensitivity of this approach and precludes the reliable comparison of findings across studies, respectively. DNA extraction is a prime target for optimization efforts because the extraction method is highly variable across eDNA studies despite being the most influential factor in detection efficiency across the entire post-collection workflow. Sequence analysis is arguably the least standardized step in the workflow, with new bioinformatics pipelines frequently emerging in the literature and being implemented with innumerable unique combinations of parameter values. The current study aimed to support the optimization and standardization of eDNA methods for fish detection by assessing two commercial DNA extraction kits manufactured by Qiagen and Macherey-Nagel on cost, time, and performance specifications and comparing the success of brook trout detection by metabarcoding across three bioinformatics pipelines, qPCR, and ddPCR. Our protocols were effective in detecting brook trout in all 20 samples analyzed. Brook trout eDNA was detected by ddPCR in nine (90%) Qiagen extracts but only seven (70%) Macherey-Nagel extracts. In comparison, detection success was equal across the two extraction kits using qPCR (70%) and metabarcoding (100%). The metabarcoding pipelines performed equally well in detecting brook trout with no significant differences in read numbers associated with the target species. Under our experimental conditions, the Qiagen kit was selected as the preferred kit due to its overall good performance and considerably lower cost despite a slightly longer extraction time.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".