Choice of DNA extraction method affects DNA metabarcoding of unsorted invertebrate bulk samples
Bibliographic record
Abstract
Characterisation of freshwater benthic biodiversity using DNA metabarcoding may allow more cost-effective environmental assessments than the current morphological-based assessment methods. DNA metabarcoding methods where sorting or pre-sorting of samples are avoided altogether are especially interesting, since the time between sampling and taxonomic identification is reduced. Due to the presence of non-target material like plants and sediments in crude samples, DNA extraction protocols become important for maximising DNA recovery and sample replicability. We sampled freshwater invertebrates from six river and lake sites and extracted DNA from homogenised bulk samples in quadruplicate subsamples, using a published method and two commercially available kits: HotSHOT approach, Qiagen DNeasy Blood & Tissue Kit and Qiagen DNeasy PowerPlant Pro Kit. The performance of the selected extraction methods was evaluated by measuring DNA yield and applying DNA metabarcoding to see if the choice of DNA extraction method affects DNA yield and metazoan diversity results. The PowerPlant Kit extractions resulted in the highest DNA yield and a strong significant correlation between sample weight and DNA yield, while the DNA yields of the Blood & Tissue Kit and HotSHOT method did not correlate with the sample weights. Metazoan diversity measures were more repeatable in samples extracted with the PowerPlant Kit compared to those extracted with the HotSHOT method or the Blood & Tissue Kit. Subsampling using Blood & Tissue Kit and HotSHOT extraction failed to describe the same community in the lake samples. Our study exemplifies that the choice of DNA extraction protocol influences the DNA yield as well as the subsequent community analysis. Based on our results, low specimen abundance samples will likely provide more stable results if specimens are sorted prior to DNA extraction and DNA metabarcoding, but the repeatability of the DNA extraction and DNA metabarcoding results was close to ideal in high specimen abundance samples.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".