Capture Probe, Metabarcoding, or Shotgun Sequencing: Which Best Reflects Local Vegetation?
Bibliographic record
Abstract
ABSTRACT Metabarcoding is the most widely applied method for studying plant communities using environmental DNA, with shotgun sequencing and capture probes being alternative methods that aim to retrieve multiple markers or genome‐wide information. Any method's ability to detect and correctly identify plant taxa varies with DNA preservation, DNA reference library, and the diversity of the local flora, making it difficult to compare results from different environments. Here we compare these three methods using lake surface‐sediments from Northern Fennoscandia with the PhyloNorway genome skim reference library (1500 taxa) that includes nearly all species of the regional flora. We also undertook vegetation surveys from around the lakes to estimate the true positive detection rate, identify false positive detections, and provide optimal filtering cut‐off thresholds for the three methods. Applying these thresholds, the rate of false positives was too high for reliable identification at the species level based on shotgun (49%) and capture probes (62%), whereas it was low for metabarcoding (5%–12%). All methods were reliable at genus and family levels after applying the optimal filtering thresholds (< 4% false positives). Our results show that in these lake sediments, metabarcoding on average detects 2.1 times as many true positive taxa as shotgun sequencing and 6.4 times as many taxa as capture probes. The proportion of a taxon's sequenced reads for the metabarcoding and shotgun methods was significantly related to the taxon's abundance category from the vegetation surveys, but this was not the case for capture probe data. We expect the false positive rate of shotgun sequencing to decrease with increasing genome completeness in the reference libraries and the method to be advantageous for highly degraded DNA with fragments too short for metabarcoding. At present, metabarcoding provides the highest detectability and taxonomic resolution for correct identification and quantification of vascular plants.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.008 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".