Capture Probe, Metabarcoding, or Shotgun Sequencing: Which Best Reflects Local Vegetation?
Bibliographic record
Abstract
ABSTRACT Metabarcoding is the most widely applied method for studying plant communities using environmental DNA, with shotgun sequencing and capture probes being alternative methods that aim to retrieve multiple markers or genome‐wide information. Any method's ability to detect and correctly identify plant taxa varies with DNA preservation, DNA reference library, and the diversity of the local flora, making it difficult to compare results from different environments. Here we compare these three methods using lake surface‐sediments from Northern Fennoscandia with the PhyloNorway genome skim reference library (1500 taxa) that includes nearly all species of the regional flora. We also undertook vegetation surveys from around the lakes to estimate the true positive detection rate, identify false positive detections, and provide optimal filtering cut‐off thresholds for the three methods. Applying these thresholds, the rate of false positives was too high for reliable identification at the species level based on shotgun (49%) and capture probes (62%), whereas it was low for metabarcoding (5%–12%). All methods were reliable at genus and family levels after applying the optimal filtering thresholds (< 4% false positives). Our results show that in these lake sediments, metabarcoding on average detects 2.1 times as many true positive taxa as shotgun sequencing and 6.4 times as many taxa as capture probes. The proportion of a taxon's sequenced reads for the metabarcoding and shotgun methods was significantly related to the taxon's abundance category from the vegetation surveys, but this was not the case for capture probe data. We expect the false positive rate of shotgun sequencing to decrease with increasing genome completeness in the reference libraries and the method to be advantageous for highly degraded DNA with fragments too short for metabarcoding. At present, metabarcoding provides the highest detectability and taxonomic resolution for correct identification and quantification of vascular plants.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.004 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".