DNA barcode accumulation curves for understudied taxa and areas
Bibliographic record
Abstract
Frequently, the diversity of umbrella taxa is invoked to predict patterns of other, less well-known, life. However, the utility of this strategy has been questioned. We tested whether a phylogenetic diversity (PD) analysis of CO1 DNA barcodes could act as a proxy for standard methods of determining sampling efficiency within and between sites, namely that an accumulation curve of barcode diversity would be similar to curves generated using morphology or nuclear genetic markers. Using taxa at the forefront of the taxonomic impediment - parasitoid wasps (Ichneumonidae, Braconidae, Cynipidae and Diapriidae), contrasted with a taxon expected to be of low diversity (Formicidae) from an area where total diversity is expected to be low (Churchill, Manitoba), we found that barcode accumulation curves based on PD were significantly different in both slope and scale from curves generated using names based on morphological data, while curves generated using nuclear genetic data were only different in scale. We conclude that these differences clearly identify the taxonomic impediment within the strictly morphological alpha-taxonomy of these hyperdiverse insects. The absence of an asymptote within the barcode PD trend of parasitoid wasps reflects the as yet incomplete sampling of the site (and more accurately its total diversity), while the morphological analysis asymptote represents a collision with the taxonomic impediment rather than complete sampling. We conclude that a PD analysis of standardized DNA barcodes can be a transparent and reproducible triage tool for the management and conservation of species and spaces.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".