The geographic and phylogenetic structure of public DNA barcode databases: an assessment using Chrysomelidae (leaf beetles)
Bibliographic record
Abstract
Introduction DNA barcoding in insects has progressed rapidly, with the ultimate goal of a complete inventory of the world’s species. However, the barcoding effort to date has been driven by a few national campaigns and leaves much of the world unsampled. This study investigates to what degree the current barcode data cover the species diversity across the globe, using the leaf beetle family Chrysomelidae as an example. Methods A recent version (June 2023) of the Barcode-of-Life database was subjected to test of sampling completeness using the barcode-to-BIN ratio and sampling coverage (SC) metric. All barcodes were placed in a phylogenetic tree of ~600 mitochondrial genomes, applying phylogenetic diversity (PD) and metrics of community phylogenetics to national barcode sets to test for sampling completeness at clade level and reveal the global structure of species diversity. Results The database included 73342 barcodes, grouped into 5310 BINs (species proxies) from 101 countries. Costa Rica contributed nearly half of all barcode sequences, while nearly 50 countries were represented by less than ten barcodes. Only five countries, Costa Rica, Canada, South Africa, Germany, and Spain, had a high sampling completeness, although collectively the barcode database covers most major taxonomic and biogeographically confined lineages. PD showed moderate saturation as more species diversity is added in a country, and community phylogenetics indicated clustering of national faunas. However, at the species level the inventory remained incomplete even in the most intensely sampled countries, and the sampling was insufficient for assessment of global species richness patterns. Discussion The sequence-based inventory in Chrysomelidae needs to be greatly expanded to include more areas and deeper local sampling before reaching a knowledge base similar to the existing Linnaean taxonomy. However, placing the barcodes into a backbone phylogenetic tree from mitochondrial genomes, a taxonomically and biogeographically highly structured pattern of global diversity emerges into which all species can be integrated via their barcodes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.039 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.005 | 0.007 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".