Barcoding a can of worms: testing <i>cox1</i> performance as a DNA barcode of Nematoda
Bibliographic record
Abstract
Accurate taxonomic identifications and species delimitations are a fundamental problem in biology. The complex taxonomy of Nematoda is primarily based on morphology, which is often dubious. DNA barcoding emerged as a handy tool to identify specimens and assess diversity, but its applications in Nematoda are incipient. We evaluated cytochrome c oxidase subunit I (cox1) efficiency as a DNA barcode for nematodes scrutinising 5241 sequences retrieved from BOLD and GenBank. The samples included genera with medical, agricultural, or ecological relevance: Anguillicola, Caenorhabditis, Heterodera, Meloidogyne, Onchocerca, Strongyloides, and Trichinella. We assessed cox1 performance through barcode gap and Probability of Correct Identification (PCI) analyses, and estimated species richness through Automatic Barcode Gap Discovery (ABGD). Each genus presented distinct gap ranges, mirroring the evolutionary diversity within Nematoda. Thus, to survey the diversity of the phylum, a careful definition of thresholds for lower taxonomic levels should be considered. PCIs were around 70% for both databases, highlighting operational biases and challenges in nematode taxonomy. ABGD inferred higher richness than the taxonomic labels informed by databases. The prevalence of specimen misidentifications and dubious species delimitations emphasise the value of integrative approaches to nematode taxonomy and systematics. Overall, cox1 is a relevant tool for integrative taxonomy of nematodes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".