DNA barcoding cannot reliably identify species of the blowfly genus<i>Protocalliphora</i>(Diptera: Calliphoridae)
Bibliographic record
Abstract
In DNA barcoding, a short standardized DNA sequence is used to assign unknown individuals to species and aid in the discovery of new species. A fragment of the mitochondrial gene cytochrome c oxidase subunit 1 is emerging as the standard barcode region for animals. However, patterns of mitochondrial variability can be confounded by the spread of maternally transmitted bacteria that cosegregate with mitochondria. Here, we investigated the performance of barcoding in a sample comprising 12 species of the blow fly genus Protocalliphora, known to be infected with the endosymbiotic bacteria Wolbachia. We found that the barcoding approach showed very limited success: assignment of unknown individuals to species is impossible for 60% of the species, while using the technique to identify new species would underestimate the species number in the genus by 75%. This very low success of the barcoding approach is due to the non-monophyly of many of the species at the mitochondrial level. We even observed individuals from four different species with identical barcodes, which is, to our knowledge, the most extensive case of mtDNA haplotype sharing yet described. The pattern of Wolbachia infection strongly suggests that the lack of within-species monophyly results from introgressive hybridization associated with Wolbachia infection. Given that Wolbachia is known to infect between 15 and 75% of insect species, we conclude that identification at the species level based on mitochondrial sequence might not be possible for many insects. However, given that Wolbachia-associated mtDNA introgression is probably limited to very closely related species, identification at the genus level should remain possible.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".