Prospects for fungus identification using <i>CO1</i> DNA barcodes, with <i>Penicillium</i> as a test case
Bibliographic record
Abstract
DNA barcoding systems employ a short, standardized gene region to identify species. A 648-bp segment of mitochondrial cytochrome c oxidase 1 (CO1) is the core barcode region for animals, but its utility has not been tested in fungi. This study began with an examination of patterns of sequence divergences in this gene region for 38 fungal taxa with full CO1 sequences. Because these results suggested that CO1 could be effective in species recognition, we designed primers for a 545-bp fragment of CO1 and generated sequences for multiple strains from 58 species of Penicillium subgenus Penicillium and 12 allied species. Despite the frequent literature reports of introns in fungal mitochondrial genomes, we detected introns in only 2 of 370 Penicillium strains. Representatives from 38 of 58 species formed cohesive assemblages with distinct CO1 sequences, and all cases of sequence sharing involved known species complexes. CO1 sequence divergences averaged 0.06% within species, less than for internal transcribed spacer nrDNA or beta-tubulin sequences (BenA). CO1 divergences between species averaged 5.6%, comparable to internal transcribed spacer, but less than values for BenA (14.4%). Although the latter gene delivered higher taxonomic resolution, the amplification and alignment of CO1 was simpler. The development of a barcoding system for fungi that shares a common gene target with other kingdoms would be a significant advance.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.011 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.003 | 0.006 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.005 | 0.002 |
| Insufficient payload (model declined to judge) | 0.006 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".