A high density <i>COX1</i> barcode oligonucleotide array for identification and detection of species of <i>Penicillium</i> subgenus <i>Penicillium</i>
Bibliographic record
Abstract
We developed a COX1 barcode oligonucleotide array based on 358 sequences, including 58 known and two new species of Penicillium subgenus Penicillium, and 12 allied species. The array was robotically spotted at near microarray density on membranes. Species and clade-specific oligonucleotides were selected using the computer programs SigOli and Array Designer. Robotic spotting allowed 768 spots with duplicate sets of perfect match and the corresponding mismatch and positive control oligonucleotides, to be printed on 2 × 6 cm(2) nylon membranes. The array was validated with hybridizations between the array and digoxigenin (DIG)-labelled COX1 polymerase chain reaction amplicons from 70 pure DNA samples, and directly from environmental samples (cheese and plants) without culturing. DNA hybridization conditions were optimized, but undesired cross-reactions were detected frequently, reflecting the relatively high sequence similarity of the COX1 gene among Penicillium species. Approximately 60% of the perfect match oligonucleotides were rejected because of low specificity and 76 delivered useful group-specific or species-specific reactions and could be used for detecting certain species of Penicillium in environmental samples. In practice, the presence of weak signals on arrays exposed to amplicons from environmental samples, which could have represented weak detections or weak cross reactions, made interpretation difficult for over half of the oligonucleotides. DNA regions with very few single nucleotide polymorphisms or lacking insertions/deletions among closely related species are not ideal for oligonucleotide-based diagnostics, and supplementing the COX1-based array with oligonucleotides derived from additional genes would result in a more robust hierarchical identification system.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".