Barcoding ciliates: a comprehensive study of 75 isolates of the genus Tetrahymena
Bibliographic record
Abstract
The mitochondrial cytochrome-c oxidase subunit 1 (cox1) gene has been proposed as a DNA barcode to identify animal species. To test the applicability of the cox1 gene in identifying ciliates, 75 isolates of the genus Tetrahymena and three non-Tetrahymena ciliates that are close relatives of Tetrahymena, Colpidium campylum, Colpidium colpoda and Glaucoma chattoni, were selected. All tetrahymenines of unproblematic species could be identified to the species level using 689 bp of the cox1 sequence, with about 11 % interspecific sequence divergence. Intraspecific isolates of Tetrahymena borealis, Tetrahymena lwoffi, Tetrahymena patula and Tetrahymena thermophila could be identified by their cox1 sequences, showing <0.65 % intraspecific sequence divergence. In addition, isolates of these species were clustered together on a cox1 neighbour-joining (NJ) tree. However, strains identified as Tetrahymena pyriformis and Tetrahymena tropicalis showed high intraspecific sequence divergence values of 5.01 and 9.07 %, respectively, and did not cluster together on a cox1 NJ tree. This may indicate the presence of cryptic species. The mean interspecific sequence divergence of Tetrahymena was about 11 times greater than the mean intraspecific sequence divergence, and this increased to 58 times when all isolates of species with high intraspecific sequence divergence were excluded. This result is similar to DNA barcoding studies on animals, indicating that congeneric sequence divergences are an order of magnitude greater than conspecific sequence divergences. Our analysis also demonstrated low sequence divergences of <1.0 % between some isolates of T. pyriformis and Tetrahymena setosa on the one hand and some isolates of Tetrahymena furgasoni and T. lwoffi on the other, suggesting that the latter species in each pair is a junior synonym of the former. Overall, our study demonstrates the feasibility of using the mitochondrial cox1 gene as a taxonomic marker for 'barcoding' and identifying Tetrahymena species and some other ciliated protists.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".