A stable phylogeny of the large‐spored <i>Metschnikowia</i> clade
Bibliographic record
Abstract
Draft genomes of 55 strains representing all known large-spored Metschnikowia species were used to construct a robust phylogeny of these yeasts found in association with flower-visiting insects. The genomes were annotated with reference to Clavispora lusitaniae. From 3016 orthologues identified, 1061 were present in all strains with enough overlap to generate alignments of 500 bp or more. We constructed trees for all those alignments and evaluated their accuracy from their ability to resolve each of 22 sets of conspecifics correctly as sister taxa. Neighbour-joining identified species membership better than maximum likelihood, as did trees based on larger gene alignments. However, correct species assignment was not predictive of a gene's ability to resolve deeper topologies, which were more reliably identified by maximum likelihood analyses of large concatenations. Specifically, 14 trees based on independent concatenations ca. 100 kb in length were topologically consistent with a tree based on a single, large concatenation (1 410 065 positions), lending a high degree of confidence to the stability of the phylogeny. A tree based on a concatenation of intergenic regions (112 136 positions) was also congruent. Again, the best predictor of phylogenetic signal quality of a gene was the size of the alignment. Bootstraps were not always good indicators of phylogenetic quality, as they were sometimes affected by clade size. A tree constructed from a presence-absence matrix of all annotated genes was remarkably congruent with sequence-based phylogenies, suggesting that gain or loss of genes is worth exploring further as a phylogenetically significant event. Copyright © 2016 John Wiley & Sons, Ltd.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".