On the monophyly of chromalveolates using a six-protein phylogeny of eukaryotes
Bibliographic record
Abstract
A global phylogeny of major eukaryotic lineages is a significant and ongoing challenge to molecular phylogenetics. Currently, there are five hypothesized major lineages or 'supergroups' of eukaryotes. One of these, the chromalveolates, represents a large fraction of protist and algal diversity. The chromalveolate hypothesis was originally based on similarities between the photosynthetic organelles (plastids) found in many of its members and has been supported by analyses of plastid-related genes. However, since plastids can move between eukaryotic lineages, it is important to provide additional support from data generated from the nuclear-cytosolic host lineage. Genes coding for six different cytosolic proteins from a variety of chromalveolates (yielding 68 new gene sequences) have been characterized so that multiple gene analyses, including all six major lineages of chromalveolates, could be compared and concatenated with data representing all five hypothesized supergroups. Overall support for much of the phylogenies is decreased over previous analyses that concatenated fewer genes for fewer taxa. Nevertheless, four of the six chromalveolate lineages (apicomplexans, ciliates, dinoflagellates and heterokonts) consistently form a monophyletic assemblage, whereas the remaining two (cryptomonads and haptophytes) form a weakly supported group. Whereas these results are consistent with the monophyly of chromalveolates inferred from plastid data, testing this hypothesis is going to require a substantial increase in data from a wide variety of organisms.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".