Phylogenomic diversity of archigregarine apicomplexans
Bibliographic record
Abstract
Gregarines are a large and diverse subgroup of Apicomplexa, a lineage of obligate animal symbionts including pathogens such as Plasmodium , the malaria parasite. Unlike Plasmodium , however, gregarines are poorly studied, despite the fact that as early-branching apicomplexans they are crucial to our understanding of the origin and evolution of all apicomplexans and their parasitic lifestyle. Exemplifying this, the earliest branch of gregarines, the archigregarines, are particularly poorly studied: around 80 species have been described from marine invertebrates, but almost all of them were assigned to a single genus, Selenidium . Most are known only from light micrographs and largely unresolved rDNA phylogenies, where they exhibit a great deal of sequence variation, and fall into four subclades. To resolve the relationships within archigregarines, we sequenced 12 single-cell transcriptomes from species representing all four known subclades, as well as one blastogregarine (which frequently branch with Selenidium ). A 190-gene phylogenomic tree confirmed four maximally supported individual clades of archigregarines and blastogregarines. These clades are discrete and distantly related, and also correlate with host identity. We propose the establishment of three novel genera of archigregarines to reflect their phylogenetic diversity and host range, and nine novel species isolated from a range of marine invertebrates.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".