Improved Modeling of Compositional Heterogeneity Supports Sponges as Sister to All Other Animals
Bibliographic record
Abstract
The relationships at the root of the animal tree have proven difficult to resolve, with the current debate focusing on whether sponges (phylum Porifera) or comb jellies (phylum Ctenophora) are the sister group of all other animals [1-5]. The choice of evolutionary models seems to be at the core of the problem because Porifera tends to emerge as the sister group of all other animals ("Porifera-sister") when site-specific amino acid differences are modeled (e.g., [6, 7]), whereas Ctenophora emerges as the sister group of all other animals ("Ctenophora-sister") when they are ignored (e.g., [8-11]). We show that two key phylogenomic datasets that previously supported Ctenophora-sister [10, 12] display strong heterogeneity in amino acid composition across sites and taxa and that no routinely used evolutionary model can adequately describe both forms of heterogeneity. We show that data-recoding methods [13-15] reduce compositional heterogeneity in these datasets and that models accommodating site-specific amino acid preferences can better describe the recoded datasets. Increased model adequacy is associated with significant topological changes in support of Porifera-sister. Because adequate modeling of the evolutionary process that generated the data is fundamental to recovering an accurate phylogeny [16-20], our results strongly support sponges as the sister group of all other animals and provide further evidence that Ctenophora-sister represents a tree reconstruction artifact. VIDEO ABSTRACT.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.019 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".