Sculpting the endomembrane system in deep time: High resolution phylogenetics of Rab GTPases
Bibliographic record
Abstract
The presence of a nucleus and other membrane-bounded intracellular compartments is the defining feature of eukaryotic cells. Endosymbiosis accounts for the origins of mitochondria and plastids, but the evolutionary ancestry of the remaining cellular compartments is incompletely documented. Resolving the evolutionary history of organelle-identity encoding proteins within the endomembrane system is a necessity for unravelling the origins and diversification of the endogenously derived organelles. Comparative genomics reveals events after the last eukaryotic common ancestor (LECA), but resolution of events prior to LECA, and a full account of the intracellular compartments present in LECA, has proved elusive. We have devised and exploited a new phylogenetic strategy to reconstruct the history of the Rab GTPases, a key family of endomembrane-specificity proteins. Strikingly, we infer a remarkably sophisticated organellar composition for LECA, which we predict possessed as many as 23 Rab GTPases. This repertoire is significantly greater than that present in many modern organisms and unexpectedly indicates a major role for secondary loss in the evolutionary diversification of the endomembrane system. We have identified two Rab paralogues of unknown function but wide distribution, and thus presumably ancient nature; RabTitan and RTW. Furthermore, we show that many Rab paralogues emerged relatively suddenly during early metazoan evolution, which is in stark contrast to the lack of significant Rab family expansions at the onset of most other major eukaryotic groups. Finally, we reconstruct higher-order ancestral clades of Rabs primarily linked with endocytic and exocytic process, suggesting the presence of primordial Rabs associated with the establishment of those pathways and giving the deepest glimpse to date into pre-LECA history of the endomembrane system.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".