Pseudogenization of the Chaperonin System in ‘<em>Candidatus</em> Phytoplasma pruni’ Reveals Insights into the Role of GroEL/Cpn60 in Phytopathogenic Mollicutes
Bibliographic record
Abstract
GroE is a chaperonin folding system consisting of GroEL (Cpn60, a 60 kDa chaperonin), and the smaller co-chaperonin GroES (Cpn10). Many “client” proteins require GroE to fold properly, including several that are essential for cell viability. Unsurprisingly then, GroE is found in nearly all bacteria and eukaryotes. Mollicutes are the only microorganisms that lack GroE in almost all cases. Only two clades of Mollicutes have retained the ancestral GroE system, or perhaps reacquired one; these exceptions include the family Acholeplasmataceae (consisting of the genera Acholeplasma and Phytoplasma). The role of GroEL in these “exceptional” Mollicutes is a source of speculation, given how many non-canonical “moonlighting” roles have been ascribed to this protein. GroEL has been suggested to play a role in pathogenesis in plant and animal pathogenic Mollicutes, by binding to host cells and facilitating invasion. However, in one further layer of exception, the phytopathogenic taxon ‘Candidatus Phytoplasma pruni’ (ribosomal group 16SrIII), was reported to lack a GroE system. This study confirms the lack of a functional GroE system in 16SrIII by providing two new, high quality, non-fragmented genome assemblies, as well as a thorough survey of other 16SrIII genomes for genes encoding GroEL/GroES, including those that may not resemble phytoplasma groEL (ie. acquired by horizontal gene transfer, HGT). We discuss the implications of a clearly phytopathogenic, invasive group of Mollicutes that nevertheless lacks GroE, in light of the presumed role of GroEL for these microorganisms. We determined that three groups of genomes of 16SrIII contain short, non-functional groEL pseudogenes, while most of the reported genomes lack any semblance of a GroE system. Examination of the new assemblies allowed us to rule out HGT as a means of GroE acquisition.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".