<i>De novo</i> genome sequence assembly of the RNAi-tractable endosymbiosis model system <i>Paramecium bursaria</i> 186b reveals factors shaping intron repertoire
Bibliographic record
Abstract
How two species engage in stable endosymbiosis is a biological quandary. The study of facultative endosymbiotic interactions has emerged as a useful approach to understand how endosymbiotic functions can arise. The ciliate protist Paramecium bursaria hosts green algae of the order Chlorellales in a facultative photo-endosymbiosis. We have recently reported RNAi as a tool for understanding gene function in Paramecium bursaria 186b, CCAP strain 1660/18 [1]. To complement this work, here we report a highly complete host genome and trans criptome sequence dataset, using both Illumina and PacBio sequencing methods to aid genome analysis and to enable the design of RNAi experiments. Our analyses demonstrate Paramecium bursaria , like other ciliates such as diverse species of Paramecia , possess numerous tiny introns. These data, combined with the alternative genetic code common to ciliates, makes gene identification and annotation challenging. To explore intron evolutionary dynamics further we show that alternative splicing leading to intron retention occurs at a higher frequency among the smaller number of longer introns, identifying a source of selection against longer introns. These data will aid the investigation of genome evolution in the Paramecia and provide additional source data for the exploration of endosymbiotic functions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".