Unveiling the functional nature of retrogenes in dinoflagellates
Bibliographic record
Abstract
Retroposition is a gene duplication mechanism that uses RNA molecules as intermediaries to generate new gene copies. Dinoflagellates are proposed as an ideal model for exploring this process due to the tagging of retrogenes with DNA-encoded remnants of the dinoflagellate-specific splice-leader motif at their 5' end. We conducted a comprehensive search for retrogenes in dinoflagellate transcriptomes to uncover their functional nature and the processes underlying their redundancy. We obtained a high-confidence set of hypothetical functional retrogenes widespread through the dinoflagellate lineage. Through annotations and gene ontology enrichment analysis, we found that the functional diversity of retrogenes reflects the most prevalent and active processes during stress periods, particularly those involving post-translational modifications and cell signalling pathways. Additionally, the significant presence of retrogenes linked to specific biological processes involved in symbiosis and toxin production underscores the role of retrogenes in adaptation. The expression profile and codon composition similar to protein-coding genes confirm the operational status of retrogenes and strengthen the idea that retrogenes recapitulate parental gene expression and function. This study provides new evidence supporting widespread gene retroposition across dinoflagellates and highlights the functional link of retrogenes with the core activity of the cell.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".