SnoRNA copy regulation affects family size, genomic location and family abundance levels
Bibliographic record
Abstract
BACKGROUND: Small nucleolar RNAs (snoRNAs) are an abundant class of noncoding RNAs present in all eukaryotes and best known for their involvement in ribosome biogenesis. In mammalian genomes, many snoRNAs exist in multiple copies, resulting from recombination and retrotransposition from an ancestral snoRNA. To gain insight into snoRNA copy regulation, we used Rfam classification and normal human tissue expression datasets generated using low structure bias RNA-seq to characterize snoRNA families. RESULTS: We found that although box H/ACA families are on average larger than box C/D families, the number of expressed members is similar for both types. Family members can cover a wide range of average abundance values, but importantly, expression variability of individual members of a family is preferred over the total variability of the family, especially for box H/ACA snoRNAs, suggesting that while members are likely differentially regulated, mechanisms exist to ensure uniformity of the total family abundance across tissues. Box C/D snoRNA family members are mostly embedded in the same host gene while box H/ACA family members tend to be encoded in more than one different host, supporting a model in which box C/D snoRNA duplication occurred mostly by cis recombination while box H/ACA snoRNA families have gained copy members through retrotransposition. And unexpectedly, snoRNAs encoded in the same host gene can be regulated independently, as some snoRNAs within the same family vary in abundance in a divergent way between tissues. CONCLUSIONS: SnoRNA copy regulation affects family sizes, genomic location of the members and controls simultaneously member and total family abundance to respond to the needs of individual tissues.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".