Conservation and architecture of housekeeping genes in the model marine diatom <i>Thalassiosira pseudonana</i>
Bibliographic record
Abstract
Housekeeping genes (HKGs) are constitutively expressed with low variation across tissues/conditions. They are thought to be highly conserved and fundamental to cellular maintenance, with distinctive genomic features. Here, we identify 1505 HKGs in the unicellular marine diatom Thalassiosira pseudonana based on an RNA-seq analysis of 232 samples taken under 12 experimental conditions over 0-72 h. We identify promising internal reference genes (IRGs) for T. pseudonana from the most stably expressed HKGs. A comparative analysis indicates < 18% of HKGs in T. pseudonana have orthologs in other eukaryotes, including other diatom species. Contrary to work on human tissues, T. pseudonana HKGs are longer than non-HKGs, due to elongated introns. More ancient HKGs tend to be shorter than more recent HKGs, and expression levels of HKGs decrease more rapidly with gene length relative to non-HKGs. Our results indicate that HKGs are highly variable across the tree of life and thus unlikely to be universally fundamental for cellular maintenance. We hypothesize that the distinct genomic features of HKGs of T. pseudonana may be a consequence of selection pressures associated with high expression and low variance across conditions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".