Characterization of TOPHAT, a novel domesticated transposable element gene
Bibliographic record
Abstract
With the advent of high-throughput sequencing, it has been revealed that protein-coding genes constitute only a small portion of the genome. Comprising a large segment of the genome are transposons, or transposable elements (TEs), which have been classically regarded as “selfish” or “junk” DNA. Research on deciphering the non-coding regions has been a more recent area of focus, and TEs have been shown to be a contributing factor in genome evolution.The activity of TEs can be induced by environmental and population factors in various organisms. As part of the VEGI project, abiotic stress screens were performed on a curated set of T-DNA insertional mutagenesis lines to identify domesticated transposable element (DTE) genes with putative functions. In this screen, one uncharacterized DTE, which we have named TOPHAT, showed significant phenotypes in multiple abiotic tolerance screens (salt, nitrogen use efficiency, freezing). TOPHAT overexpression lines were created in a wild-type background to further characterize this DTE gene. In a parallel RNA-sequencing experiment, expression analyses confirmed the osmotic stress phenotype in the knockout line and identified a large group of pathogen defence genes that are expressed constitutively in the overexpression genotype. Proteins encoded by these genes play critical roles in pathogen recognition and the activation of defence responses. We propose that TOPHAT is a DTE that causes changes in the expression of a host of genes through a synergistic co-regulation with other genes in its gene family. Among those are genes which are associated with abiotic and biotic resistance and response. TOPHAT and the closely related SLEEPER genes may be acting collectively to repress a suite of stress-related genes to ensure and maintain transiency of expression. We hope our study characterizing this novel DTE can further our knowledge of the importance of DTEs in genome evolution.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".