Long Terminal Repeats Are Used as Alternative Promoters for the Endothelin B Receptor and Apolipoprotein C-I Genes in Humans
Bibliographic record
Abstract
To examine the potential regulatory involvement of retroelements in the human genome, we screened the transcribed sequences of GenBank and expressed sequence tag data bases with long terminal repeat (LTR) elements derived from different human endogenous retroviruses. These screenings detected human transcripts containing LTRs belonging to the human endogenous retrovirus-E family fused to the apolipoprotein CI (apoC-I) and the endothelin B receptor (EBR) genes. However, both genes are known to have non-LTR (native) promoters. Initial reverse transcription-polymerase chain reaction experiments confirmed and authenticated the presence of transcripts from both the native and LTR promoters. Using a 5'-rapid amplification of cDNA ends protocol, we showed that the alternative transcripts of apoC-I and EBR are initiated and promoted by the LTRs. The LTR-apoC-I fusion and native apoC-I transcripts are present in many of the tissues tested. As expected, we found apoC-I preferentially expressed in liver, where about 15% of the transcripts are derived from the LTR promoter. Transient transfections suggest that the expression is not dependent on the LTR itself, but the presence of the LTR increases activity of the apoC-I promoter from both humans and baboons. The native EBR-driven transcripts were also detected in many tissues, whereas the LTR-driven transcripts appear limited to placenta. In contrast to the LTR of apoC-I, the EBR LTR promotes a significant proportion of the total EBR transcripts, and transient transfection results indicate that the LTR acts as a strong promoter and enhancer in a placental cell line. This investigation reports two examples where LTR sequences contribute to increased transcription of human genes and illustrates the impact of mobile elements on gene and genome evolution.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".