Identification of Novel DNA Sequence Motifs that Modulate Transcription in T cells
Bibliographic record
Abstract
Abstract Considerable progress has been made towards associating transcription factor binding sites (TFBS) with cell-type-specific gene expression, however, the full repertoire of DNA sequence motifs that regulate transcription remains unknown. Improving our understanding of transcriptional regulation is especially important in T cells, given the enormous potential of genetically engineered T cells as an emerging class of therapeutics. Here, we report results from a comprehensive and unbiased survey investigating whether there are novel motifs enriched in regulatory regions of genes with the highest constitutive and selective expression across diverse T- cell subsets. Using computational and experimental methods, we identified 2,036 novel motifs and 629 previously curated TFBS that are enriched, both individually and in specific combinations, in the regulatory regions of genes exhibiting T-cell-specific gene expression. We then used the self-transcribing active regulatory region sequencing (STARR-seq) assay to evaluate all possible three- way combinations of a subset of 18 candidate motifs to test their ability to modulate transcription in immortalized lymphoblastic cell lines of T-cell origin (Jurkat E6) versus myeloid origin (K562). Our results revealed novel motifs that modulate gene transcription in T cells, with some exhibiting stronger regulatory effects than TFBS for TFs with established roles in T cells. The regulatory activity of these novel motifs was influenced by the motif’s orientation, position, and copy number. Overall, these results highlight our incomplete understanding of the relationship between sequence composition and T-cell gene regulation and indicate that previously annotated TFBS represent only a subset of motifs capable of modulating gene transcription in T cells.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".