Interaction of bZIP and bHLH Transcription Factors with the G-box
Bibliographic record
Abstract
Transcription factors are proteins that regulate transcription of genes by binding to specific DNA sequences proximal to the gene. The specificity and affinity of protein-DNA recognition is critical for proper gene regulation. This thesis explores the mechanisms of binding to the sequence 5’CACGTG, a common recognition sequence both in plants where it is known as the G-box and in mammalian cells where it is termed the E-box. This sequence is of clinical interest because it is the target of the transcription factor Myc, an oncogene linked to many cancers. A number of alpha-helical proteins with different dimerization elements, from the basic region-leucine zipper (bZIP), basic region helix-loop-helix leucine zipper (bHLHZ) and basic region helix-loop-helix-PAS (bHLH-PAS) protein families, are capable of binding to this sequence. The basic regions of all these protein families contain residues that contact DNA and determine DNA sequence specificity while the other subdomains are responsible for dimerization specificity. First, the influence of protein-DNA contacts on sequence specificity of the plant bZIP protein EmBP-1 was probed by point mutations in the basic region. Residues that contact the DNA outside the core G-box sequence and residues that contact the phosphate backbone were found to be important for sequence specificity. Second, the impact of the dimerization subdomains of bHLHZ protein Max, the required heterodimerization partner of the Myc protein, and bHLH-PAS protein Arnt was probed by mutation, deletion and inter-family subdomain swapping studies. All studied protein families are intrinsically disordered, forming structure upon dimerization and DNA binding. The dimerization domains were found to indirectly influence DNA binding by affecting folding, dimerization ability or proper orientation of the basic regions relative to DNA. Lastly, a new strategy for selection of G-box binding proteins in the Yeast One-hybrid system is explored. Together, these studies broaden our understanding of the structure-function relationship of the DNA-binding activities of these closely related families of transcription factors. The creation and characterization of mutants with altered specificity, affinity and dimerization specificity may also be useful for biotechnology applications.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".