Epigenotype–genotype–phenotype correlations in <i>SETD1A</i> and <i>SETD2</i> chromatin disorders
Bibliographic record
Abstract
Germline pathogenic variants in two genes encoding the lysine-specific histone methyltransferase genes SETD1A and SETD2 are associated with neurodevelopmental disorders (NDDs) characterized by developmental delay and congenital anomalies. The SETD1A and SETD2 gene products play a critical role in chromatin-mediated regulation of gene expression. Specific methylation episignatures have been detected for a range of chromatin gene-related NDDs and have impacted clinical practice by improving the interpretation of variant pathogenicity. To investigate if SETD1A and/or SETD2-related NDDs are associated with a detectable episignature, we undertook targeted genome-wide methylation profiling of > 2 M CpGs using a next-generation sequencing-based assay. A comparison of methylation profiles in patients with SETD1A variants (n = 6) did not reveal evidence of a strong methylation episignature. A review of the clinical and genetic features of the SETD2 patient group revealed that, as reported previously, there were phenotypic differences between patients with truncating mutations (n = 4, Luscan-Lumish syndrome; MIM:616831) and those with missense codon 1740 variants [p.Arg1740Trp (n = 4) and p.Arg1740Gln (n = 2)]. Both SETD2 subgroups demonstrated a methylation episignature, which was characterized by hypomethylation and hypermethylation events, respectively. Within the codon 1740 subgroup, both the methylation changes and clinical phenotype were more severe in those with p.Arg1740Trp variants. We also noted that two of 10 cases with a SETD2-NDD had developed a neoplasm. These findings reveal novel epigenotype-genotype-phenotype correlations in SETD2-NDDs and predict a gain-of-function mechanism for SETD2 codon 1740 pathogenic variants.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".