Variably methylated regions in the newborn epigenome: environmental, genetic and combined influences
Bibliographic record
Abstract
Abstract Background Epigenetic processes, including DNA methylation (DNAm), are among the mechanisms allowing integration of genetic and environmental factors to shape cellular function. While many studies have investigated either environmental or genetic contributions to DNAm, few have assessed their integrated effects. We examined the relative contributions of prenatal environmental factors and genotype on DNA methylation in neonatal blood at variably methylated regions (VMRs), defined as consecutive CpGs showing the highest variability of DNAm in 4 independent cohorts (PREDO, DCHS, UCI, MoBa, N=2,934). Results We used Akaike’s information criterion to test which factors best explained variability of methylation in the cohort-specific VMRs: several prenatal environmental factors (E) including maternal demographic, psychosocial and metabolism related phenotypes, genotypes in cis (G), or their additive (G+E) or interaction (GxE) effects. G+E and GxE models consistently best explained variability in DNAm of VMRs across the cohorts, with G explaining the remaining sites best. VMRs best explained by G, GxE or G+E, as well as their associated functional genetic variants (predicted using deep learning algorithms), were located in distinct genomic regions, with different enrichments for transcription and enhancer marks. Genetic variants of not only G and G+E models, but also of variants in GxE models were significantly enriched in genome wide association studies (GWAS) for complex disorders. Conclusion Genetic and environmental factors in combination best explain DNAm at VMRs. The CpGs best explained by G, G+E or GxE are functionally distinct. The enrichment of GxE variants in GWAS for complex disorders supports their importance for disease risk.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".