CpGPT: a Foundation Model for DNA Methylation
Bibliographic record
Abstract
DNA methylation is a type of epigenetic modification that plays a significant role in development, aging, and disease. Despite extensive research, how genome-wide DNA methylation patterns collectively encode and influence complex phenotypes such as aging and disease remains difficult to characterize with conventional approaches. Foundation models are a class of machine learning model that leverage vast quantities of data to make sense of complex data types, such as genome sequences or single-cell transcriptomes. Here, we present the Cytosine-phosphate-Guanine Pretrained Transformer (CpGPT), a novel foundation model pretrained on CpGCorpus, a novel database with more than 2,000 DNA methylation datasets encompassing over 150,000 samples from diverse conditions. CpGPT leverages an improved transformer architecture to learn comprehensive representations of methylation patterns, allowing it to impute and reconstruct genome-wide methylation profiles from limited input data. By capturing sequence, positional, and epigenetic contexts, CpGPT outperforms specialized models when finetuned for aging-related tasks, including the state-of-the-art GrimAge2 and PCGrimAge for mortality and morbidity estimation. The model is highly adaptable and can impute beta values across different methylation platforms, tissue types, mammalian species, and even single-cell data. As a foundation model, CpGPT can be leveraged as a new tool for biological discovery in the field of epigenetics. The open-source code and model can be found at \url{http://github.com/lucascamillomd/CpGPT}.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".