Simultaneous epigenomic profiling and regulatory activity measurement using e2MPRA
Bibliographic record
Abstract
regulatory elements (CREs) have a major effect on phenotypes including disease. They are identified in a genome-wide manner by analyzing the binding of transcription factors (TFs), various co-factors and histone modifications in DNA using assays such as ChIP-seq, Cut&Tag and ATAC-seq. However, these assays are descriptive and require high-throughput technologies, such as massively parallel reporter assays (MPRAs), to test the functional activity and variant effect on these sequences. Currently, technologies that can simultaneously analyze both the regulatory function of a specific sequence and the TFs, cofactors and epigenomic modifications that determine it do not exist. Here, we developed enrichment followed by epigenomic profiling MPRA (e2MPRA), a novel technology that utilizes lentivirus-based MPRA to enrich for the integration of specific CREs into the genome followed by Cut&Tag or ATAC-seq targeted specifically for these sequences. This method allows to simultaneously analyze in a high-throughput manner regulatory activity, protein binding and epigenetic modification of thousands of candidate CREs and their variants. We demonstrate that e2MPRA can be used to dissect the epigenetic functions of TF motifs arranged in synthetic enhancers, as well as to analyze the effect of enhancer sequence variants on epigenetic modifications. In summary, this technology will increase our understanding of the regulatory code, its effect on the epigenome and how its alteration can lead to a variety of phenotypes including human disease.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".