MétaCan
Menu
Back to cohort
Record W2735049306 · doi:10.1080/15592294.2017.1335849

A MBD-seq protocol for large-scale methylome-wide studies with (very) low amounts of DNA

2017· article· en· W2735049306 on OpenAlexaff
Karolina A. Åberg, Robin F. Chan, Andrey A. Shabalin, Min Zhao, Gustavo Turecki, Nicklas Heine Staunstrup, Anna Starnawska, Ole Mors, Lin Xie, Edwin JCG van den Oord

Bibliographic record

VenueEpigenetics · 2017
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicEpigenetics and DNA Methylation
Canadian institutionsMcGill UniversityDouglas Mental Health University Institute
FundersNational Institute of Mental HealthLundbeckfondenNational Institutes of HealthVirginia Commonwealth University
KeywordsDNA methylationBisulfite sequencingBiologyComputational biologyCpG siteProtocol (science)DNA sequencingDNAGenomicsgenomic DNAGenomeMethylationComputer scienceGeneticsGene

Abstract

fetched live from OpenAlex

We recently showed that, after optimization, our methyl-CpG binding domain sequencing (MBD-seq) application approximates the methylome-wide coverage obtained with whole-genome bisulfite sequencing (WGB-seq), but at a cost that enables adequately powered large-scale association studies. A prior drawback of MBD-seq is the relatively large amount of genomic DNA (ideally >1 µg) required to obtain high-quality data. Biomaterials are typically expensive to collect, provide a finite amount of DNA, and may simply not yield sufficient starting material. The ability to use low amounts of DNA will increase the breadth and number of studies that can be conducted. Therefore, we further optimized the enrichment step. With this low starting material protocol, MBD-seq performed equally well, or better, than the protocol requiring ample starting material (>1 µg). Using only 15 ng of DNA as input, there is minimal loss in data quality, achieving 93% of the coverage of WGB-seq (with standard amounts of input DNA) at similar false/positive rates. Furthermore, across a large number of genomic features, the MBD-seq methylation profiles closely tracked those observed for WGB-seq with even slightly larger effect sizes. This suggests that MBD-seq provides similar information about the methylome and classifies methylation status somewhat more accurately. Performance decreases with <15 ng DNA as starting material but, even with as little as 5 ng, MBD-seq still achieves 90% of the coverage of WGB-seq with comparable genome-wide methylation profiles. Thus, the proposed protocol is an attractive option for adequately powered and cost-effective methylome-wide investigations using (very) low amounts of DNA.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.004
metaresearch head score (Gemma)0.005
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: Bench or experimental
GenreCandidate signal: Protocol · Consensus signal: none
Teacher disagreement score0.013
Threshold uncertainty score0.043

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0040.005
Meta-epidemiology (narrow)0.0020.002
Meta-epidemiology (broad)0.0020.002
Bibliometrics0.0020.002
Science and technology studies0.0020.001
Scholarly communication0.0020.001
Open science0.0020.002
Research integrity0.0020.006
Insufficient payload (model declined to judge)0.0130.015

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.034
GPT teacher head0.355
Teacher spread0.321 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreProtocol

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations63
Published2017
Admission routes1
Has abstractyes

Explore more

Same venueEpigeneticsSame topicEpigenetics and DNA MethylationFrench-language works237,207