MétaCan
Menu
Back to cohort
Record W2582660596 · doi:10.1101/095463

Umap and Bismap: quantifying genome and methylome mappability

2016· preprint· en· W2582660596 on OpenAlexafffund
Mehran Karimzadeh, Carl Ernst, Anshul Kundaje, Michael M. Hoffman

Bibliographic record

VenuebioRxiv (Cold Spring Harbor Laboratory) · 2016
Typepreprint
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicEpigenetics and DNA Methylation
Canadian institutionsMcGill UniversityPrincess Margaret Cancer CentreUniversity of Toronto
FundersUniversity of TorontoPrincess Margaret Cancer FoundationNatural Sciences and Engineering Research Council of CanadaUniversity Health Network
KeywordsGenomeBisulfite sequencingComputational biologyGeneticsBiologyDNA sequencingGenomicsWhole genome sequencingBisulfiteHybrid genome assemblyHuman genomeReference genomeCancer genome sequencingDNA methylationGene

Abstract

fetched live from OpenAlex

Abstract Motivation Short-read sequencing enables assessment of genetic and biochemical traits of individual genomic regions, such as the location of genetic variation, protein binding, and chemical modifications. Every region in a genome assembly has a property called mappability which measures the extent to which it can be uniquely mapped by sequence reads. In regions of lower mappability, estimates of genomic and epigenomic characteristics from sequencing assays are less reliable. At best, sequencing assays will produce misleadingly low numbers of reads in these regions. At worst, these regions have increased susceptibility to spurious mapping from reads from other regions of the genome with sequencing errors or unexpected genetic variation. Bisulfite sequencing approaches used to identify DNA methylation exacerbate these problems by introducing large numbers of reads that map to multiple regions. While many tools consider mappability during the read mapping process, subsequent analysis often loses this information. Both to correct assumptions of uniformity in downstream analysis, and to identify regions where the analysis is less reliable, it is necessary to know the mappability of both ordinary and bisulfite-converted genomes. Results We introduce the Umap software for identifying uniquely mappable regions of any genome. Its Bismap extension identifies mappability of the bisulfite-converted genome. With a read length of 24 bp, 18.7% of the unmodified genome and 33.5% of the bisulfite-converted genome is not uniquely mappable. This complicates interpretation of functional genomics experiments using short-read sequencing, especially in regulatory regions. For example, 81% of human CpG islands overlap with regions that are not uniquely mappable. Similarly, in some ENCODE ChIP-seq datasets, up to 50% of peaks overlap with regions that are not uniquely mappable. We also explored differentially methylated regions from a case-control study and identified regions that were not uniquely mappable. In the widely used 450K methylation array, 4,230 probes are not uniquely mappable. Genome mappability is higher with longer sequencing reads, but most publicly available ChIP-seq and reduced representation bisulfite sequencing datasets have shorter reads. Therefore, uneven and low mappability remains a concern in a majority of existing data. Availability A Umap and Bismap track hub for human genome assemblies GRCh37/hg19 and GRCh38/hg38, and mouse assemblies GRCm37/mm9 and GRCm38/mm10 is available at http://bismap.hoffmanlab.org for use with the UCSC and Ensembl genome browsers. We have deposited in Zenodo the current version of our software ( https://doi.org/10.5281/zenodo.800648 ) and the mappability data used in this project ( https://doi.org/10.5281/zenodo.800645 ). In addition, the software ( https://bitbucket.org/hoffmanlab/umap ) is freely available under the GNU General Public License, version 3 (GPLv3). Contact michael.hoffman@utoronto.ca

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.003
metaresearch head score (Gemma)0.012
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.005
Threshold uncertainty score0.015

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0030.012
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0050.004
Science and technology studies0.0010.001
Scholarly communication0.0020.002
Open science0.0010.002
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0040.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.021
GPT teacher head0.249
Teacher spread0.228 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations48
Published2016
Admission routes2
Has abstractyes

Explore more

Same venuebioRxiv (Cold Spring Harbor Laboratory)Same topicEpigenetics and DNA MethylationFrench-language works237,207