MétaCan
Menu
Back to cohort
Record W2101407398 · doi:10.1101/gr.092072.109

A probabilistic approach for SNP discovery in high-throughput human resequencing data

2009· article· en· W2101407398 on OpenAlexafffund
Rose Hoberman, Joana Dias, Bing Ge, Eef Harmsen, Michael B. Mayhew, Dominique J. Verlaan, Tony Kwan, Ken Dewar, Mathieu Blanchette, Tomi Pastinen

Bibliographic record

VenueGenome Research · 2009
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicGenetic Associations and Epidemiology
Canadian institutionsMcGill University Health CentreMcGill University and Génome Québec Innovation CentreMcGill University
FundersGenome Canada
KeywordsInternational HapMap ProjectBiologyGeneticsComputational biology1000 Genomes ProjectFalse discovery rateSNP genotypingGenotypingDNA sequencingProbabilistic logicIdentification (biology)GenotypeSingle-nucleotide polymorphismComputer scienceArtificial intelligenceGene

Abstract

fetched live from OpenAlex

New high-throughput sequencing technologies are generating large amounts of sequence data, allowing the development of targeted large-scale resequencing studies. For these studies, accurate identification of polymorphic sites is crucial. Heterozygous sites are particularly difficult to identify, especially in regions of low coverage. We present a new strategy for identifying heterozygous sites in a single individual by using a machine learning approach that generates a heterozygosity score for each chromosomal position. Our approach also facilitates the identification of regions with unequal representation of two alleles and other poorly sequenced regions. The availability of confidence scores allows for a principled combination of sequencing results from multiple samples. We evaluate our method on a gold standard data genotype set from HapMap. We are able to classify sites in this data set as heterozygous or homozygous with 98.5% accuracy. In de novo data our probabilistic heterozygote detection ("ProbHD") is able to identify 93% of heterozygous sites at a <5% false call rate (FCR) as estimated based on independent genotyping results. In direct comparison of ProbHD with high-coverage 1000 Genomes sequencing available for a subset of our data, we observe >99.9% overall agreement for genotype calls and close to 90% agreement for heterozygote calls. Overall, our data indicate that high-throughput resequencing of human genomic regions requires careful attention to systematic biases in sample preparation as well as sequence contexts, and that their impact can be alleviated by machine learning-based sequence analyses allowing more accurate extraction of true DNA variants.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.004
metaresearch head score (Gemma)0.001
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.948
Threshold uncertainty score0.467

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0040.001
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0010.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.216
GPT teacher head0.434
Teacher spread0.217 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations25
Published2009
Admission routes2
Has abstractyes

Explore more

Same venueGenome ResearchSame topicGenetic Associations and EpidemiologyFrench-language works237,207