MétaCan
Menu
Back to cohort
Record W2256016639 · doi:10.1038/nature19057

Analysis of protein-coding genetic variation in 60,706 humans

2016· article· en· W2256016639 on OpenAlexafffund
Monkol Lek, Konrad J. Karczewski, Eric Vallabh Minikel, Kaitlin E. Samocha, Eric Banks, Timothy R. Fennell, Anne O’Donnell‐Luria, James S. Ware, Andrew Hill, Beryl B. Cummings, Taru Tukiainen, Daniel P. Birnbaum, Jack A. Kosmicki, Laramie E. Duncan, Karol Estrada, Fengmei Zhao, James Zou, Emma Pierce‐Hoffman, Joanne Berghout, D.N. Cooper, Nicole Deflaux, Mark A. DePristo, Ron Do, Jason Flannick, Menachem Fromer, Laura D. Gauthier, Jackie Goldstein, Namrata Gupta, Daniel P. Howrigan, Adam Kieżun, Mitja Kurki, Ami Levy Moonshine, Pradeep Natarajan, Lorena Orozco, Gina M. Peloso, Ryan Poplin, Manuel A. Rivas, Valentín Ruano-Rubio, Samuel A. Rose, Douglas M. Ruderfer, Khalid Shakir, Peter D. Stenson, Christine Stevens, Brett Thomas, Grace Tiao, Maria T. Tusie-Luna, Ben Weisburd, Hong‐Hee Won, Dongmei Yu, David Altshuler, Diego Ardissino, Michael Boehnke, John Danesh, Stacey Donnelly, Roberto Elosúa, José C. Florez, Stacey Gabriel, Gad Getz, Stephen J. Glatt, Christina M. Hultman, Sekar Kathiresan, Markku Laakso, Steven A. McCarroll, Mark I. McCarthy, Dermot McGovern, Ruth McPherson, Benjamin M. Neale, Aarno Palotie, Shaun Purcell, Danish Saleheen, Jeremiah M. Scharf, Pamela Sklar, Patrick F. Sullivan, Jaakko Tuomilehto, Ming T. Tsuang, Hugh Watkins, James G. Wilson, Mark J. Daly, Daniel G. MacArthur

Bibliographic record

VenueNature · 2016
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicGenomics and Rare Diseases
Canadian institutionsUniversity of Ottawa
FundersNational Institute of Diabetes and Digestive and Kidney DiseasesWellcome TrustNational Human Genome Research InstituteNational Institute of Mental HealthNational Heart, Lung, and Blood InstituteNational Institute on AgingMedical Research CouncilEuropean CommissionCanadian Institutes of Health ResearchNational Institute of Neurological Disorders and StrokeNational Institute of General Medical SciencesNational Institute for Health and Care ResearchU.S. Public Health ServiceBritish Heart Foundation
KeywordsVariation (astronomy)Genetic variationCoding (social sciences)Evolutionary biologyComputational biologyBiologyGeneticsStatisticsGeneMathematicsPhysics

Abstract

fetched live from OpenAlex

Large-scale reference data sets of human genetic variation are critical for the medical and functional interpretation of DNA sequence changes. Here we describe the aggregation and analysis of high-quality exome (protein-coding region) DNA sequence data for 60,706 individuals of diverse ancestries generated as part of the Exome Aggregation Consortium (ExAC). This catalogue of human genetic diversity contains an average of one variant every eight bases of the exome, and provides direct evidence for the presence of widespread mutational recurrence. We have used this catalogue to calculate objective metrics of pathogenicity for sequence variants, and to identify genes subject to strong selection against various classes of mutation; identifying 3,230 genes with near-complete depletion of predicted protein-truncating variants, with 72% of these genes having no currently established human disease phenotype. Finally, we demonstrate that these data can be used for the efficient filtering of candidate disease-causing variants, and for the discovery of human ‘knockout’ variants in protein-coding genes. Exome sequencing data from 60,706 people of diverse geographic ancestry is presented, providing insight into genetic variation across populations, and illuminating the relationship between DNA variants and human disease. As part of the Exome Aggregation Consortium (ExAC) project, Daniel MacArthur and colleagues report on the generation and analysis of high-quality exome sequencing data from 60,706 individuals of diverse ancestry. This provides the most comprehensive catalogue of human protein-coding genetic variation to date, yielding unprecedented resolution for the analysis of very rare variants across multiple human populations. The catalogue is freely accessible and provides a critical reference panel for the clinical interpretation of genetic variants and the discovery of disease-related genes.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.002
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.002
Threshold uncertainty score0.005

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.002
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.002
Science and technology studies0.0000.000
Scholarly communication0.0010.000
Open science0.0000.001
Research integrity0.0010.000
Insufficient payload (model declined to judge)0.0010.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.004
GPT teacher head0.230
Teacher spread0.227 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations10,311
Published2016
Admission routes2
Has abstractyes

Explore more

Same venueNatureSame topicGenomics and Rare DiseasesFrench-language works237,207