MétaCan
Menu
Back to cohort
Record W2216458865 · doi:10.1038/ncomms9111

Improved imputation of low-frequency and rare variants using the UK10K haplotype reference panel

2015· article· en· W2216458865 on OpenAlexafffund
Jie Huang, Bryan Howie, Shane McCarthy, Yasin Memari, Klaudia Walter, Josine L. Min, Petr Danecek, Giovanni Malerba, Elisabetta Trabetti, Hou‐Feng Zheng, Saeed Al Turki, Antoinette Amuzu, Carl A. Anderson, Richard Anney, Dinu Antony, María Soler Artigas, Muhammad Ayub, Senduran Bala, Jeffrey C. Barrett, Inês Barroso, Phil Beales, Marianne Benn, Jamie Bentham, Shoumo Bhattacharya, Ewan Birney, Douglas Blackwood, Martin Bobrow, Elena G. Bochukova, Patrick Bolton, Rebecca Bounds, Chris Boustred, Gerome Breen, Mattia Calissano, Keren Carss, Juan P. Casas, John C. Chambers, Ruth Charlton, Krishna Chatterjee, Lu Chen, Antonio Ciampi, Sebahattin Çırak, Peter Clapham, Gail Clement, Guy Coates, Massimiliano Cocca, David Collier, Catherine Cosgrove, Tony Cox, Nick Craddock, Lucy Crooks, Sarah Curran, David Curtis, Allan Daly, Ian N.M. Day, Aaron Day-Williams, George Dedoussis, Thomas A. Down, Yuanping Du, Cornelia M. van Duijn, Ian Dunham, Sarah Edkins, Rosemary Ekong, Peter Ellis, David M. Evans, I. Sadaf Farooqi, David Fitzpatrick, Paul Flicek, James Floyd, A. Reghan Foley, Christopher S. Franklin, Marta Futema, Louise Gallagher, Paolo Gasparini, Tom R. Gaunt, Matthias Geihs, Daniel H. Geschwind, Celia M.T. Greenwood, Heather Griffin, Detelina Grozeva, Xiaosen Guo, Xueqin Guo, Hugh Gurling, Deborah Hart, Audrey E. Hendricks, Peter Holmans, Tim Hubbard, Steve E. Humphries, Matthew E. Hurles, Pirro G. Hysi, Valentina Iotchkova, Aaron Isaacs, David K. Jackson, Yalda Jamshidi, Jon Johnson, Christopher Joyce, Konrad J. Karczewski, Jane Kaye, Thomas Keane, John P. Kemp, Karen Kennedy, Alastair Kent, Julia M. Keogh, Farrah Khawaja, Marcus E. Kleber, Margriet van Kogelenberg, Anja Kolb-Kokocinski, Jaspal S. Kooner, Geneviève Lachance, Claudia Langenberg, Cordelia Langford, Daniel J. Lawson, Irene Lee, Monkol Lek, Rui Li, Jieqin Liang, Hong Lin, Ryan Liu, Jouko Lönnqvist, Luís R. Lopes, Margarida Lopes, Jian’an Luan, Daniel G. MacArthur, Massimo Mangino, Gaëlle Marenne, Winfried März, John Maslen, Angela Matchan, Iain Mathieson, Peter McGuffin, Andrew M. McIntosh, Andrew G. McKechanie, Andrew McQuillin, Sarah Metrustry, Nicola Migone, Hannah M. Mitchison, Alireza Moayyeri, James Morris, Richard Morris, Dawn Muddyman, Francesco Muntoni, Børge G. Nordestgaard, Kate Northstone, Michael O‘Donovan, Stephen O’Rahilly, Alexandros Onoufriadis, Karim Oualkacha, Michael J. Owen, Aarno Palotie, Kalliope Panoutsopoulou, Victoria Parker, Jeremy Parr, Lavinia Paternoster, Tiina Paunio, Felicity Payne, Stewart J. Payne, John R. B. Perry, Olli Pietiläinen, Vincent Plagnol, Rebecca C. Pollitt, Sue Povey, Michael A. Quail, Lydia Quaye, Lucy Raymond, Karola Rehnström, Cheryl K. Ridout, Susan M. Ring, Graham R. S. Ritchie, Nicola Roberts, Rachel L. Robinson, David B. Savage, Peter Scambler, Stephan Schiffels, Miriam Schmidts, Nadia Schoenmakers, Richard H. Scott, Robert A. Scott, Robert K. Semple, Eva Serra, Sally I. Sharp, Adam Shaw, Hashem A. Shihab, So–Youn Shin, David Skuse, Kerrin S. Small, Carol Smee, George Davey Smith, Lorraine Southam, Olivera Spasić-Bošković, Timothy D. Spector, David St Clair, Beaté St Pourcain, Jim Stalker, Elizabeth Stevens, Jianping Sun, Gabriela Surdulescu, Jaana Suvisaari, Petros Syrris, Ioanna Tachmazidou, Rohan Taylor, Jing Tian, Martin D. Tobin, Daniela Toniolo, Michela Traglia, Anne Tybjærg‐Hansen, Ana M. Valdes, Anthony M. Vandersteen, Anette Varbo, Parthiban Vijayarangakannan, Peter M. Visscher, Louise V. Wain, James Walters, Guangbiao Wang, Jun Wang, Yu Wang, Kirsten Ward, Eleanor Wheeler, Peter H. Whincup, Tamieka Whyte, Hywel Williams, Kathleen A. Williamson, Crispian Wilson, Scott G. Wilson, Kim Wong, Changjiang Xu, Jian Yang, Gianluigi Zaza, Eleftheria Zeggini, Feng Zhang, Pingbo Zhang, Weihua Zhang, Giovanni Gambaro, J. Brent Richards, Richard Durbin, Nicholas J. Timpson, Jonathan Marchini, Nicole Soranzo

Bibliographic record

VenueNature Communications · 2015
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicGenetic Associations and Epidemiology
Canadian institutionsUniversité du Québec à MontréalQueen's UniversityMcGill UniversityJewish General Hospital
FundersEconomic and Social Research CouncilMedical Research CouncilJewish General HospitalEuropean CommissionUniversity College LondonCanadian Institutes of Health ResearchNational Institute for Health and Care ResearchQuébec Consortium for Drug DiscoveryBritish Heart FoundationWellcome TrustLondon School of Hygiene and Tropical Medicine
KeywordsHaplotypeImputation (statistics)Allele frequencyGeneticsComputational biologyComputer scienceBiologyGenotypeGeneMissing dataMachine learning

Abstract

fetched live from OpenAlex

Imputing genotypes from reference panels created by whole-genome sequencing (WGS) provides a cost-effective strategy for augmenting the single-nucleotide polymorphism (SNP) content of genome-wide arrays. The UK10K Cohorts project has generated a data set of 3,781 whole genomes sequenced at low depth (average 7x), aiming to exhaustively characterize genetic variation down to 0.1% minor allele frequency in the British population. Here we demonstrate the value of this resource for improving imputation accuracy at rare and low-frequency variants in both a UK and an Italian population. We show that large increases in imputation accuracy can be achieved by re-phasing WGS reference panels after initial genotype calling. We also present a method for combining WGS panels to improve variant coverage and downstream imputation accuracy, which we illustrate by integrating 7,562 WGS haplotypes from the UK10K project with 2,184 haplotypes from the 1000 Genomes Project. Finally, we introduce a novel approximation that maintains speed without sacrificing imputation accuracy for rare variants.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.001
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.865
Threshold uncertainty score0.228

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.001
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.065
GPT teacher head0.333
Teacher spread0.268 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations385
Published2015
Admission routes2
Has abstractyes

Explore more

Same venueNature CommunicationsSame topicGenetic Associations and EpidemiologyFrench-language works237,207