MétaCan
Menu
Back to cohort
Record W4390110942 · doi:10.1101/2023.12.20.23300294

Rare disease gene association discovery from burden analysis of the 100,000 Genomes Project data

2023· preprint· en· W4390110942 on OpenAlexaff
Valentina Cipriani, Letizia Vestito, Emma Magavern, Julius O.B. Jacobsen, Gavin Arno, Elijah R. Behr, Katherine A. Benson, Marta Bértoli, Detlef Böckenhauer, Michael R. Bowl, Kate Burley, Li F. Chan, Patrick F. Chinnery, Peter J. Conlon, Marcos Costa, Alice E. Davidson, Sally J. Dawson, Elhussein A. Elhassan, Sarah E. Flanagan, Marta Futema, Daniel P. Gale, Sonia García-Ruiz, M. Cecilia Gonzalez Corcia, Helen Griffin, Sophie Hambleton, Amy R. Hicks, Henry Houlden, Richard S. Houlston, Sarah Howles, Robert Kleta, Iris Lekkerkerker, Siying Lin, Petra Lišková, Hannah M. Mitchison, Heba Morsy, Andrew Mumford, Ruxandra Neatu, Edel A. O’Toole, Albert Ong, Alistair T. Pagnamenta, Shamima Rahman, Neil Rajan, Peter N Robinson, Mina Ryten, Omid Sadeghi‐Alavijeh, John A. Sayer, Claire L. Shovlin, Jenny C. Taylor, Omri Teltsh, Ian Tomlinson, Arianna Tucci, Clare Turnbull, Albertien M. van Eerde, James S. Ware, Laura M Watts, Andrew R. Webster, Sarah K. Westbury, Sean Zheng, Mark J. Caulfield, Damian Smedley

Bibliographic record

VenuemedRxiv · 2023
Typepreprint
Languageen
FieldMedicine
TopicCardiomyopathy and Myosin Studies
Canadian institutionsUniversité de MontréalMcGill UniversityCentre Hospitalier Universitaire Sainte-Justine
FundersMedical Research CouncilUniversity College London Hospitals NHS Foundation TrustGrantová Agentura České RepublikyMichael J. Fox Foundation for Parkinson's ResearchDepartment of Health and Social CareNational Institutes of HealthCancer Research UKWellcome TrustNational Institute of Child Health and Human DevelopmentNational Institute for Health and Care Research
KeywordsDiseaseGenetic associationGenome-wide association studyMedicineBioinformaticsLocus heterogeneityIn silicoGeneticsBiologyPhenotypeGenetic heterogeneityGenePathologySingle-nucleotide polymorphismGenotype

Abstract

fetched live from OpenAlex

Abstract To discover rare disease-gene associations, we developed a gene burden analytical framework and applied it to rare, protein-coding variants from whole genome sequencing of 35,008 cases with rare diseases and their family members recruited to the 100,000 Genomes Project (100KGP). Following in silico triaging of the results, 88 novel associations were identified including 38 with existing experimental evidence. We have published the confirmation of one of these associations, hereditary ataxia with UCHL1 , and independent confirmatory evidence has recently been published for four more. We highlight a further seven compelling associations: hypertrophic cardiomyopathy with DYSF and SLC4A3 where both genes show high/specific heart expression and existing associations to skeletal dystrophies or short QT syndrome respectively; monogenic diabetes with UNC13A with a known role in the regulation of β cells and a mouse model with impaired glucose tolerance; epilepsy with KCNQ1 where a mouse model shows seizures and the existing long QT syndrome association may be linked; early onset Parkinson’s disease with RYR1 with existing links to tremor pathophysiology and a mouse model with neurological phenotypes; anterior segment ocular abnormalities associated with POMK showing expression in corneal cells and with a zebrafish model with developmental ocular abnormalities; and cystic kidney disease with COL4A3 showing high renal expression and prior evidence for a digenic or modifying role in renal disease. Confirmation of all 88 associations would lead to potential diagnoses in 456 molecularly undiagnosed cases within the 100KGP, as well as other rare disease patients worldwide, highlighting the clinical impact of a large-scale statistical approach to rare disease gene discovery.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.007
metaresearch head score (Gemma)0.022
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.007
Threshold uncertainty score0.039

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0070.022
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.002
Bibliometrics0.0060.005
Science and technology studies0.0010.000
Scholarly communication0.0020.001
Open science0.0010.003
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0030.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.075
GPT teacher head0.323
Teacher spread0.248 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2023
Admission routes1
Has abstractyes

Explore more

Same venuemedRxivSame topicCardiomyopathy and Myosin StudiesFrench-language works237,207