MétaCan
Menu
← Back to cohort
Record W3009856650 · doi:10.1101/2020.01.12.20016691

metPropagate: network-guided propagation of metabolomic information for prioritization of neurometabolic disease genes

2020· preprint· en· W3009856650 on OpenAlexafffund
Emma Graham, Phillip A. Richmond, Maja Tarailo‐Graovac, Udo F. H. Engelke, Leo A. J. Kluijtmans, Karlien L. M. Coene, Ron A. Wevers, Wyeth W. Wasserman, Clara D.M. van Karnebeek, Sara Mostafavi

Bibliographic record

VenuemedRxiv · 2020
Typepreprint
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicBioinformatics and Genomic Networks
Canadian institutionsAlberta Children's HospitalUniversity of CalgaryBC Children's HospitalUniversity of British Columbia
FundersNatural Sciences and Engineering Research Council of CanadaCanadian Institutes of Health ResearchBC Children's HospitalMichael Smith Health Research BCStichting MetakidsChildren's Hospital Foundation
KeywordsCandidate geneComputational biologyGeneExomeExome sequencingPrioritizationPhenotypeBioinformaticsGeneticsBiologyMedicine

Abstract

fetched live from OpenAlex

ABSTRACT Many inborn errors of metabolism (IEMs) are amenable to treatment, therefore early diagnosis before irreversible damage occurs is imperative. Despite recent advances, the genetic basis of many metabolic phenotypes remains unknown. For discovery purposes, Whole Exome Sequencing (WES) variant prioritization coupled with phenotype-guided clinical and bioinformatics expertise is currently the primary method used to identify novel disease-causing variants; however, it can be challenging to identify the causal candidate gene given the large number of plausible variants. Using untargeted metabolomics (UM) to prioritize metabolically relevant candidate genes is a promising approach to diagnosing known or novel IEMs in a single patient. Here, we present a network-based bioinformatics approach, metPropagate, that uses UM data from a single patient and a group of controls to prioritize candidate genes. We validate metProp on 107 patients with diagnosed IEMs and 11 patients with novel IEMs diagnosed through the TIDE gene discovery project at BC Children’s Hospital. The metPropagate method ranks candidate genes by considering the network of interactions between them. This is done by using a graph smoothing algorithm called label propagation to quantify the metabolic disruption in genes’ local neighbourhood. metPropagate was able to prioritize the causative gene in the top 20th percentile of candidate genes for 91% of patients with known IEM disorders. For novel IEMs, metPropagate placed the causative gene in the top 20 th percentile in 9/11 patients. Using metPropagate, the causative gene was ranked higher than Exomiser’s phenotype-based ranking in 6/11 patients. The results of this study indicate that for diagnostic and gene discovery purposes, network-based analysis of metabolomics data can lend support to WES gene-discovery methods by providing an additional mode of evidence to help identify causal genes.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.004
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: Simulation or modeling
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.008
Threshold uncertainty score0.017

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.004
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0030.001
Science and technology studies0.0010.000
Scholarly communication0.0010.001
Open science0.0010.001
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0040.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.019
GPT teacher head0.253
Teacher spread0.234 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2020
Admission routes2
Has abstractyes

Explore more

Same venuemedRxiv→Same topicBioinformatics and Genomic Networks→French-language works237,207→