Increased Diagnostic Yield by Reanalysis of Whole Exome Sequencing Data in Mitochondrial Disease
Bibliographic record
Abstract
Background: The genetic diagnosis of mitochondrial disorders is complicated by its genetic and phenotypic complexity. Next generation sequencing techniques have much improved the diagnostic yield for these conditions. A cohort of individuals with multiple respiratory chain deficiencies, reported in the literature 10 years ago, had a diagnostic rate of 60% by whole exome sequencing (WES) but 40% remained undiagnosed. Objective: We aimed to identify a genetic diagnosis by reanalysis of the WES data for the undiagnosed arm of this 10-year-old cohort of patients with suspected mitochondrial disorders. Methods: The WES data was transferred and processed by the RD-Connect Genome-Phenome Analysis Platform (GPAP) using their standardized pipeline. Variant prioritisation was carried out on the RD-Connect GPAP. Results: Singleton WES data from 14 individuals was reanalysed. We identified a possible or likely genetic diagnosis in 8 patients (8/14, 57%). The variants identified were in a combination of mitochondrial DNA (n = 1, MT-TN), nuclear encoded mitochondrial genes (n = 2, PDHA1, and SUCLA2) and nuclear genes associated with nonmitochondrial disorders (n = 5, PNPLA2, CDC40, NBAS and SLC7A7). Variants in both the NBAS and CDC40 genes were established as disease causing after the original cohort was published. We increased the diagnostic yield for the original cohort by 15% without generating any further genomic data. Conclusions: In the era of multiomics we highlight that reanalysis of existing WES data is a valid tool for generating additional diagnosis in patients with suspected mitochondrial disease, particularly when more time has passed to allow for new bioinformatic pipelines to emerge, for the development of new tools in variant interpretation aiding in reclassification of variants and the expansion of scientific knowledge on additional genes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.016 | 0.044 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.005 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".