Estimating the Prevalence of GNE Myopathy Using Population Genetic Databases
Bibliographic record
Abstract
GNE myopathy (GNEM) is a rare autosomal recessive disorder characterized by progressive skeletal muscle wasting starting in early adulthood. The prevalence of GNEM is estimated to range between one and nine cases per million individuals, but the accuracy of these estimates is limited by underdiagnosis, misdiagnosis, and bias introduced by founder allele frequencies. As GNEM is a recessive disorder, unaffected carriers of single damaging variants can be expected to be found in the healthy population, providing an alternative method for estimating prevalence. We aim to estimate the prevalence of GNEM using allele frequencies obtained from healthy population genetic databases. We performed a review to establish a complete list of all known pathogenic GNEM variants from both literature and variant databases. We then developed standardized filtering steps using in silico tools to predict the pathogenicity of unreported GNE variants of uncertain clinical significance and validated our pathogenicity inferences using Mendelian Approach to Variant Effect pRedICtion built in Keras (MAVERICK) and AlphaMissense. We calculated conservative and liberal disease prevalence estimates using allele frequencies from the Genome Aggregation Database (gnomAD) population database by employing methodologies based on the assumptions of the Hardy–Weinberg Equilibrium. We additionally calculated estimates for disease prevalence removing the contribution of unique variant combinations that either do not cause myopathy in humans or result in embryonic lethality. We present the most comprehensive list of reported pathogenic GNE variants to date, together with additional variants predicted as pathogenic by in silico methods. We provide additional pathogenicity scores for these variants using new pathogenicity prediction tools and present a set of estimates for GNEM prevalence based on the different assumptions. Our most conservative estimate suggested a prevalence of 18.46 cases per million, while our most liberal estimate places the prevalence at 95.42 cases per million. When accounting for variant severity, this range drops to 11.00–87.68 cases per million. Our findings indicate that the true global prevalence of GNEM is greater than previous predictions underscoring that this condition is considerably more widespread than previously believed.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".