GENESIS: Gene-Specific Machine Learning Models for Variants of Uncertain Significance Found in Catecholaminergic Polymorphic Ventricular Tachycardia and Long QT Syndrome-Associated Genes
Bibliographic record
Abstract
Background: Cardiac channelopathies such as catecholaminergic polymorphic tachycardia and long QT syndrome predispose patients to fatal arrhythmias and sudden cardiac death. As genetic testing has become common in clinical practice, variants of uncertain significance (VUS) in genes associated with catecholaminergic polymorphic ventricular tachycardia and long QT syndrome are frequently found. The objective of this study was to predict pathogenicity of catecholaminergic polymorphic ventricular tachycardia-associated RYR2 VUS and long QT syndrome-associated VUS in KCNQ1 , KCNH2 , and SCN5A by developing gene-specific machine learning models and assessing them using cross-validation, cellular electrophysiological data, and clinical correlation. Methods: The GENe-specific EnSemble grId Search framework was developed to identify high-performing machine learning models for RYR2 , KCNQ1 , KCNH2 , and SCN5A using variant- and protein-specific inputs. Final models were applied to datasets of VUS identified from ClinVar and exome sequencing. Whole cell patch clamp and clinical correlation of selected VUS was performed. Results: The GENe-specific EnSemble grId Search models outperformed alternative methods, with area under the receiver operating characteristics up to 0.87, average precisions up to 0.83, and calibration slopes as close to 1.0 (perfect) as 1.04. Blinded voltage-clamp analysis of HEK293T cells expressing 2 predicted pathogenic variants in KCNQ1 each revealed an ≈80% reduction of peak Kv7.1 current compared with WT. Normal Kv7.1 function was observed in KCNQ1-V241I HEK cells as predicted. Though predicted benign, loss of Kv7.1 function was observed for KCNQ1-V106D HEK cells. Clinical correlation of 9/10 variants supported model predictions. Conclusions: Gene-specific machine learning models may have a role in post-genetic testing diagnostic analyses by providing high performance prediction of variant pathogenicity.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".