Clinical applications of and molecular insights from RNA sequencing in a rare disease cohort
Bibliographic record
Abstract
BACKGROUND: RNA sequencing (RNA-seq) is emerging as a valuable tool for identifying disease-causing RNA transcript aberrations that cannot be identified by DNA-based testing alone. Previous studies demonstrated some success in utilizing RNA-seq as a first-line test for rare inborn genetic conditions. However, DNA-based testing (increasingly, whole genome sequencing) remains the standard initial testing approach in clinical practice. The indications for RNA-seq after a patient has undergone DNA-based sequencing remain poorly defined, which hinders broad implementation and funding/reimbursement. METHODS: In this study, we identified four specific and familiar clinical scenarios, and investigated in each the diagnostic utility of RNA-seq on clinically accessible tissues: (i) clarifying the impact of putative intronic or exonic splice variants (outside of the canonical splice sites), (ii) evaluating canonical splice site variants in patients with atypical phenotypes, (iii) defining the impact of an intragenic copy number variation on gene expression, and (iv) assessing variants within regulatory elements and genic untranslated regions. RESULTS: These hypothesis-driven RNA-seq analyses confirmed a molecular diagnosis and pathomechanism for 45% of participants with a candidate variant, provided supportive evidence for a DNA finding for another 21%, and allowed us to exclude a candidate DNA variant for an additional 24%. We generated evidence that supports two novel Mendelian gene-disease associations (caused by variants in PPP1R2 and MED14) and several new disease mechanisms, including the following: (1) a splice isoform switch due to a non-coding variant in NFU1, (2) complete allele skew from a transcriptional start site variant in IDUA, and (3) evidence of a germline gene fusion of MAMLD1-BEND2. In contrast, RNA-seq in individuals with suspected rare inborn genetic conditions and negative whole genome sequencing yielded only a single new potential diagnostic finding. CONCLUSIONS: In summary, RNA-seq had high diagnostic utility as an ancillary test across specific real-world clinical scenarios. The findings also underscore the ability of RNA-seq to reveal novel disease mechanisms relevant to diagnostics and treatment.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".