Resolving the dark matter of <i>ABCA4</i> for 1,054 Stargardt disease probands through integrated genomics and transcriptomics
Bibliographic record
Abstract
ABSTRACT Missing heritability in human diseases represents a major challenge. Although whole-genome sequencing enables the analysis of coding and non-coding sequences, substantial costs and data storage requirements hamper its large-scale use to (re)sequence genes in genetically unsolved cases. The ABCA4 gene implicated in Stargardt disease (STGD1) has been studied extensively for 22 years, but thousands of cases remained unsolved. Therefore, single molecule molecular inversion probes were designed that enabled an automated and cost-effective sequence analysis of the complete 128-kb ABCA4 gene. Analysis of 1,054 unsolved STGD and STGD-like probands resulted in bi-allelic variations in 448 probands. Twenty-seven different causal deep-intronic variants were identified in 117 alleles. Based on in vitro splice assays, the 13 novel causal deep-intronic variants were found to result in pseudo-exon (PE) insertions (n=10) or exon elongations (n=3). Intriguingly, intron 13 variants c.1938-621G>A and c.1938-514G>A resulted in dual PE insertions consisting of the same upstream, but different downstream PEs. The intron 44 variant c.6148-84A>T resulted in two PE insertions that were accompanied by flanking exon deletions. Structural variant analysis revealed 11 distinct deletions, two of which contained small inverted segments. Uniparental isodisomy of chromosome 1 was identified in one proband. Integrated complete gene sequencing combined with transcript analysis, identified pathogenic deep-intronic and structural variants in 26% of bi-allelic cases not solved previously by sequencing of coding regions. This strategy serves as a model study that can be applied to other inherited diseases in which only one or a few genes are involved in the majority of cases.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".