Comprehensive prediction of mRNA splicing effects of BRCA1 and BRCA2 variants
Bibliographic record
Abstract
Variants of uncertain significance (VUS) in the BRCA1 and BRCA2 genes potentially affecting coding sequence as well as normal splicing activity have confounded predisposition testing in breast cancer. Here, we apply information theory to analyze BRCA1/2 mRNA splicing mutations categorized as VUS. The method was validated for 31 of 36 mutations known to cause missplicing in BRCA1/2 and all 26 that do not alter splicing. All single-nucleotide variants in the Breast Cancer Information Resource (BIC; Breast Cancer Information Core Database; http://research.nhgri.nih.gov/bic; last access June 1, 2010) were then analyzed. Information analysis is similar in sensitivity to other predictive methods; however, the thermodynamic basis of the theory also enables splice-site affinity to be determined accurately, which is important for assessing mutations that render natural splice sites partially functional and competition between cryptic and natural splice sites. We report 299 of 2,071 single-nucleotide BIC mutations that are predicted to significantly weaken natural sites and/or strengthen cryptic splice sites, 171 of which are not designated as splicing mutations in the database. Splicing alterations are predicted for 68 of 690 BRCA1 and 60 of 958 BRCA2 mutations designated as VUS. These analyses should be useful in prioritizing suspected mutations for downstream expression studies and for predicting aberrantly spliced isoforms generated by these mutations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".