CAG size-specific risk estimates for intermediate allele repeat instability in Huntington disease
Bibliographic record
Abstract
INTRODUCTION: New mutations for Huntington disease (HD) occur due to CAG repeat instability of intermediate alleles (IA). IAs have between 27 and 35 CAG repeats, a range just below the disease threshold of 36 repeats. While they usually do not confer the HD phenotype, IAs are prone to paternal germline CAG repeat instability. Consequently, they may expand into the HD range upon transmission to the next generation, producing a new mutation. Quantified risk estimates for IA repeat instability are extremely limited but needed to inform clinical practice. METHODS: Using small-pool PCR of sperm DNA from Caucasian men, we examined the frequency and magnitude of CAG repeat instability across the entire range of intermediate CAG sizes. The CAG size-specific risk estimates generated are based on the largest sample size ever examined, including 30 IAs and 18 198 sperm. RESULTS: Our findings demonstrate a significant risk of new mutations. While all intermediate CAG sizes demonstrated repeat expansion into the HD range, alleles with 34 and 35 CAG repeats were associated with the highest risk of a new mutation (2.4% and 21.0%, respectively). IAs with ≥33 CAG repeats showed a dramatic increase in the frequency of instability and a switch towards a preponderance of repeat expansions over contractions. CONCLUSIONS: These data provide novel insights into the origins of new mutations for HD. The CAG size-specific risk estimates inform clinical practice and provide accurate risk information for persons who receive an IA predictive test result.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.022 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".