I14 Predicting individual disease progression in huntington’s disease using mixed models and transition models
Bibliographic record
Abstract
Background and aim We aimed to compare different statistical methods for predicting clinical progression, among individuals with Huntington’s disease (HD). Methods We compared the following methods: (1) Multiple Linear Regression, (2) Linear Mixed Models with different covariance structures (no correlation, autoregressive(AR)), and (3) Transition Models (Panel Models). We built models to predict the most commonly used clinical trial outcomes: (1) motor function (measured by the Total Motor Score on the UHDRS) and (2) daily function (measured by the Total Functional Capacity). Potential predictors considered included Cytosine-Adenine-Guanine repeat expansion length, age of onset, and years since diagnosis. We used genetic and longitudinal clinical data collected for the Cooperative Observational Research Trial (COHORT) study, led by the Huntington Study Group (United States, Canada, and Australia, 2006–2011). Data was randomly split into a training set (2/3) and testing set (1/3). We selected predictors for models using the Bayesian Information Criterion. We compared the predictive accuracy of different models using the Mean Squared Prediction Error. Results Overall, models that took into account longitudinal functional scores for individuals (Mixed Models with AR covariate structure, and Transition Model) performed better than Multiple Linear Regression. However the Prediction Intervals were wide. Under 17% of participants had data for four or more visits, limiting our ability to use higher order AR structures for the Mixed and Transition Models. Conclusions Mixed and Transition models improved prediction of individual disease progression in HD compared to Multiple Linear Regression, although current level of predictive accuracy is not high enough for clinical use.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.022 | 0.034 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.005 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".