MRI-based clinical trials in relapsing–remitting MS: new sample size calculations based on a longitudinal model
Bibliographic record
Abstract
BACKGROUND: Sample sizes for magnetic resonance imaging (MRI)-based clinical trials in multiple sclerosis (MS) generally assume that lesion counts are reasonably described by the negative binomial (NB) model. OBJECTIVE: This study aimed to assess the appropriateness of the NB model for lesion count data and to provide sample sizes for placebo-controlled, MRI-based clinical trials in relapsing-remitting MS using a more realistic model. METHODS: The fit of the NB model in each arm of five MS clinical trials was assessed using Pearson's chi-squared statistic. Required sample sizes associated with various tests of treatment effect were estimated by simulating data from a new, longitudinal model for repeated lesion count data on individual patients. RESULTS: Evidence (p < 0.05) against the NB model was found in at least one arm of four of the five trials. If a trial is designed using this model but the resulting clinical data do not follow its assumptions then this trial can be seriously under-powered for assessing differences in mean lesion counts. CONCLUSION: Sample sizes based on the longitudinal model are more realistic and often smaller than those previously reported using the NB model.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.020 | 0.146 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".