MRI-based clinical trials in relapsing–remitting MS: new sample size calculations based on a longitudinal model
Bibliographic record
Abstract
BACKGROUND: Sample sizes for magnetic resonance imaging (MRI)-based clinical trials in multiple sclerosis (MS) generally assume that lesion counts are reasonably described by the negative binomial (NB) model. OBJECTIVE: This study aimed to assess the appropriateness of the NB model for lesion count data and to provide sample sizes for placebo-controlled, MRI-based clinical trials in relapsing-remitting MS using a more realistic model. METHODS: The fit of the NB model in each arm of five MS clinical trials was assessed using Pearson's chi-squared statistic. Required sample sizes associated with various tests of treatment effect were estimated by simulating data from a new, longitudinal model for repeated lesion count data on individual patients. RESULTS: Evidence (p < 0.05) against the NB model was found in at least one arm of four of the five trials. If a trial is designed using this model but the resulting clinical data do not follow its assumptions then this trial can be seriously under-powered for assessing differences in mean lesion counts. CONCLUSION: Sample sizes based on the longitudinal model are more realistic and often smaller than those previously reported using the NB model.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.169 | 0.275 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.002 |
| Scholarly communication | 0.001 | 0.003 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".