Comparative effectiveness in multiple sclerosis: A methodological comparison
Bibliographic record
Abstract
BACKGROUND: In the absence of evidence from randomised controlled trials, observational data can be used to emulate clinical trials and guide clinical decisions. Observational studies are, however, susceptible to confounding and bias. Among the used techniques to reduce indication bias are propensity score matching and marginal structural models. OBJECTIVE: To use the comparative effectiveness of fingolimod vs natalizumab to compare the results obtained with propensity score matching and marginal structural models. METHODS: Patients with clinically isolated syndrome or relapsing remitting MS who were treated with either fingolimod or natalizumab were identified in the MSBase registry. Patients were propensity score matched, and inverse probability of treatment weighted at six monthly intervals, using the following variables: age, sex, disability, MS duration, MS course, prior relapses, and prior therapies. Studied outcomes were cumulative hazard of relapse, disability accumulation, and disability improvement. RESULTS: 4608 patients (1659 natalizumab, 2949 fingolimod) fulfilled inclusion criteria, and were propensity score matched or repeatedly reweighed with marginal structural models. Natalizumab treatment was associated with a lower probability of relapse (PS matching: HR 0.67 [95% CI 0.62-0.80]; marginal structural model: 0.71 [0.62-0.80]), and higher probability of disability improvement (PS matching: 1.21 [1.02 -1.43]; marginal structural model 1.43 1.19 -1.72]). There was no evidence of a difference in the magnitude of effect between the two methods. CONCLUSIONS: The relative effectiveness of two therapies can be efficiently compared by either marginal structural models or propensity score matching when applied in clearly defined clinical contexts and in sufficiently powered cohorts.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.008 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".