Interpreting change on the Symbol Digit Modalities Test in people with relapsing multiple sclerosis using the reliable change methodology
Bibliographic record
Abstract
BACKGROUND: The Symbol Digit Modalities Test (SDMT) is increasingly utilized in clinical trials. A SDMT score change of 4 points is considered clinically important, based on association with employment anchors. Optimal thresholds for statistically reliable SDMT changes, accounting for test reliability and measurement error, are yet to be applied to individual cases. OBJECTIVE: The aim of this study was to derive a statistically reliable marker of individual change on the SDMT. METHODS: This prospective, case-control study enrolled 166 patients with multiple sclerosis (MS). SDMT scores at baseline, relapse, and 3-month follow-up were compared between relapsing and stable patient groups. Using data from the stable group and three previously published studies, candidate thresholds for reliable decline were calculated and validated against other tests and a clinically meaningful anchor-cognitive relapse. RESULTS: Candidate thresholds for reliable decline at the 80% confidence level varied between 6 and 11 points. An SDMT change of 8 or more raw score points was deemed to offer the best balance of discriminatory power and external validity for estimating cognitive decline. CONCLUSION: This study illustrates the feasibility and usefulness of reliable change methodology for identifying statistically meaningful cognitive decline that could be implemented to identify change in individual patients, for both clinical management and clinical trial outcomes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.027 | 0.050 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".