A Physician-Completed Digital Tool for Evaluating Disease Progression (Multiple Sclerosis Progression Discussion Tool): Validation Study
Bibliographic record
Abstract
BACKGROUND: Defining the transition from relapsing-remitting multiple sclerosis (RRMS) to secondary progressive multiple sclerosis (SPMS) can be challenging and delayed. A digital tool (MSProDiscuss) was developed to facilitate physician-patient discussion in evaluating early, subtle signs of multiple sclerosis (MS) disease progression representing this transition. OBJECTIVE: This study aimed to determine cut-off values and corresponding sensitivity and specificity for predefined scoring algorithms, with or without including Expanded Disability Status Scale (EDSS) scores, to differentiate between RRMS and SPMS patients and to evaluate psychometric properties. METHODS: Experienced neurologists completed the tool for patients with confirmed RRMS or SPMS and those suspected to be transitioning to SPMS. In addition to age and EDSS score, each patient's current disease status (disease activity, symptoms, and its impacts on daily life) was collected while completing the draft tool. Receiver operating characteristic (ROC) curves determined optimal cut-off values (sensitivity and specificity) for the classification of RRMS and SPMS. RESULTS: Twenty neurologists completed the draft tool for 198 patients. Mean scores for patients with RRMS (n=89), transitioning to SPMS (n=47), and SPMS (n=62) were 38.1 (SD 12.5), 55.2 (SD 11.1), and 69.6 (SD 12.0), respectively (P<.001, each between-groups comparison). Area under the ROC curve (AUC) including and excluding EDSS were for RRMS (including) AUC 0.91, 95% CI 0.87-0.95, RRMS (excluding) AUC 0.88, 95% CI 0.84-0.93, SPMS (including) AUC 0.91, 95% CI 0.86-0.95, and SPMS (excluding) AUC 0.86, 95% CI 0.81-0.91. In the algorithm with EDSS, the optimal cut-off values were ≤51.6 for RRMS patients (sensitivity=0.83; specificity=0.82) and ≥58.9 for SPMS patients (sensitivity=0.82; specificity=0.84). The optimal cut-offs without EDSS were ≤46.3 and ≥57.8 and resulted in similar high sensitivity and specificity (0.76-0.86). The draft tool showed excellent interrater reliability (intraclass correlation coefficient=.95). CONCLUSIONS: The MSProDiscuss tool differentiated RRMS patients from SPMS patients with high sensitivity and specificity. In clinical practice, it may be a useful tool to evaluate early, subtle signs of MS disease progression indicating the evolution of RRMS to SPMS. MSProDiscuss will help assess the current level of progression in an individual patient and facilitate a more informed physician-patient discussion.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.053 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".