Remote Testing Apps for Multiple Sclerosis Patients: Scoping Review of Published Articles and Systematic Search and Review of Public Smartphone Apps
Bibliographic record
Abstract
Background: Many apps have been designed to remotely assess clinical status and monitor symptom evolution in persons with multiple sclerosis (MS). These may one day serve as an adjunct for in-person assessment of persons with MS, providing valuable insight into the disease course that is not well captured by cross-sectional snapshots obtained from clinic visits. Objective: This study sought to review the current literature surrounding apps used for remote monitoring of persons with MS. Methods: A scoping review of published articles was conducted to identify and evaluate the literature published regarding the use of apps for monitoring of persons with MS. PubMed/Medline, EMBASE, CINAHL, and Cochrane databases were searched from inception to January 2022. Cohort studies, feasibility studies, and randomized controlled trials were included in this review. All pediatric studies, single case studies, poster presentations, opinion pieces, and commentaries were excluded. Studies were assessed for risk of bias using the Scottish Intercollegiate Guidelines Network, when applicable. Key findings were grouped in categories (convergence to neurological exam, feasibility of implementation, impact of weather, and practice effect), and trends are presented. In a parallel systematic search, the Canadian Apple App Store and Google Play Store were searched to identify relevant apps that are available but have yet to be formally studied and published in peer-reviewed publications. Results: We included 18 articles and 18 apps. Although many MS-related apps exist, only 10 apps had published literature supporting their use. Convergence between app-based testing and the neurological exam was examined in 12 articles. Most app-based tests focused on physical disability and cognition, although other domains such as ambulation, balance, visual acuity, and fatigue were also evaluated. Overall, correlations between the app versions of standardized tests and their traditional counterparts were moderate to strong. Some novel app-based tests had a stronger correlation with clinician-derived outcomes than traditional testing. App-based testing correlated well with the Multiple Sclerosis Functional Composite but less so with the Expanded Disability Status Scale; the latter correlated to a greater extent with patient quality of life questionnaire scores. Conclusions: Although limited by a small number of included studies and study heterogeneity, the findings of this study suggest that app-based testing demonstrates adequate convergence to traditional in-person assessment and may be used as an adjunct to and perhaps in lieu of specific neurological exam metrics documented at clinic visits, particularly if the latter is not readily accessible for persons with MS.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.070 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.006 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".