Simulation of Clinical Visits as a Novel Approach to Evaluate Digital Health in Multiple Sclerosis: Simulation Study
Bibliographic record
Abstract
Background: Efforts are being made to integrate digital health technologies into clinical care for multiple sclerosis (MS) to improve patient monitoring. Efficiently probing how they might impact clinical care could streamline digital tool development. The Floodlight digital tool, comprising 5 smartphone sensor-based tests, was used to generate health-related data on patient function and symptoms in a clinical simulation. Objective: The study had 3 objectives: (1) assess the utility of simulated clinical encounters as a research methodology for exploring the introduction of digital health technologies into clinical practice in MS, (2) confirm the fidelity of the simulated environment and patient cases developed and understand what metrics (eg, workflow, comprehensive evaluation) could be generated, and (3) generate insights into the utility of digitally collected data, including usability, clinical decision contribution, and impact on workflows, in clinical practice. Methods: A total of 2 patient cases consisting of clinical, radiological, and digital health data were developed with clinician input. US-based neurologists prepared for and conducted 2 simulated teleconsultations each, with an actor briefed on case profiles. Floodlight data were available, via the Floodlight MS™ Health Care Professional Portal, for 1 of the 2 consultations. Participant neurologists completed interviews and surveys assessing the fidelity of the cases presented, user experience and workflow metrics, patient concerns identified, care decisions made, and confidence in making decisions. Results: All 10 neurologists indicated that the simulations were high-fidelity representations of real consultations. Using the Floodlight technology for the first time, median time taken to prepare for and conduct the consultation was ~1.7-2 minutes longer, with slightly greater mental effort reported by participants, compared with not using the tool. The Floodlight MS Health Care Professional Portal scored an "above average" 79 on the System Usability Scale and an "acceptable" Net Promoter Score of 10. In total, 6 of the 10 neurologists "strongly agreed" that it was easier and quicker to identify patient concerns when they had access to the patient-generated Floodlight data to prepare for their encounters than when they did not. Overall, more care and management decisions were taken when the digital tool was used (37 vs 29). Of those 37 decisions, Floodlight data were reported as a trigger for 20 decisions, always in combination with other elements including patient history (20/20) and clinical exam findings (9/20). Conclusions: These findings advance our understanding of clinical simulation as a method for evaluating digital tools and other innovative technologies for MS care. High-fidelity patient cases could be provided for the mock teleconsultations, and the simulated clinical environment was useful for evaluating usability and utility of a new digital tool-yielding preliminary evidence on how digital data could be accessed and utilized by neurologists to support routine MS care.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.030 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".