Validation of a Musculoskeletal Digital Assessment Routing Tool: Protocol for a Pilot Randomized Crossover Noninferiority Trial
Bibliographic record
Abstract
BACKGROUND: Musculoskeletal conditions account for 16% of global disability, resulting in a negative effect on millions of patients and an increasing demand for health care use. Digital technologies to improve health care outcomes and efficiency are considered a priority; however, innovations are rarely tested with sufficient rigor in clinical trials, which is the gold standard for clinical proof of safety and efficacy. We have developed a new musculoskeletal digital assessment routing tool (DART) that allows users to self-assess and be directed to the right care. DART requires validation in a real-world setting before implementation. OBJECTIVE: This pilot study aims to assess the feasibility of a future trial by exploring the key aspects of trial methodology, assessing the procedures, and collecting exploratory data to inform the design of a definitive randomized crossover noninferiority trial to assess DART safety and effectiveness. METHODS: We will collect data from 76 adults with a musculoskeletal condition presenting to general practitioners within a National Health Service (NHS) in England. Participants will complete both a DART assessment and a physiotherapist-led triage, with the order determined by randomization. The primary analysis will involve an absolute agreement intraclass correlation (A,1) estimate with 95% CI between DART and the clinician for assessment outcomes signposting to condition management pathways. Data will be collected to allow the analysis of participant recruitment and retention, randomization, allocation concealment, blinding, data collection process, and bias. In addition, the impact of trial burden and potential barriers to intervention delivery will be considered. The DART user satisfaction will be measured using the system usability scale. RESULTS: A UK NHS ethics submission was done during June 2021 and is pending approval; recruitment will commence in early 2022, with data collection anticipated to last for 3 months. The results will be reported in a follow-up paper in 2022. CONCLUSIONS: This study will inform the design of a randomized controlled crossover noninferiority study that will provide evidence concerning mobile health DART system clinical signposting in an NHS setting before real-world implementation. Success should produce evidence of a safe, effective system with good usability, potentially facilitating quicker and easier patient access to appropriate care while reducing the burden on primary and secondary care musculoskeletal services. This rigorous approach to mobile health system testing could be used as a guide for other developers of similar applications. TRIAL REGISTRATION: ClinicalTrials.gov NCT04904029; http://clinicaltrials.gov/ct2/show/NCT04904029. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): PRR1-10.2196/31541.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".