An Artificial Intelligence–Based App for Self-Management of Low Back and Neck Pain in Specialist Care: Process Evaluation From a Randomized Clinical Trial
Bibliographic record
Abstract
BACKGROUND: Self-management is endorsed in clinical practice guidelines for the care of musculoskeletal pain. In a randomized clinical trial, we tested the effectiveness of an artificial intelligence-based self-management app (selfBACK) as an adjunct to usual care for patients with low back and neck pain referred to specialist care. OBJECTIVE: This study is a process evaluation aiming to explore patients' engagement and experiences with the selfBACK app and specialist health care practitioners' views on adopting digital self-management tools in their clinical practice. METHODS: App usage analytics in the first 12 weeks were used to explore patients' engagement with the SELFBACK app. Among the 99 patients allocated to the SELFBACK interventions, a purposive sample of 11 patients (aged 27-75 years, 8 female) was selected for semistructured individual interviews based on app usage. Two focus group interviews were conducted with specialist health care practitioners (n=9). Interviews were analyzed using thematic analysis. RESULTS: Nearly one-third of patients never accessed the app, and one-third were low users. Three themes were identified from interviews with patients and health care practitioners: (1) overall impression of the app, where patients discussed the interface and content of the app, reported on usability issues, and described their app usage; (2) perceived value of the app, where patients and health care practitioners described the primary value of the app and its potential to supplement usual care; and (3) suggestions for future use, where patients and health care practitioners addressed aspects they believed would determine acceptance. CONCLUSIONS: Although the app's uptake was relatively low, both patients and health care practitioners had a positive opinion about adopting an app-based self-management intervention for low back and neck pain as an add-on to usual care. Both described that the app could reassure patients by providing trustworthy information, thus empowering them to take actions on their own. Factors influencing app acceptance and engagement, such as content relevance, tailoring, trust, and usability properties, were identified. TRIAL REGISTRATION: ClinicalTrials.gov NCT04463043; https://clinicaltrials.gov/study/NCT04463043.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.046 | 0.058 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.004 | 0.004 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.008 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".