Availability, Quality, and Evidence-Based Content of mHealth Apps for the Treatment of Nonspecific Low Back Pain in the German Language: Systematic Assessment
Bibliographic record
Abstract
BACKGROUND: Nonspecific low back pain (NSLBP) carries significant socioeconomic relevance and leads to substantial difficulties for those who are affected by it. The effectiveness of app-based treatments has been confirmed, and clinicians are recommended to use such interventions. As 88.8% of the German population uses smartphones, apps could support therapy. The available apps in mobile app stores are poorly regulated, and their quality can vary. Overviews of the availability and quality of mobile apps for Australia, Great Britain, and Spain have been compiled, but this has not yet been done for Germany. OBJECTIVE: We aimed to provide an overview of the availability and content-related quality of apps for the treatment of NSLBP in the German language. METHODS: A systematic search for apps on iOS and Android was conducted on July 6, 2022, in the Apple App Store and Google Play Store. The inclusion and exclusion criteria were defined before the search. Apps in the German language that were available in both stores were eligible. To check for evidence, the apps found were assessed using checklists based on the German national guideline for NSLBP and the British equivalent of the National Institute for Health and Care Excellence. The quality of the apps was measured using the Mobile Application Rating Scale. To control potential inaccuracies, a second reviewer resurveyed the outcomes for 30% (3/8) of the apps and checked the inclusion and exclusion criteria for these apps. The outcomes, measured using the assessment tools, are presented in tables with descriptive statistics. Furthermore, the characteristics of the included apps were summarized. RESULTS: In total, 8 apps were included for assessment. Features provided with different frequencies were exercise tracking of prefabricated or adaptable workout programs, educational aspects, artificial intelligence-based therapy or workout programs, and motion detection. All apps met some recommendations by the German national guideline and used forms of exercises as recommended by the National Institute for Health and Care Excellence guideline. The mean value of items rated as "Yes" was 5.75 (SD 2.71) out of 16. The best-rated app received an answer of "Yes" for 11 items. The mean Mobile Application Rating Scale quality score was 3.61 (SD 0.55). The highest mean score was obtained in "Section B-Functionality" (mean 3.81, SD 0.54). CONCLUSIONS: Available apps in the German language meet guideline recommendations and are mostly of acceptable or good quality. Their use as a therapy supplement could help promote the implementation of home-based exercise protocols. A new assessment tool to obtain ratings on apps for the treatment of NSLBP, combining aspects of quality and evidence-based best practices, could be useful. TRIAL REGISTRATION: Open Science Framework Registries sq435; https://osf.io/sq435.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.034 | 0.157 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.008 | 0.006 |
| Bibliometrics | 0.022 | 0.013 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.003 | 0.004 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".