Usability Assessment Methods for Mobile Apps for Physical Rehabilitation: Umbrella Review
Bibliographic record
Abstract
BACKGROUND: Usability has been touted as one determiner of success of mobile health (mHealth) interventions. Multiple systematic reviews of usability assessment approaches for different mHealth solutions for physical rehabilitation are available. However, there is a lack of synthesis in this portion of the literature, which results in clinicians and developers devoting a significant amount of time and effort in analyzing and summarizing a large body of systematic reviews. OBJECTIVE: This study aims to summarize systematic reviews examining usability assessment instruments, or measurements tools, in mHealth interventions including physical rehabilitation. METHODS: An umbrella review was conducted according to a published registered protocol. A topic-based search of PubMed, Cochrane, IEEE Xplore, Epistemonikos, Web of Science, and CINAHL Complete was conducted from January 2015 to April 2023 for systematic reviews investigating usability assessment instruments in mHealth interventions including physical exercise rehabilitation. Eligibility screening included date, language, participant, and article type. Data extraction and assessment of the methodological quality (AMSTAR 2 [A Measurement Tool to Assess Systematic Reviews 2]) was completed and tabulated for synthesis. RESULTS: A total of 12 systematic reviews were included, of which 3 (25%) did not refer to any theoretical usability framework and the remaining (n=9, 75%) most commonly referenced the ISO framework. The sample referenced a total of 32 usability assessment instruments and 66 custom-made, as well as hybrid, instruments. Information on psychometric properties was included for 9 (28%) instruments with satisfactory internal consistency and structural validity. A lack of reliability, responsiveness, and cross-cultural validity data was found. The methodological quality of the systematic reviews was limited, with 8 (67%) studies displaying 2 or more critical weaknesses. CONCLUSIONS: There is significant diversity in the usability assessment of mHealth for rehabilitation, and a link to theoretical models is often lacking. There is widespread use of custom-made instruments, and preexisting instruments often do not display sufficient psychometric strength. As a result, existing mHealth usability evaluations are difficult to compare. It is proposed that multimethod usability assessment is used and that, in the selection of usability assessment instruments, there is a focus on explicit reference to their theoretical underpinning and acceptable psychometric properties. This could be facilitated by a closer collaboration between researchers, developers, and clinicians throughout the phases of mHealth tool development. TRIAL REGISTRATION: PROSPERO CRD42022338785; https://www.crd.york.ac.uk/prospero/#recordDetails.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.101 | 0.291 |
| Meta-epidemiology (narrow) | 0.004 | 0.003 |
| Meta-epidemiology (broad) | 0.011 | 0.017 |
| Bibliometrics | 0.051 | 0.027 |
| Science and technology studies | 0.003 | 0.004 |
| Scholarly communication | 0.010 | 0.010 |
| Open science | 0.004 | 0.007 |
| Research integrity | 0.004 | 0.004 |
| Insufficient payload (model declined to judge) | 0.008 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".