AI-Based Gait Analysis System for Rehabilitation: Usability Evaluation of Human-Technology Interaction
Bibliographic record
Abstract
BACKGROUND: Artificial intelligence (AI)-based gait analysis systems are increasingly applied in rehabilitation settings for objective and quantitative assessment of gait function. However, despite their potential, clinical adoption remains limited due to insufficient consideration of usability, user experience, and integration into actual clinical workflows. OBJECTIVE: This study aimed to conduct a formative evaluation of a prototype AI-based gait analysis system (MediStep M). METHODS: A mixed methods formative usability evaluation was conducted with 5 licensed physical therapists. Qualitative data were collected through focus group interviews, and quantitative usability was measured using the system usability scale (SUS). A scenario-based usability assessment was applied to identify user interface challenges, workflow issues, and potential design improvements. RESULTS: Participants identified major usability barriers, including limited accessibility of the power button, absence of battery status indicators, burdensome manual calibration, and insufficient clinical detail in the gait analysis reports. They also emphasized the need for wireless operation, improved portability, and integration with hospital electronic medical record systems. The mean SUS score was 57 (grade D), indicating suboptimal usability and the need for iterative design refinements. CONCLUSIONS: Although AI-based gait analysis systems hold promise for enhancing rehabilitation outcomes, key usability challenges must be resolved before clinical implementation. Improvements in hardware portability, automated calibration, data management, and user interface design are essential to ensure safety, efficiency, and clinical applicability. These findings provide evidence-based insights to guide iterative development and promote user-centered innovation in AI-based rehabilitation technologies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.011 | 0.022 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".