Comparability of methods for remote assessment of gait quality: An example from people with Parkinson’s Disease (PD)
Bibliographic record
Abstract
Background: Advancements in technologies for the analysis of gait enable efficient, cost-effective, and remote gait analysis that allows personalized, real-time assessments within the comfort of a person’s home. However, in clinical practice, gait analysis remains anchored to traditional observational methods that focus on performance in a controlled environment, failing to accurately assess a person’s true motor competence. For instance, a person with PD at an early stage may be able to cover an optimal distance while walking when assessed in a clinic but may not necessarily exhibit an efficient gait pattern, increasing the risk of falls. It is thus important to assess gait capacity and gait quality parameters in daily life. Although digital technologies are poised to fill this gap, their limited integration into clinical practice is due to uncertainties related to their comparability and reliability with methods used in clinical practice.Objective: The overall aim is to contribute evidence as to the comparability of 3 methods of remote gait assessment in individuals with Parkinson’s Disease. Specifically, the purpose is to estimate the extent to which values on gait metrics are similar/different among 3 different methods of assessing gait quality (1) Observational analysis by physiotherapists; (2) Wearable sensor - Heel2toeTM sensor; (3) Pose estimation – MediaPipe Pose. Secondarily, the aim is to identify challenges encountered with each of these methods.Methods: A cross-sectional, multiple case series study was conducted remotely recruiting adult members of Parkinson Quebec with mild to moderate gait deficits. After screening for eligibility, 20 participants submitted videos of them performing a modified TUG test at home/in their neighbourhood with the Heel2ToeTM sensor. A checklist for observational analysis of gait specific to PD was developed for this study. Each video was subsequently analyzed by six raters using the checklist who were allotted videos at random. The same videos were then analyzed using a customized program with the MediaPipe Pose library. Results: Crude agreement over individual items on the checklist ranged between 71-100%. Scores created by summing the item scores (maximum 35) given by the raters yielded an ICC of 0.78 indicating reliability sufficient to compare groups of people but not sufficient for within-individual change. The quality of the videos significantly affected the overall agreement, where a video with ‘excellent’ quality had an estimated score of 13.7 (0.0013) points higher than a video of poor quality. Agreement on ‘excellent’ quality videos was 96%. The values from the wearable sensor and observational ratings were compared pairwise to the inter-rater agreement results. The observational ratings agreed with the wearable sensor on accurately detecting the heel strike 64% of the time and 28.5% of the time on detecting the push-off. Agreed 35.7% of the time on detecting foot clearance and 85% of the time on detecting cadence. A Spearman’s rank correlation of 0.32 (p = 0.260) for heel strike and -0.22 (p = 0.386) for push-off was calculated from the comparison between pose estimation and wearable sensor; a correlation coefficient of -0.28 (p = 0.225) for heel strike and 0.15 (p = 0.514) for push-off was calculated from the comparison between pose estimation and observational ratings. Thus, values for heel strike and push-off obtained from pose estimation had a weak correlation with both the wearable sensor and observational ratings. Conclusion: A combination of digital technologies for remote gait analysis, such as wearable sensors and pose estimation, can detect subtle nuances in gait impairments that may be overlooked by the human eye, offering greater accuracy and reducing variability among raters
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.022 | 0.066 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".