Evaluating the Accuracy and Reliability of Real-World Digital Mobility Outcomes in Older Adults After Hip Fracture: Cross-Sectional Observational Study
Bibliographic record
Abstract
BACKGROUND: Algorithms estimating real-world digital mobility outcomes (DMOs) are increasingly validated in healthy adults and various disease cohorts. However, their accuracy and reliability in older adults after hip fracture, who often walk slowly for short durations, is underexplored. OBJECTIVE: This study examined DMO accuracy and reliability in a hip fracture cohort considering walking bout (WB) duration, physical function, days since surgery, and walking aid use. METHODS: In total, 19 community-dwelling participants were real-world monitored for 2.5 hours using a lower back wearable device and a reference system combining inertial modules, distance sensors, and pressure insoles. A total of 6 DMO estimates from 164 WBs from 58% (11/19) of the participants (aged 71-90 years; assessed 32-390 days after surgery; Short Physical Performance Battery [SPPB] scores of 3-12; gait speed range 0.39-1.34 m/s) were assessed against the reference system at the WB and participant level. We stratified by WB duration (all WBs, WBs of >10 seconds, WBs of 10-30 seconds, and WBs of >30 seconds) and lower versus higher SPPB scores and observed whether days since surgery and walking aid use affected DMO accuracy and reliability. RESULTS: Across WBs, walking speed and distance ranged from 0.25 to 1.29 m/s and from 1.7 to 436.5 m, respectively. Estimation of walking speed, cadence, stride duration, number of steps, and distance stratified by WB duration showed intraclass correlation coefficients (ICCs) ranging from 0.50 to 0.99 and mean relative errors (MREs) from -6.9% to 12.8%. Stride length estimation showed poor reliability, with ICCs ranging from 0.30 to 0.49 and MREs from 6.1% to 13.2%. Walking speed and distance ICCs in the higher-SPPB score group ranged from 0.85 to 0.99, and MREs ranged from -10.1% to -1.7%. In the lower-SPPB score group, walking speed and distance ICCs ranged from 0.17 to 0.99, and MREs ranged from 13.5% to 32.6%. There was no discernible effect of time since surgery or walking aid use. CONCLUSIONS: In total, 5 accurate and reliable real-world DMOs were identified in older adults after hip fracture: walking speed, cadence, stride duration, number of steps, and distance. Accuracy and reliability of most DMOs improved when excluding WBs of <10 seconds and were higher for WBs of >30 seconds than for WBs of 10 to 30 seconds and for participants with higher physical function. DMOs capture daily gait as early as 1 month after surgery also in people using walking aids. However, as most WBs in this cohort were short, there was a trade-off between improving accuracy and reliability by excluding short WBs and losing a substantial amount of data. These results have important implications for establishing the clinical validity of DMOs and evaluating the effects of interventions on daily-life gait, thereby facilitating the design of optimal care pathways.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.012 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".