MétaCan
Menu
Back to cohort
Record W4409169824 · doi:10.2196/67792

Evaluating the Accuracy and Reliability of Real-World Digital Mobility Outcomes in Older Adults After Hip Fracture: Cross-Sectional Observational Study

2025· article· en· W4409169824 on OpenAlexvenueno aff
Martin Aursand Berge, Anisoara Paraschiv-Ionescu, Cameron Kirk, Arne Küderle, M. Encarna Micó-Amigo, Clemens Becker, Andrea Cereatti, Silvia Del Din, Monika Engdal, Judith García‐Aymerich, Karoline Blix Grønvik, Clint Hansen, Jeffrey M. Hausdorff, Jorunn L. Helbostad, Carl-Philipp Jansen, Lars Gunnar Johnsen, Jochen Klenk, Sarah Koch, Walter Maetzler, Dimitrios Megaritis, Arne Mueller, Lynn Rochester, Lars Schwickert, Kristin Taraldsen, Beatrix Vereijken

Bibliographic record

VenueJMIR Formative Research · 2025
Typearticle
Languageen
FieldMedicine
TopicHip and Femur Fractures
Canadian institutionsnot available
Fundersnot available
KeywordsPreprintReliability (semiconductor)Hip fractureCross-sectional studyMedicineGerontologyComputer scienceOsteoporosisWorld Wide Web

Abstract

fetched live from OpenAlex

BACKGROUND: Algorithms estimating real-world digital mobility outcomes (DMOs) are increasingly validated in healthy adults and various disease cohorts. However, their accuracy and reliability in older adults after hip fracture, who often walk slowly for short durations, is underexplored. OBJECTIVE: This study examined DMO accuracy and reliability in a hip fracture cohort considering walking bout (WB) duration, physical function, days since surgery, and walking aid use. METHODS: In total, 19 community-dwelling participants were real-world monitored for 2.5 hours using a lower back wearable device and a reference system combining inertial modules, distance sensors, and pressure insoles. A total of 6 DMO estimates from 164 WBs from 58% (11/19) of the participants (aged 71-90 years; assessed 32-390 days after surgery; Short Physical Performance Battery [SPPB] scores of 3-12; gait speed range 0.39-1.34 m/s) were assessed against the reference system at the WB and participant level. We stratified by WB duration (all WBs, WBs of >10 seconds, WBs of 10-30 seconds, and WBs of >30 seconds) and lower versus higher SPPB scores and observed whether days since surgery and walking aid use affected DMO accuracy and reliability. RESULTS: Across WBs, walking speed and distance ranged from 0.25 to 1.29 m/s and from 1.7 to 436.5 m, respectively. Estimation of walking speed, cadence, stride duration, number of steps, and distance stratified by WB duration showed intraclass correlation coefficients (ICCs) ranging from 0.50 to 0.99 and mean relative errors (MREs) from -6.9% to 12.8%. Stride length estimation showed poor reliability, with ICCs ranging from 0.30 to 0.49 and MREs from 6.1% to 13.2%. Walking speed and distance ICCs in the higher-SPPB score group ranged from 0.85 to 0.99, and MREs ranged from -10.1% to -1.7%. In the lower-SPPB score group, walking speed and distance ICCs ranged from 0.17 to 0.99, and MREs ranged from 13.5% to 32.6%. There was no discernible effect of time since surgery or walking aid use. CONCLUSIONS: In total, 5 accurate and reliable real-world DMOs were identified in older adults after hip fracture: walking speed, cadence, stride duration, number of steps, and distance. Accuracy and reliability of most DMOs improved when excluding WBs of <10 seconds and were higher for WBs of >30 seconds than for WBs of 10 to 30 seconds and for participants with higher physical function. DMOs capture daily gait as early as 1 month after surgery also in people using walking aids. However, as most WBs in this cohort were short, there was a trade-off between improving accuracy and reliability by excluding short WBs and losing a substantial amount of data. These results have important implications for establishing the clinical validity of DMOs and evaluating the effects of interventions on daily-life gait, thereby facilitating the design of optimal care pathways.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.005
metaresearch head score (Gemma)0.012
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.005
Threshold uncertainty score0.024

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0050.012
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.001
Science and technology studies0.0000.000
Scholarly communication0.0010.001
Open science0.0010.001
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.124
GPT teacher head0.521
Teacher spread0.397 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations5
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueJMIR Formative ResearchSame topicHip and Femur FracturesFrench-language works237,207