Commercially available activity monitors such as the fitbit charge and apple watch show poor validity in patients with gait aids after total knee arthroplasty
Bibliographic record
Abstract
PURPOSE: The aim of this study is to determine the validity of consumer grade step counter devices during the early recovery period after knee replacement surgery. METHODS: Twenty-three participants wore a Fitbit Charge or Apple Watch Series 4 smart watch and performed a walking test along a 50-metre hallway. There were 9 males and 14 females included in the study with an average age of 68.5 years and BMI of 32. Each patient wore both the Fitbit Charge and Apple Watch while completing the walking test and an observer counted the ground truth value using a thumb-push tally counter. This test was repeated pre-operatively with no gait aid, immediately post operatively with a walker, at 6 weeks follow up with a cane and at 6 months with no gait aid. Bland-Altman plots were performed for all walking tests to compare the agreement between measurement techniques. RESULTS: Mean overall agreement of step count for pre-operative and at 6 months for subjects walking without gait aids was excellent for both the Apple Watch vs. actual and Fitbit vs. actual with bias values ranging from - 0.87 to 1.36 with limits of agreement (LOA) ranging between - 10.82 and 15.91. While using a walker both devices showed extremely little agreement with the actual step count with bias values between 22.5 and 24.37 with LOA between 11.7 and 33.3. At 6 weeks post-op while using a cane, both the Apple Watch and Fitbit devices had a range of bias values between - 2.8 and 5.73 with LOA between - 13.51 and 24.97. CONCLUSIONS: These devices show poor validity in the early post operative setting, especially with the use of gait aids, and therefore results should be interpreted with caution.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".