WEARABLE SENSOR MONITORING OF WALKING ON DIFFERENT SURFACES AS A DIGITAL OUTCOME: DEEP LEARNING MODEL PERFORMANCE WITH SENSOR AND CLASS REDUCTION
Bibliographic record
Abstract
The remote monitoring of walking behavior, such as step counts, using wearables is increasingly common in research and clinical settings. However, contextless step counts have limited value for assessing orthopedic patients recovering from trauma or surgery. Instead, step analysis on specific surfaces, like stairs or slopes, is more relevant for evaluating rehabilitation progress and tailoring therapeutic decision making. Algorithms that classify walking surfaces from inertial measurement unit (IMU) signals have been developed, but the lack of standardized and practical methods for IMU data collection and analysis has hindered the creation of robust, generalizable models. This study investigates whether simplifying IMU-based gait monitoring through deep learning (DL) models—via sensor reduction (fewer sensors/signals) or surface class grouping—can maintain or improve classification performance. Data were sourced from an open-science multi-modal Gait Database (Losing et. al, 2022), comprising 20 subjects (5 female, 15 male; 18–69 years) walking on different urban surfaces wearing lycra suits embedded with 17 IMU sensors. Our baseline DL model was limited to lower-body sensors (pelvis, thighs, shanks & feet) which we deemed a feasible, but still encompassing setup. The original surface classes (n=5) included flat walking, stairs up/down, and slopes up/down. Using a previously validated Bi-Directional CNN-LSTM model with batch normalization (Vinco et. al, 2024), we tested variations which included: (a) acceleration signals only, (b) a single pelvis sensor, and (c) surface class reduction to two groups: i) by type (flat, slope, stairs) and ii) by elevation (flat, up, down). Simplifications aimed to maximize usability by enhancing patient compliance (single pelvis sensor) or battery life (acceleration-only signals). Using the baseline sensor arrangement and both the acceleration and gyroscope signals, the model achieved 84.01% accuracy (F1: 0.77–0.98). Using only acceleration, lower-body sensors reached 65.75% accuracy (F1: 0.53–0.96), and the pelvis sensor alone scored 67.08% (F1: 0.63–0.75). With surface type grouping, lower-body sensors achieved 86.86% accuracy (F1: 0.74–0.94); pelvis sensor, 65.80%. Grouping by elevation improved lower-body sensor accuracy to 89.69% (F1: 0.83–0.92) and pelvis sensor to 85.27% (F1: 0.69–0.94). High classification accuracy (>85%) was achieved, even with a single pelvis sensor and grouped classes. While sensor reduction decreased performance, grouping by elevation produced comparable results. These findings suggest practical, patient-compliant solutions for gait monitoring with potential clinical applications. Including gyroscope data appears effective towards model performance especially when many surface types need to be classified.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".