Usability Study of Mainstream Wearable Fitness Devices: Feature Analysis and System Usability Scale Evaluation
Bibliographic record
Abstract
BACKGROUND: Wearable devices have the potential to promote a healthy lifestyle because of their real-time data monitoring capabilities. However, device usability is a critical factor that determines whether they will be adopted on a large scale. Usability studies on wearable devices are still scarce. OBJECTIVE: This study aims to compare the functions and attributes of seven mainstream wearable devices and to evaluate their usability. METHODS: The wearable devices selected were the Apple Watch, Samsung Gear S, Fitbit Surge, Jawbone Up3, Mi Band, Huawei Honor B2, and Misfit Shine. A mixed method of feature comparison and a System Usability Scale (SUS) evaluation based on 388 participants was applied; the higher the SUS score, the better the usability of the product. RESULTS: For features, all devices had step counting, an activity timer, and distance recording functions. The Samsung Gear S had a unique sports track recording feature and the Huawei Honor B2 had a unique wireless earphone. The Apple Watch, Samsung Gear S, Jawbone Up3, and Fitbit Surge could measure heart rate. All the devices were able to monitor sleep, except the Apple Watch. For product characteristics, including attributes such as weight, battery life, price, and 22 functions such as step counting, activity time, activity type identification, sleep monitoring, and expandable new features, we found a very weak negative correlation between the SUS scores and price (r=-.10, P=.03) and devices that support expandable new features (r=-.11, P=.02), and a very weak positive correlation between the SUS scores and devices that support the activity type identification function (r=.11, P=.02). The Huawei Honor B2 received the highest score of mean 67.6 (SD 16.1); the lowest Apple Watch score was only 61.4 (SD 14.7). No significant difference was observed among brands. The SUS score had a moderate positive correlation with the user's experience (length of time the device was used) (r=.32, P<.001); participants in the medical and health care industries gave a significantly higher score (mean 61.1, SD 17.9 vs mean 68.7, SD 14.5, P=.03). CONCLUSIONS: The functions of wearable devices tend to be homogeneous and usability is similar across various brands. Overall, Mi Band had the lowest price and the lightest weight. Misfit Shine had the longest battery life and most functions, and participants in the medical and health care industries had the best evaluation of wearable devices. The perceived usability of mainstream wearable devices is unsatisfactory and customer loyalty is not high. A consumer's SUS rating for a wearable device is related to their personal situation instead of the device brand. Device manufacturers should put more effort into developing innovative functions and improving the usability of their products by integrating more cognitive behavior change techniques.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.003 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".