From zero- to few-shot: deep temporal learning of wrist EMG enables scalable cross-user gesture recognition
Bibliographic record
Abstract
Abstract Objective. Wrist electromyography (EMG) is emerging as an enticing wearable input modality for human–machine interaction. Traditionally recorded from the forearm for use in transradial prostheses, wrist-based EMG sensors are now being integrated into devices such as watches and wristbands for hand gesture recognition (HGR). Consumer familiarity with wrist-worn devices makes wrist EMG a compelling option, but the need for individualized user calibration remains a challenge. Approach. This study therefore evaluated various cross-user models to reduce the calibration burden and compared wrist- and forearm-based models. Eight different machine learning architectures were evaluated across 33 users, using varying amounts of data from the end user. Main results. A temporal convolutional network-bidirectional long short-term memory architecture, applied for the first time to EMG classification, was found to significantly ( p < 0.05) outperform other tested machine learning architectures. An inter-day feature set combined with Z -score normalization achieved the best performance when classifying five gestures (plus a rest class) using either wrist or forearm EMG. Consistent with other recent results, wrist EMG consistently outperformed forearm EMG in all analyses, including within- and across-user comparisons ( p < 0.05). In cross-user models, wrist EMG demonstrated a zero-shot performance of 78.2%, compared to 71.6% for forearm EMG ( p < 0.05). Introducing one calibration repetition from the end user increased one-shot performance of wrist EMG to 91.6%, compared to 86.9% for forearm EMG ( p < 0.05). Adding further training repetitions boosted wrist EMG performance to 98.3%, compared to 97.4% for forearm EMG. Significance. These findings provide new evidence supporting the viability of wrist EMG for cross-user HGR models that generalize to new users with minimal calibration, suggesting promising potential for its broader adoption in wearable devices.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".